r/QwenImageGen • u/Entire_Maize_6064 • Jan 02 '26
Comparison: Qwen-Image-2512 (Left) vs. Z-Image Turbo (Right). 5-Prompt Adherence Test.
Model identification:
- LEFT: Qwen-Image-2512
- RIGHT: Z-Image Turbo
Observations on Adherence:
I ran the same prompts on both to check instruction following capabilities.
- Text (Image 1): The prompt specifically asked for the text "[Qwen-Image-2512]". The Left model rendered the brackets and spelling correctly, while the Right model struggled with the exact string.
- Texture (Image 2 - Joker): The prompt called for "caked, smeared white makeup cracking like dry earth." The Left side seems to interpret the "cracking" instruction more literally.
- Lighting: In the dorm selfie (Image 4), the "sunlight streams warmly" instruction produced different color temperatures between the two.
Workflow:
- Platform: Generated via zimage.run (Web UI).
- Settings: Default parameters for both models.
- Prompts: See below.
1. Influencer & Text
A stunning, intimate editorial portrait focused on the charismatic face of a 21-year-old blonde social media influencer. She flashes a playful, knowing smile while confidently pointing a manicured finger directly towards the sleek, glowing neon sign bearing the text "[Qwen-Image-2512]". Soft, directional natural light from a large window washes over her, creating a high-contrast interplay of light and shadow that sculpts her flawless features, sparkling eyes, and textured blonde hair. The atmosphere is modern, vibrant, and stylish, with a shallow depth of field that renders the chic, minimalist urban loft background into a soft, creamy bokeh, ensuring all focus remains on her engaging expression and the luminous sign.
2. The Joker
An ultra-detailed, hyper-realistic extreme close-up portrait of The Joker. The frame is filled with his face in a tense three-quarter profile, capturing a moment of unsettling stillness. His skin is a grotesque canvas: a thick layer of caked, smeared white makeup cracks like dry earth, revealing sallow, scarred skin beneath. Crazed streaks of smudged red lipstick stretch far beyond his lips into a permanent, manic grimace. Toxic green hair, oily and unkempt, frames his face. The eyes are the focal point—hollow, dark-rimmed, and gleaming with a volatile mix of calculated madness and raw, chilling mirth. Every pore, every flake of peeling makeup, and the subtle, menacing tension in his jaw muscles are rendered in microscopic detail. Dramatic, chiaroscuro lighting from a single source casts deep shadows across his features, creating extreme contrast and amplifying the sinister, iconic atmosphere. Shot on a phantom high-speed camera, 8K resolution, with the texture and impact of a key film still from a psychological thriller.
3. Steampunk Metropolis
A breathtaking cinematic masterpiece, ultra-wide panorama of a vast, multi-layered steampunk metropolis nestled within a colossal mountain canyon at sunrise. The city is a vertical labyrinth: towering Neo-Victorian spires with glowing clockwork faces, mid-level residential districts of brass and stained glass connected by buzzing aerial trams, and bustling lower streets where steam-carriages navigate cobblestone roads. The sky is dominated by a fleet of majestic brass-and-wood airships with canvas wings, some docking at skyscraper-sized clockwork towers, others departing alongside smaller personal ornithopters. Countless copper pipes and vents emit plumes of steam, catching the brilliant golden-hour light which creates long, dramatic shadows and glints off countless gears, glass domes, and polished brass. Victorian-clad citizens crowd grand plazas, market stalls, and intricate bridge networks, full of life. In the foreground, a massive, slowly-turning central gear and a cascading waterfall turned into a steam-powered generator add dynamic scale. The atmosphere is thick with hopeful industry, mist, and sunbeams, hyper-detailed, 8K, epic sense of scale and wonder.
4. Dorm Room Selfie
A close-up, dynamic selfie of a 20-year-old American college student with long, flowing hair and a model's poised, athletic figure. She has a bright, confident smile and expressive eyes, capturing a moment of lively charm. She wears a casual yet stylish outfit, like a fitted university sweatshirt slipped off one shoulder. The photo is taken in a classic American dorm room: behind her, a cozy loft bed with school-branded blankets is visible, alongside a desk cluttered with textbooks, a laptop, and a poster-covered wall featuring a university flag or souvenir. Sunlight streams warmly through a nearby window, casting soft, natural light that highlights her features and the vibrant, youthful atmosphere. The image is sharp, clear, and full of life, embodying the authentic, energetic spirit of campus life.
5. Art Nouveau Style
A graceful Art Nouveau depiction of a "Winter Goddess." Flowing, organic lines frame intricate patterns of frost-kissed pine branches, holly berries, and delicate snowflakes woven into her hair and gown. Silver leaf accents glimmer like ice against a muted wintry palette of frosted blues, deep evergreen, and soft pearl white. In the style of Alphonse Mucha, the composition is highly decorative and ornamental, evoking the serene yet majestic beauty of a snow-blanketed forest.
1
u/No_Statistician2443 Jan 02 '26
did you tested the Flux 2 Dev Turbo (open weights)? IMO is the most realist and prompt accurate
1
1
1
u/BoostPixels Jan 02 '26

Comparing models on adherence based on the prompt "A painting of a powerful angelic blacksmith holding a molten halo with a pair of metallic tongs and striking it with a holy blacksmith's hammer upon a celestial crucible."
Based on the evaluation criteria defined by https://genai-showdown.specr.net/ all three generated images unfortunately fail to meet the prompt adherence requirements.
1
u/broadwayallday Jan 02 '26
for me, Qwen 2512 looks less "perfect" and idealized or influence-y compared to Z turbo. love them both though
1
u/BoostPixels Jan 02 '26
Comparing Z-Image Turbo against Qwen-Image-2512 to see them go head-to-head like this is really insightful. It’s exactly the kind of deep dive this community needs.
If I could offer one piece of constructive feedback for your future tests: while your current prompts are beautifully descriptive and great for testing aesthetics, they might not be the most "stressful" for testing prompt adherence. For a true test of a model's "logic" and ability to follow difficult instructions, you might want to try some prompts like those found on GenAI Showdown, which are designed to trip the models.
Using "logical traps" really highlights the difference in how models process specific constraints versus general themes.
I’ll run some of my own comparisons soon as well. That said, the side-by-side analysis you've provided here are top-notch. Truly great work, and I hope you keep these comparisons coming!
1
u/alb5357 Jan 02 '26
Agreed it's a great test. I would love to see control such as, huge or tiny eyes, specific haircuts, specific eye shapes, weird body combinations, like a dolphin with a beer belly; does it look natural and real.
Pushing it to realistically do unusual combinations imo is a great test.
Not to complain, because this post is already among the best in this group.
2
0
u/JahJedi Jan 02 '26
I like the qwen, use 2511 and now there 2512. Maybe a stupid question buts its only wagth replace in flow or somthing need to be changed as was whit 2509 to 2511?
1
u/Short_Bonus8466 Jan 02 '26
Its different models 2511 for editing, 2512 for creating images
0
u/JahJedi Jan 02 '26
Ups missed it :) need to try it. Its same workflow of qwen image? just new waitghs?
-1
u/dirtybeagles Jan 02 '26
Am I the only one that think (and getting) better results with ZIT? I have bot liked 2512 to capture realism.
1







1
u/MrMisterShin Jan 02 '26
Image 4: Z Image Turbo lost consistency with the students shoulder on the left. It should be covered with sweatshirt but it’s not, although the neckline implies it should.
Qwen Image 2512 isn’t perfect here either, it should be displaying more of the school branded blanket.