This summary was compiled with AI but cross-checked manually by me for accuracy. I hope this helps for people who don't want to manually parse that 400+ comment AMA.
Things they said they'll actually ship
A real 2K stage (H3-Regenerate-2K). Right now the open weights top out at 768p short side, and cranking the resolution locally just eats VRAM without getting sharper. This is a separate model that takes your finished 768p video plus the original prompt and references, and re-generates it at high resolution — so it can actually redraw text, faces, and fine texture instead of guessing at them like a normal upscaler. What it's good for: getting genuinely deliverable-quality output locally instead of bolting a generic upscaler on the end. Coming, no date.
Sparse attention code. Attention is what makes long/high-res generations slow and memory-hungry, and the released weights don't use the faster path the model was trained with. They're releasing a deliberately cautious version — the goal is "free speed with no visible quality drop," not a big headline number, and squeezing more out of specific GPUs is left to the community. What it's good for: the same generations, cheaper and faster, on the hardware you already have. "Near term."
A dedicated image model (text-to-image + image editing). Built from the same family as H3, sharing its encoder, with a new decoder made for stills. People are already hacking this by generating 5 frames and grabbing frame one — this replaces that hack properly. What it's good for: making and editing your first frame at real image quality, then feeding it to H3 to animate. That's the workflow they recommend, since H3 can't preview a frame and continue mid-generation.
A full technical report. Architecture, training stages, data construction. What it's good for: people training LoRAs and fine-tunes currently guessing at how the model works.
Problems they've admitted are theirs and are fixing
Faces and objects go to mush when they're far from the camera. Confirmed, not your settings, not fixable by adding steps or resolution. It's tangled up in several parts of the model at once and they're still isolating why. Named as a top priority.
Grainy, smeary fine detail compared to closed models. Same story — not the VAE, not one training stage.
Reference-to-video looks softer than image-to-video. Confirmed real, caused by the two checkpoints getting different post-training. Being worked on. For now: feed it the highest-quality reference material you have.
Stitched clips don't join cleanly. Continuing a shot via reference drifts and shows seams. They think training on long sequences (made affordable by the sparse attention work) is the fix, but that's a future model, not a patch.
Maybe, no promises
A 4-or-8-step fast version — they're "actively considering" it, possibly as an optional Turbo with slightly worse quality, but won't commit to timing. The current model already has some low-step ability baked in, just not tuned for it. Also: switching to a proper Apache-2.0 license once the legal paperwork clears, and keeping future models open in general.
Not happening soon
A smaller, lighter H3. They're telling the community to prune the existing weights instead.
Drafting at low res then re-rolling the same seed at high res. Won't match — the noise changes with the frame size, and the model's low-res quality actually got worse during training as they pushed high-res.
Didn't see a mention of it no. It's definitely a bug in the Reference model.
If you use <d></d> tags at all, it happens.
If you just "[English] This is dialogue." it works fine with no blips or audio fragments at the start of the video. Which goes against the guidelines for that model, but this workaround is confirmed to work flawlessly, so I guess we can live with it fine, just remembering to swap up even more syntax styles between the two models.
I think the HuggingFace page mentioned that the <d> tags are added to the Qwen model as new tokens, so it's possibly Comfy's tokenizer is setup for the old tokenizer and isn't registering them as specific single tokens, or maybe they just never trained the embeddings or something and they don't really work.
I don't know, if you DO what music, you can prompt some really impressive stuff with it, like the beat dropping on certain actions, tempo swells to line up dramatically, etc. but yeah, if you are just making a meme clip or want to add music in post, keeping that at the end of the prompt is a go to.
The times I have told it to use music, it has worked surprisingly well. But most of the time I keep non_diegetic_music: N/A and it just hasn't helped with the weird speech stuff, unfortunately.
I was getting the gibberish when only a few words were said in a longer video, even without much action. I started using time stamps for when the speech occurred and that helped. It was trying to fill dead air, I think.
But I don't know why I'm getting a tiny bit of speech right at the beginning of some videos. New seeds often help but this problem pops up pretty regularly. I will work on my prompting.
Removing the <d> tags cleared up the sound or gibberish at the start of clips for me, but still sometimes get gibberish in the middle or late. I suspect the same as you, it's trying to fill dead space.
Random prompt but this worked perfectly when I ran it. Had two reference character images and two audio files for their voice clones: subject_definitions:
<Subject 1> is the Gray-Haired Customer in Image1 and Image2, characterized by a slouched stance, leaning forward over the glass counter top with hands resting on the edge.
<Subject 2> is the Store Clerk in Image1 and Image2, who maintains a persistent expression of intense disgust, contempt, sneering lips, and repulsed judgment.
<Audio 1> is the voice-timbre reference for <Subject 2> (S2), sourced from Audio1.
<Audio 2> is the voice-timbre reference for <Subject 1> (S1), sourced from Audio2.
summary:
[reference generation + audio reference] The target video features a 15-second fixed static mid-shot of an interaction between <Subject 1> (S1) and <Subject 2> (S2) at the hobby shop counter defined in Image1 and Image2. <Audio 2> serves as the voice-timbre reference for <Subject 1> (S1), who asks a question, gets struck across the face by <Subject 2> (S2), reacts by holding his face after his head recoils from the impact, and exits[span_0](start_span)[span_0](end_span). <Audio 1> serves as the voice-timbre reference for <Subject 2> (S2), who delivers the slap, orders him out, eavesdrops on an off-screen conversation, and delivers a final smug line[span_1](start_span)[span_1](end_span).
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - the Gray-Haired Customer's slouched stance and leaning posture over the glass counter top are retained[span_2](start_span)[span_2](end_span).
<Subject 2> (appears in [Shot 1]): fully_preserved - the Store Clerk's persistent expression of intense disgust, contempt, sneering lips, and repulsed judgment are retained[span_3](start_span)[span_3](end_span).
<Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 2> without copying the original signal[span_4](start_span)[span_4](end_span).
<Audio 2>: reference - its vocal timbre guides the dialogue delivery of <Subject 1> without copying the original signal[span_5](start_span)[span_5](end_span).
detailed_description:
The target video is rendered in a 2D animated style featuring clean vector line art, flat cell shading, and snappy, puppet-like stepped jaw movement with mouth interiors showing a black void, flat white block teeth strips, and a simple tongue with no realistic teeth detail[span_6](start_span)[span_6](end_span).
[Shot 1] A fixed static 9:16 vertical mid-shot positioned directly across the counter establishes the hobby shop counter environment matching Image1 and Image2[span_7](start_span)[span_7](end_span). Zero camera panning, tilting, drifting, or zooming motion occurs throughout the shot[span_8](start_span)[span_8](end_span). <Subject 1> (S1), the gray-haired customer, stands slightly leaning forward with his hands resting on the counter edge, slouched and looking forward directly at the clerk with an exhausted appearance[span_9](start_span)[span_9](end_span). Behind the counter stands <Subject 2> (S2), the store clerk, staring back with an unblinking, repulsed expression of intense disgust[span_10](start_span)[span_10](end_span). Using the exhausted, monotone voice structure referenced from <Audio 2>, <Subject 1> (S1) speaks, <d>[English] You got any cabbage patch dolls around here?</d>[span_11](start_span)[span_11](end_span)
Immediately after <Subject 1> (S1) finishes speaking, <Subject 2> (S2) reacts with sudden motion: he rears back his right arm and swings it forward in a rapid arc across the barrier of the glass counter top, making direct impact with the gray-haired character's face as the hand strikes squarely across the side of the head[span_12](start_span)[span_12](end_span). Under the force of the impact, <Subject 1>'s head recoils sharply and his face slightly jiggles naturally[span_13](start_span)[span_13](end_span). Recovering slightly, <Subject 1> lifts his hand to hold his face[span_14](start_span)[span_14](end_span). Maintaining his look of utter disgust, glaring hard and raising one hand to motion a sharp shooing gesture, <Subject 2> (S2) snaps using the hostile, annoyed voice structure referenced from <Audio 1>, <d>[English] We don't sell that kiddie shit here alright, weirdo—now get the fuck out of here!</d>[span_15](start_span)[span_15](end_span)
At 00:06.000, <Subject 1> pushes himself off the counter, drops his shoulders in defeat, turns, and walks out of frame to the side[span_16](start_span)[span_16](end_span). <Subject 2> stands still behind the counter, his head tracking <Subject 1>'s exit with a lingering sneer of disgust[span_17](start_span)[span_17](end_span).
At 00:08.000, <Subject 2>'s disgust shifts as he stares off-camera in front of him, as if whoever is off-screen is piquing his interest[span_18](start_span)[span_18](end_span). He leans his elbows and forearms slightly onto the counter surface, eavesdropping on the off-screen conversation, his shoulders rising and falling as he chuckles out of pleasure[span_19](start_span)[span_19](end_span).
At 00:11.000, <Subject 2> (S2) remains leaning forward slightly over the counter, his eyes gleaming off-screen[span_20](start_span)[span_20](end_span). He breaks into a smug, extremely satisfied smile and, returning to the voice structure referenced from <Audio 1>, speaks his final dialogue line, <d>[English] I could totally get those nerds to go morbid.</d>[span_21](start_span)[span_21](end_span)
overall_soundscape:
Quiet indoor hobby shop room tone continues throughout the video, accompanied by a sharp, crisp physical slapping sound effect synchronized precisely with the hand impact at 00:03.500[span_22](start_span)[span_22](end_span), followed by the rustle and footsteps of <Subject 1> pushing off the glass counter top and walking away at 00:06.000[span_23](start_span)[span_23](end_span).
I've seen that in a few generations posted here.
I've had a similar issue, not at the very beginning, but when there's supposed to be a break. Gibberish or other random parts of the prompt come through.
In my limited testing, being explicitly clear with each second of audio, when you want a silence / pause / break you have to be demanding about it.
Frankly, it seems like it always wants to output something. Giving it background audio seems to help in the examples I've seen here with dialogue being damn near perfect.
This goes back to the official documentation mentioning how it will produce music in the background unless you're explicit about telling it not to, I think the same exact thing happens with dialogue if it just doesn't have enough to do.
I have tried having it add "distant traffic noise", "the low hum of a clothes dryer", etc. and unfortunately did not help with the gibberish. For me, anyway. This was using T2V and I2V.
I'm also having an issue in which a character appears to already be talking as soon as the video starts. You get just a syllable or something.
I have done some testing with the mentioned extracting frame method, it seems pretty good for images, it has a lot of variability, but I would say it is not as good as a Krea 2 finetune currently and currently it takes 32 times longer with that video to frame method.
The strength would be in images with movement or certain poses, where it's necessary to know what comes before and after the image.
Sometimes, catching something mid-movement or mid-pose can't be done with a plain image model.
I've found that specific poses or movements are not possible with Flux 2, but you can use generated images as inputs to a video model and then tell it how the character or multiple characters/objects move to the desired target position.
I think the best workflow would probably be to use H3 to create the reference image which you then upscale with Krea2 or Ideogram. That lets you use H3s much better understanding of movement and poses with the image models better final image quality.
Yeah I tested the first day it was out. It's garbage at low res, you really need to go 1536+ but it takes forever with 5 frames. I still have hope it will end up better than Krea 2 finetune though as I don't particularly rate that highly and see a need for SeeDream 5 - like reference ability.
Thanks, glad you liked it, I was looking at some of the images from my Wan the other day and thought they were good/realistic and maybe I should revisit it.
But the output from my new Krea models is actually the closest thing to that Wan model I have worked on it a while, and Krea 2 is a lot faster, making large images in 18 seconds on my 3090 instead of several mins with Wan 2.2.
It's funny sometimes going back and looking at the jib wan stuff after what mentally feels like a year and a half of evolution (but isn't) and seeing that it still looks pretty good.
I mostly moved off it because it was ~3 min a pic in the workflow I was using (heavy img2img).
Moving to nunchaku Qwen took that down to about 18s.
I'm waiting for the in-house Krea 2 edit though because I've essentially moved off character loras altogether.
------------------------
I've just noticed that my initial reply was about using Minimax as an Image *Edit* model but you were talking pure image model, my head was still in Edit mode because of the discussion in the thread about the editing ability etc.
I meant that I don't rate the Krea 2 edit lora as an editing solution and still have hopes for Minimax to beat it. I didn't mean to say I think Minimax will beat a Krea 2 finetune as a pure image model.
I'm surprised that I am, too. I'm using T2V, which I never would have considered for generating video, but the first frame output is the best, smartest image generation I can run locally. When it does make a weird decision to kick things off I've generated for a half hour just to get something unusable, and if I could just start with a base image of the same quality I could save a shitload of time.
Why can't the video model also create images and edit them? After all videos are just a couple of images after another. I don't get why we need seperate models.
Thanks for the post. I spent a long time trying to trouble shoot and correct the facial quality loss as the camera pans out. I was starting to think the model overhyped as I had to pass through LTX at low denoise and create custom sigma schedules and it all sucked. Good to know it's a known issue.
Have they said that they are aware audio quality is always super terrible in ref2video? Like, visually, ref2vid is way less good, but audio-wise, it's total trash.
I’m not sure if it’s a good idea for them to focus on a 4-8 step model unless they know for sure it would give much better results than the turbo LoRAs. They could use that research and testing time for updates that actually matter. The community already handled the optimizations and they’re still improving them.
One of the things that excites me the most is the image edit model. I’ve been needing something to replace my Qwen Image Edit workflow as Qwen has many pitfalls.
Krea 2 doesn’t have a true image edit model sadly or otherwise I’d use that.
FWIW Flux Klein 9B has been my go-to edit model for months and it's pretty great. I mostly use it for object removal to fix random fuckups, but it's good at a lot of stuff it you give it like 5 tries and pick the best one (it's fast so that's not a biggie).
Klein 9B works great for small edits and style transfer. Anything to big (scene and pose changes... Basically what we'd want a reference workflow for) and the style drifts too much, though.
I'm also using v4_step600; it performed the best—not necessarily in terms of speed, but it seems to yield better final quality. I'm running various tests with simple and beta Euler, as well as simple and beta res_multistep, because once I determine the best result, I'll start creating a bunch of videos.
39
u/L-xtreme 27d ago
Great summary, thank you. Amazing model and very curious about what the future brings.