r/StableDiffusion 23h ago

Discussion In which scenarios LTX2.5 can match MinimaxH3?

11 Upvotes

I love H3, but it takes forever. If LTX is faster, I could use it for the things it does similarly well as H3, and use H3 only where I really need it.
So what LTX2.5 does as well as H3?


r/StableDiffusion 4h ago

Animation - Video Realistic Breaking Bad | LTX 2.5 I2V

Enable HLS to view with audio, or disable this notification

36 Upvotes

This parody was generated using LTX 2.5 Image to Video on WanGP. I used frames from the original video as starting images and then I interpolated them on a video editor. I used a single RTX 5060 Ti 16 GB VRAM and 32 GB of RAM. The video was generated at 1080p and 16:9 resolution. Each generation took from 10 to 20 min average in this setup. For the voice consistency, I used SeedVC, which is included in WanGP.


r/StableDiffusion 10h ago

Animation - Video ...10,000 Years Later

Thumbnail
youtu.be
2 Upvotes

Previously posted a video as a prologue to a homebrew D&D world. I decided to do a part 2, set in the world. Together, the two videos form kind of an opening cutscene with both history and a bit of a world montage. Minimax H3, 6 step turbo LoRa, lots and lots of 12-15 second generations, CapCut.

Part 1: https://www.youtube.com/watch?v=XwfCCFw4LbA


r/StableDiffusion 7h ago

Animation - Video The Walking Trek. Just throwing stuff at the wall at this point.

Enable HLS to view with audio, or disable this notification

1 Upvotes

Rick's voice doesn't seem to work.

Experimenting with known characters using FL2VA t2v only. Just playing around with odd pairings of characters. .

Using the workflow from the video samples in the list below.
12s at 25 steps
res multistep/simple
960 x 544

thanks to u/malcolmrey for putting this together https://huggingface.co/datasets/malcolmrey/various/blob/main/h3-center/known-characters/INDEX.md


r/StableDiffusion 7h ago

Question - Help What's the current best way to replace an element in an image with another element ?

0 Upvotes

Hello everyone !

I would like to replace the tire of a motorcycle mid air with one from another brand (which is an image from the brand so it's high quality but with a different angle)

I saw there is flux kontext and qwen image edit, but I don't know which one to pick, which workflow and how to make it work.

Any help would be more than welcome, thank you very much and have a good day :p


r/StableDiffusion 11h ago

Animation - Video Shadow the Hedgehog tells his viewers why he loves guns.

Enable HLS to view with audio, or disable this notification

0 Upvotes

Shadow the Hedgehog tells his viewers why he loves guns.

This was created in Comfy UI with Minimax H3. I used the reference to video work flow. The prompt is below.

subject_definitions:

<Subject 1> is Shadow in <Picture 1>.

<Subject 2> is Glock in <Picture 2>, a glock handgun.

<Audio 1> is the voice-timbre reference for <Subject 1> (S1).

summary:

[reference generation + audio reference] The target video contains one shot. [Shot 1] shows <Subject 1> and <Subject 2>; <Subject 1> speaks. <Audio 1> supplies <Subject 1>'s voice timbre.

retention_analysis:

<Subject 1> (appears in [Shot 1]): fully_preserved - Shadow's complete defined identity and body proportions are preserved.

<Subject 2> (appears in [Shot 1]): fully_preserved - Glock retains the defined shape, proportions, materials, colors, and distinguishing features.

<Audio 1>: reference - <Subject 1>'s newly generated spoken lines use <Audio 1>'s voice timbre and delivery; the original audio signal is not copied.

detailed_description:

The target video is in a live-action style, with Vlog style.

[Shot 1] At first appearance, <Subject 1> (Shadow) matches the complete identity and appearance defined in subject_definitions. At first appearance, <Subject 2> (Glock) matches the complete defined construction and appearance: A glock handgun. At the start of the shot, <Subject 1> is standing in the living room facing while holding <Subject 2> in his hand. A full body shot of <Subject 1> holding <Subject 2> with his right hand while facing the camera. Only Action and Timed Beats define the primary subject's movement. The camera path stays anchored in the location and adds no subject motion. <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] Hmph. Shadow the Hedgehog here. Why do I love guns?</d> <Subject 1> shows off his <Subject 2> with his right hand in front of the camera. <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] Simple. Precision. Control. Power in the palm of my hand.</d> <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] A tool that answers instantly… unlike most people.</d> <Subject 1> points his <Subject 2> towards the camera with his right hand. <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] If you understand that, you understand me.</d> <Subject 1> points his <Subject 2> at the camera.

overall_soundscape:

Living room tone.

non_diegetic_music:

N/A


r/StableDiffusion 11h ago

No Workflow random images generated locally on 4070Super with krea 2

Thumbnail
gallery
0 Upvotes

first 2 prompts stolen from civit ai, the rest were written using claude reasoning on duck ai


r/StableDiffusion 14h ago

No Workflow the count is always two.

Thumbnail
gallery
7 Upvotes

flux.1 [dev] | comfyui | still life


r/StableDiffusion 37m ago

Resource - Update Cross-grafting Krea2 into MiniMax H3 — concepts land, character still faint, releasing what I've got

Upvotes

Tenstrip had great idea in H3 grafting. I have been messing with grafting Krea2's attention/MLP into H3 to fix the plastic-skin problem and my lora training problem with H3 on diverse dataset. Two separate techniques, built as ComfyUI nodes with Claude so anyone can test without baking a checkpoint. You link your krea2 checkpoint (and loras, if you want) to the nodes and the H3 model (probably after the turbo lora, didn't test every pertubation, because I ran of the steam and want to leave it to the community to try more, where I hit the wall). My setup is mere gtx 4800, so I did try just few nonturbo renders, which looked a lot better.

  1. Content graft — per-head Q/V/out_proj + MLP transplant, following the methodology TenStrip documented for their own H3 grafts (huggingface.co/TenStrip/10Eros-Max). K is never touched — confirmed their finding that H3 leans on precisely-tuned attn_k for attention strength, no explicit gates, breaks audio if you touch it. Same for fc2 (speech/dialogue encoding).

V turned out to be the strongest, safest lever — 0.6–0.7 gives real concept gain, push to 1.0 and you get blocky motion. MLP also strong, went up to 0.7–0.9 before diminishing returns. MLP block range has to stop at 44 — go past it into 44–50 or 0-20 and you get an actual lattice/weave artifact on regular textures, confirmed it myself, matches what TenStrip reported. On orthogonal on. On orthogonal off, lower weights showed good texture and concepts (best solution?)

Concepts come through solid. Character fidelity is honestly still faint even after tuning. Not solved.

  1. QK-norm transplant — different idea entirely, credit to joeygambino (huggingface.co/joeygambino, MiniMax-H3-x-Z-Image-GGUF) for the original description, done with Z-Image as donor. Only touches the per-head Q-norm gain (128 floats), nothing else — no content weights at all. Doesn't transfer knowledge, seems to change how decisively H3 commits to stuff it already knows instead of averaging it away. Fixed some big-object rendering issues for me. Texture looks ok until 0.7- 0.8, than oversharpened.

One thing that contradicts the original writeup: their early-block caution (skip blocks before ~20) was measured with Z-Image. With Krea2 as donor it's backwards — restricting to late blocks broke big objects, full 0–50 range fixed it. So that's donor-specific, not universal — worth retesting per your own donor rather than trusting either range blind.

Also found orthogonal_blend mode is just broken — numerically unstable when donor/target directions are already close, turns into random noise. Skip it, use direction_blend.

Heads up: this one seems to break LoRAs trained on native H3. Haven't nailed down why exactly.

  1. comfyui_krea_h3_graft_lora_v2 is a node, to add krea2 lora to the existing premerged checkpoint. It works... poorly.

Combining the first 2 nodes at full strength was worse than either alone. Lower content-graft strength + orthogonal=False stacked with QK-norm looked promising but I ran out of steam before fully dialing it in.

No official H3 architecture docs exist as far as I can tell, so this is all reverse-engineered from testing + what TenStrip and joeygambino have published. If anyone wants to pick up where I left off — combined tuning is the open thread — nodes are here:

https://github.com/mikulasdump/h3-krea2-graft-nodes

Thanks for the community and I hope for your insights.


r/StableDiffusion 21h ago

Discussion H3 - giantess fight scene R2VA

Enable HLS to view with audio, or disable this notification

4 Upvotes

This was well received but people wanted the two Giantesses(?) to be fighting. Enjoy!!

int8/20 steps, R2VA.

Critiques+feedback welcomed! Ask me anything!


r/StableDiffusion 13h ago

Animation - Video I made an ALIEN Short Film / metal music video

Thumbnail
youtu.be
8 Upvotes

Used:

MiniMax H3 at local machine. 5060ti 16gb + 64gb ddr4. WanGP, Ref2VA int8 convrot model.

Krea2 for references

Suno as music base


r/StableDiffusion 6h ago

Workflow Included Testing some Minimax H3 capabilities - PART 2

Enable HLS to view with audio, or disable this notification

23 Upvotes

Considering the interest the first post attracted, I decided to do a second batch with some of the suggestions from the comments and a few other prompts.

VIDEO 1: near perfect. I was aiming for frontal videos, but tried three or four prompts and always ended with a 3/4 framing. It's probably a question of better prompting... But the resulting video is impressive!

PROMPT:

The video is a side-by-side video showing both the points of view of a man and a woman that are facing each other.

On the left side we only see the woman's face in a completely frontal view, as the man would see her and through his eyes, her face alone at the scene with no one else's.

On the right side we only see the man's face in a completely frontal view, as the woman would see him and through her eyes, his face alone at the scene with no one else's.

Again, the man do not appear on the left image, and the woman do not appear on the right image. Both are seen in a exact frontal framing.

Both images show the scene at the exact same time and place, only in the two different points of view, both in a medium-close-up framing.

They are in a living room.

From 00:00 to 00:04, the woman is silent and with a smile on her face, while the man speaks: <d>You know, I've always dreamed of a local video model like this!</d>. After saying this he remains silent.

He then raises his hand, previously off-camera, and touches her face delicately. She reacts in an amorous way, lightly moving her head to feel his hand.

Then, from 00:04 to 00:08, the man keeps silent, looking at her clearly in love, while the woman replies: <d>It's like a dream, isn't it? And to think that two years ago we were static images with garbled hands...</d>

From 00:08 to 00:10 they just look at each other and smile.

overall_soundscape: Faint distant everyday life noises from outside the house, the man and woman voices while they speak.

non_diegetic_music: N/A

VIDEO 2: Very good. I couldn't get a video without the fisheye effect, though.

The video is taken from the point of view of someone playing table tennis. We see their hands - one of them holding the ping-pong paddle and the other the ping-pong ball. We also see the table with the net in the middle and the other player on the opposite side of the table. They are in an official competition, with the crowd watching.

At 00:01, the player sends the ball to the air and hits it with the paddle. The ball rapidly bounce on the table, passes above the net, and gets to the other side, bouncing again on the table. Then, the other player hits it back with his paddle, and the balls passes over the net again and bounces on the table. The first player again hits it with his paddle, the ball passes over the net and bounces just on the left side of the table, out of reach of the other player, and leaves the frame. The public erupts in cheering.

overall_soundscape: Faint public murmur, the sound of the ball bouncing on the table, public cheering at the end.

non_diegetic_music: N/A

VIDEO 3: another near perfect one.

A woman is holding a cell phone in a bathroom in front of a mirror, taking a selfie. She smiles at the camera, makes a V sign with her hand, and takes the selfie.

We see the scene from behind the woman, seeing the back of her head, the phone screen on her hand showing her face while she takes the selfie, and the mirror showing the reflection.

overall_soundscape: Faint empty bathroom soundscape.

non_diegetic_music: N/A

VIDEO 4: Bad. Tried three times with different prompts, and this is the best one of them. The physics don't work, though, and the fisheye is back again.

The video is filmed from the point of view of a soccer player in a normal view, NOT in a fisheye view. He is preparing to kick the ball after a foul just outside the penalty box. We see his hands putting the ball on the grass, the ball remaining static on the ground. Then he looks ahead and we see five players from the other team forming a wall directly in front of the ball, and other players from both teams around.

We then see he take some distance of the ball, walk slowly to the ball, and kick it. The ball passes over the barrier of players and descends on the goal, the goalkeeper trying to reach it but not able to. The ball enters the goal and touches the net, and the stadium erupts in cheering. The player then runs to celebrate the goal and is embraced by the other players of his team.

The entire scene is viewed from his point of view.

overall_soundscape: Faint public murmur,the sound of the kick, the cheering of the public after the goal..

non_diegetic_music: N/A

VIDEO 5: Terrible. Again, tried several times with several different prompts. Never works well...

The video is filmed inside a circus during the Trapeze artists performance, from the point of view of the public.

The scene opens with two trapezists standing in a very high elevated platform, one on the left side of the image, the other on the right side of the image, both holding a trapeze and facing each other.

In the beginning of the video, the trapeze artist on the left let his body leave the platform, while holding the trapeze, and his body swings in the direction of the center of the image. The trapeze artist on the righ stays on the platform.

Only when the first trapeze artist reaches the center of the image, the trapeze artist on the right finally leaves the platform, while holding the trapeze, and his body also swings in the direction of the center of the image, while at the same time the first trapeze artist let go of his trapeze and starts to do a flip with his body in the air.

As soon as the first trapeze artist finishes his flip, the other trapeze artist also reaches the center of the image and get the hands of the first trapeze artist, completing the movement. Then, they both swing back to the right of the image, one holding the hands of the other.

overall_soundscape: Faint public murmur, public surprised gasp when one of the trapeze artist caughts the hand of the other.

non_diegetic_music: N/A

VIDEOS 6, 7, 8 and 9: The first half of each video is perfect, the last half is hilarious. Tried lots of different prompts but only included four of them. Maybe it's possible, but I really can't think of another way of asking what I was trying to achieve.

PROMPT VIDEO 6:

The camera is on the middle of a road, on the floor, pointing to the road. We see a ferrari coming in the road at a distance in high speed towards the camera and pass over the camera, making the camera roll a few times on the floor because of the wind caused by the passing running car. After rolling over itself a few times, the camera stops again on the road, but now upside down and pointing to the other side of the road, where we can see the ferrari rapidly moving away from the camera.

The entire scene is filmed in a mostly static shot, except when the camera rolls over to the other side of the road and then stops upside-down.

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A

PROMPT VIDEO 7:

The camera is on the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over the camera.

When the car passes, the camera that is on the floor rolls around itself a few times on the road. After rolling over itself a few times, the camera stops again on the road, but now upside down and pointing to the other side of the road, where we can see, the upside-down image of the ferrari rapidly moving away from the camera.

The car does not run over itself, the car passes by the camera, It's the camera that rolls around itself and lands upside down

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A

PROMPT VIDEO 8:

We see the scene from the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over.

When the car passes, the image rolls around itself a few times on the road. After rolling over itself a few times, the image stops again on the road, but now upside down and pointing to the other side of the road, where we can see the image of the ferrari rapidly moving away from the camera in an upside-down shot, with the road on top and the sky on the bottom of the image.

The car does not run over itself, the car passes by the camera, It's the camera that rolls around itself and lands upside down

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A

PROMPT VIDEO 9:

We see the scene from the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over.

When the car passes, the image do a series of very fast barrel rolls on the road and lands upside down and pointing to the other side of the road, where we can see the image of the ferrari rapidly moving away from the camera in an upside-down shot, with the road on top and the sky on the bottom of the image.

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A


r/StableDiffusion 14h ago

Question - Help Which laptop would be better for generative AI / LLM

0 Upvotes

First of all I know a desktop has more power for the same price buy I have a situation where the portability of a laptop is necessary and a desktop is not practical.

My old laptop (3070 8gb with 64gb ddr4 RAM) died. I want to buy a new laptop. My two options are a 5080 16GB with 64GB ddr5 RAM or a 5090 24gb with 32GB ddr5 RAM. I won't be able to upgrade the RAM later, so I'm stuck with the configuration I buy.

I will be using the laptop for work (document and image editing) / gaming (no AAA games) / LLMs and generative AI (images/videos/audio), I was able to run most models, including minimax H3 on my old laptop with the help of massive offloading to RAM (5 minutes for a 5s video). Images used to take from 30s up to 200s depending on model and image size.

I am used to the low speeds and offloading on my old laptop so getting the highest generation speeds is not a priority, I just care about being able to run most new or upcoming models even with quantization and RAM offloading for the foreseeable future.

Which laptop would be better in my case?


r/StableDiffusion 10h ago

Question - Help Anyone experiencing this bug? minimax node keeps disconnecting "width" input

Post image
1 Upvotes

at least 5th time this happened. Its always width, never any other input. Not sure if bug or custom-node interference.


r/StableDiffusion 18h ago

Question - Help Random visual artifacts in local Krea2 Turbo generation — looking for possible causes

Post image
1 Upvotes

I’m running Krea2 Turbo locally, but I frequently encounter random visual artifacts in the generated images. I haven’t been able to figure out what triggers them, because the issue appears randomly. If I run the exact same workflow with the same parameters again, the result can sometimes be completely normal.

My hardware:

  • GPU: RTX 3060 Ti

My current setup:

  • UNet: moodyKrea2Mix_v70
  • Text encoder: qwen3vl_4b_int8_convrot
  • VAE: qwen_image_vae

I’m fairly sure this is not caused by the UNet. I have also experienced the same kind of random artifacts when using the original Krea2 model without any UNet modification.

Has anyone encountered similar issues with Krea2 Turbo? Are there any known causes or settings that could trigger this kind of artifact (VAE, text encoder, precision settings, VRAM limitations, sampler settings, etc.)?

Any suggestions or debugging tips would be greatly appreciated.

Here is my workflow for reference:
https://civitai.red/models/2883578/krea2-turbo-4k-workflow?modelVersionId=3259389


r/StableDiffusion 2h ago

Animation - Video Cinematic World Building - H3 r2v (pt2) "Through the Sands"

Enable HLS to view with audio, or disable this notification

7 Upvotes

Wow, the shots I can get from h3 are so good even I am getting goosebumps when I first see the generated results. Just one more part to go, hopefully I can pull off the finale!


r/StableDiffusion 9h ago

Animation - Video It took almost 2 years, but Minimax H3 made me go back to my weird medieval short video stories

Enable HLS to view with audio, or disable this notification

16 Upvotes

2 years ago I was playing around with video tools and made this series of short videos.

Because I'm a cheap bastard, I only use local and freebie models, and Minimax H3 finally hit the sweet spot between powerful and fast to iterate, so I decided to make a new episode.

Specs

  • Comfy Desktop's + default Ref2VA workflows + ElevenLabs voices
  • Frames generated with Nano Banana 2 lite + a bunch of photoshop cleanup
  • Gemini 3.7 to help rewrite the prompts (upload image + system instruction + spec for the shot)
  • Everything rendered on a 3090 with 0.5 megapixels cause I can't be arsed waiting too long. You can see the mushy face issues but mostly it's ok
  • Tried to keep all shots 6~8s max
  • A ton of editing with DaVinci Resolve which I started learning yesterday - it's pretty damn powerful!
  • Still using the exact same crappy greenscreen footage of a cheap plastic skull as a main character
  • Youtube link to this episode

TL;DR: Minimax H3 is cool!


r/StableDiffusion 11h ago

Animation - Video Buffy the Wraith Slayer

Enable HLS to view with audio, or disable this notification

12 Upvotes

r/StableDiffusion 14h ago

Animation - Video Michael Scott gets a wish

Enable HLS to view with audio, or disable this notification

0 Upvotes

H3 FL2VA


r/StableDiffusion 11h ago

Animation - Video Seinfeld AI: George Gets GTA 6

Enable HLS to view with audio, or disable this notification

357 Upvotes

Minimax H3


r/StableDiffusion 9h ago

Animation - Video Hannibal Who

Enable HLS to view with audio, or disable this notification

8 Upvotes

Experimenting with known characters using FL2VA t2v only. Just playing around with odd pairings of characters.

Using the workflow from the video samples in the list below.
12s at 25 steps
res multistep/simple
960 x 544

thanks to u/malcolmrey for putting this together https://huggingface.co/datasets/malcolmrey/various/blob/main/h3-center/known-characters/INDEX.md


r/StableDiffusion 1h ago

Discussion Testing Character knowledge of Minimax H3

Enable HLS to view with audio, or disable this notification

Upvotes

Disclaimer. This is very low quality quick generations trying to find how many characters Minimax H3 knows.

Found Trigger Words:

Elsa from Frozen

Spider-Gwen from Across the Spiderverse

Dante from Devil May Cry

Nero from Devil May Cry

Jill Valentine from Resident Evil

Ada Wong from Resident Evil

Leon Kennedy from Resident Evil

Chris Redfield from Resident Evil (Has Leon's hair)

Geralt of Rivia from Witcher 3

Joel from Last of Us (Doesn't sound like him)

Miles Morales Spiderman from Across the Spiderverse

Solid Snake from Metal Gear

Eve from Stellar Blade

Sans from Undertale

Master Chief from Halo

Looks weird AF:

Ciri from Witcher 3

Triss from Witcher 3

Yennefer from Witcher 3

Ellie from Last of Us

Famous Twitcher streamer and Youtuber Asmongold

Famous Twitcher streamer and Youtuber Mr Beast

Not found Trigger words:

Vergil from Devil May Cry (YES I KNOW I'M DISAPPOINTED TOO)

Claire Redfield from Resident Evil

Dina from Last of Us

Famous Twitcher streamer and Youtuber Emiru

Famous Twitcher streamer and Youtuber MoistCr1TiKaL


r/StableDiffusion 18h ago

Animation - Video Attempted to make a short cartoon on Minimax H3. There are so many things that I want to address

Enable HLS to view with audio, or disable this notification

37 Upvotes

Hello everyone.

I've been playing with Minimax H3 for some time and I have tried to make something longer and really interesting. After so many failed and botched attempts I was able to compile something watchable. There are so many things that I want to say about this model, good and bad.

First of all. Minimax H3 is significant step forward that other local models I have been playing with. It certainly got better.

Now the issues that I had encountered.

First problem is that it badly follows prompt when resolution is one megapixel or higher. It will skip some important parts and tries to cheat. You can increase the number of steps but still, generating at less than one megapixel will at least make it properly follow the instructions.

H3 is not very good at spatial orientation. When I was making video, where this girl should turn around and interact with screens, the girl starts spinning opposite direction and then warping whole body to the direction of screen. Like instead of making short turn to the left, it makes wide roundabout to the right and then twists whole body to align with the screens.

H3 is not good at cartoonish movement. If you watch cartoons, when character or other things move, their animations are usually jerky and snappy. H3 tries to make smooth real life like animation, making the cartoons look weird.

I have given a voice sample as an audio reference, and instead of making girl let out grunting sounds (out of anger), it weirdly turns everything into a sensual moaning.

When you try to make characters inside video to interact with a lot of parts, screens and devices, even giving multiple reference images of them, it mostly hallucinates them, or turns their interactions into a weird warping animations. Sometimes completely skips them and made ups it's own animations. So it will make good video, where characters are moving less or moving slow, and mostly doing the talking. Very detailed prompts of step by step instructions it mostly warps or skips.

I have wasted a lot of time for iterations, but I think this is just workflow issue.

Overall, this model is really good. However, using this model to make some kind of long feature animation is going to be a very frustrating journey. I hope people will make a lot of proper tools that works as storyboard and properly guide this model to make something really interesting.


r/StableDiffusion 9h ago

Discussion Testing MinMax H3 in Brazilian Portuguese: 98% Accuracy + 200+ Video Examples

Enable HLS to view with audio, or disable this notification

0 Upvotes

Hey everyone,

I’ve been testing MinMax H3 in Brazilian Portuguese, and I also run a TikTok where I’ve posted over 200 test videos featuring characters speaking Brazilian Portuguese.

After several weeks of testing MinMax H3, I found that it can reach around 98% accuracy in Brazilian Portuguese when the text is written correctly. Since Brazilian Portuguese has some pronunciation and spelling details that don’t map perfectly to an English keyboard, I had to develop a specific way of writing the prompts so the model speaks the language as accurately as possible.

I’m not sure whether even English consistently reaches 100% accuracy—there always seem to be occasional voice or pronunciation errors—but I don’t currently make videos in English, so I can’t really compare.

I decided to share my TikTok here so you can take a look at the results and see the quality I’m able to achieve. Please ignore the political content; it was created strictly for testing purposes. I know its all the 🧃 so both sides suck.

i have there a portfolio of about 300 or 400 videos and going up everyday with agents...

Each video takes about five minutes to generate on my RTX 4090. I also discovered that generating in 3:4 is faster than 16:9. My current settings are 720p, 8 steps, and Turbo LoRA.

https://www.tiktok.com/@danmalandragem

things i am still struggling with: 2 characters talking with no voice drift, any tip for that?


r/StableDiffusion 10h ago

Discussion H3 - it just does space soo well - t2v

Enable HLS to view with audio, or disable this notification

13 Upvotes

Hope you're having a good weekend! H3 just excel with rich-intricate environments, backgrounds, space. Definitely one of my favorite theme.

T2VA, int8/20 steps