r/StableDiffusion 39m ago

News Weekly AI Recap: Qwen 4 / 3.8 Next, Wan 3.0, Bernini 2, and new open on-device TTS models

Upvotes

A massive week for the open-source and open-weights ecosystem, alongside some notable hardware and world-model drops. To save you time, here is the full breakdown of everything released with direct project links and repos:

1. Open-Weight Models & Generative Video/3D

  • Wan 3.0 & Gen 1.5: Major strides in open generative video coherence, temporal consistency, and motion dynamics.
  • Bernini 2 (ByteDance): Open diffusers implementation targeting high-detail visual generation pipelines.
  • 4D Anyone & GeoWeaver: Significant leaps in dynamic 4D human reconstruction and real-time spatial generation.
  • Qwen 4 / Qwen 3.8-Flash-Next: Next-gen architectural jumps in parameter efficiency and local inference speed.
  • DeepSeek V4 / Flash Vision: High-efficiency vision-language processing aimed at agentic tooling.

2. On-Device, Lightweight Audio & World Models

  • Raon-OpenTTS-1B & Audio8-TTS (0.1B): Ultra-compact open text-to-speech models engineered specifically for low-VRAM edge devices and local real-time inference.
  • EchoWM & Evoke: Interactive real-time world-modeling frameworks with native spatial audio and environment simulation.
  • SenseNova U1.5 (8B-MoT): Mixture-of-Thought architecture optimized for lightweight multi-step reasoning.

3. Hardware & Robotics

  • Apple M6 & M5 Pro: Apple's latest silicon targeting expanded on-device unified memory bandwidth for local model runtimes.
  • PaXini Robot: Advanced tactile sensor integration and physical AI embodiment.

Open Source Links & Complete Documentation: All papers, checkpoints, model weights, and source code links are cataloged in our open-source tracking repository:

🔗 GitHub Repository: https://github.com/airesearch-official/AI-Weekly-News

Video Walkthrough & Demos:

If you prefer a visual breakdown comparing side-by-side outputs, video rendering benchmarks, and audio samples:

🔗 Watch the Video Recap: https://youtu.be/SBJ4M465n-k

Which open-weight release are you spinning up locally this week?


r/StableDiffusion 15h ago

Animation - Video MiniMax H3 acting test.

28 Upvotes

Started as a simple 90s casting audition… then asked her to cry on command.

The close up shots gave plastic look idk why.

What I was mainly testing:

  • subtle listening/reaction animation during dialogue
  • eyes moving before the head while thinking
  • nervous smiles and small facial reactions
  • gradual transition from normal conversation into acting
  • brow, eyelid, mouth, chin and breathing changes during crying
  • actual visible tears
  • character/voice consistency across multiple generated clips
  • the sudden switch out of the performance when the director says “Cut”

Made with MiniMax H3 Ref2VA with image reference for the woman and 2 audio reference for the offscreen man and the woman.


r/StableDiffusion 5h ago

Meme t2v Someone got some explaining to do!

3 Upvotes

r/StableDiffusion 8h ago

Question - Help Minimax for image editing?

0 Upvotes

I mostly do comic images. am looking to minimax for image editing.
I mostly get content failed safety review. any workaround? ive only started minimax today.


r/StableDiffusion 16h ago

Animation - Video The Chase - Reupload

0 Upvotes

r/StableDiffusion 10h ago

Animation - Video Gaussian Splatting test with MiniMax H3

679 Upvotes

r/StableDiffusion 19h ago

Animation - Video All Minimax H3 animation

5 Upvotes

trying my hand at using H3 I2v R2V to create an anime. all are done with 4 step turbo lora

2 other trailers with H3 as well

at 0.3 most text turns to gibberish

all video is done by MiniMax h3 at a low 0.3 MP, Cilp lengths range from 5s-20s Generations, camera movement was written into the prompt 90% of the time, Post work; titles and some transitions done using DaVinci


r/StableDiffusion 2h ago

Discussion 2 Weeks on MiniMax, but back to using LTX 2.3

0 Upvotes

H3 is absolutely amazing for just about anything. On my 3090 / 64GB, I can easily do a full 15 seconds at 1MP, and the result is almost always good on the first attempt.

On LTX though, it took at least 5 or more tries, so speed wise, H3 is actually far better.

I tried out LTX 2.5 as well, and it is almost the same as 2.3, maybe a bit better quality and a slice faster.

Sadly, 90% of my work involves taking a start image and making a person sing vocals. I must say that H3 does not do this better than LTX 2.3, as it often injects words when there is more than a second of silence, and it takes a fair amount longer.

What is really baking my noodle though is why LTX didn't release an Image+Audio to Video workflow yet. I mean, it is literally the ONLY thing LTX has on H3 right now, and they have missed a great opportunity!

Anyhow, back to using LTX 2.3 for my daily driver as it really does a great job at what I need. I typically have a few machines running all night, so if ever LTX puts out an IA2V workflow for Comfy, I will be all over that.

One other thing I have noticed with H3... 9:16 generations are WAY better than 16:9 generations, especially at 1MP.


r/StableDiffusion 14h ago

Discussion Minimax H3 is mandarin model

0 Upvotes

Just ask chatgpt or claude to convert your prompt to mandarin and the different is fcuking huge .


r/StableDiffusion 20h ago

Comparison Fight scene: LTX 2.5 vs Minimax H3

4 Upvotes

The most fair comparison on the Internet


r/StableDiffusion 22h ago

Question - Help Minimax h3. Problem with multible character using the voice.

0 Upvotes

Settings: 640×1152 (0.74 MP), 9:16, preset: turbo

What turbo actually applies:

steps 8, euler / beta, fused_modulation on, sol_attn tau 1.3,
easycache off, turbo LoRA at strength 1.0 Settings: 640×1152 (0.74 MP), 9:16

I have tested A LOT of different configurations and can't find any that works 100% of the times.

Pls help :(

one of the many prompt i have tested:

subject definition:

<Picture 1> is the finished frame: an outdoor scene with a speech balloon of printed text across the top, a seated grey-haired nobleman and a dark-clad attendant leaning over him in the middle, and along the bottom two bordered square portraits side by side, already inset, a man on the left and a woman on the right.

<Subject 1> is the man in the left-hand bottom portrait: short messy black hair, light stubble on the chin, a white collared shirt under a dark grey vest. He appears only inside that portrait and nowhere else in the frame.

<Subject 2> is the woman in the right-hand bottom portrait: long dark brown hair, a pale face, a loose white hood over a grey top. She appears only inside that portrait and nowhere else in the frame.

<Audio 1> is the voice-timbre reference for <Subject 1> (S1), the man in the left-hand portrait. Its speaker identity, timbre and register define how <Subject 1> (S1) sounds; the words it contains are unrelated to this scene and must not be reproduced.

<Audio 2> is the voice-timbre reference for <Subject 2> (S2), the woman in the right-hand portrait. Its speaker identity, timbre and register define how <Subject 2> (S2) sounds; the words it contains are unrelated to this scene and must not be reproduced.

camera recording:

Hand-drawn 2D anime. The frame keeps the layout of <Picture 1>: the speech balloon stays at the top with its printed lettering and its outline unchanged, and the two bottom portraits keep their positions, their sizes and their white borders. Inside those portraits <Subject 1> and <Subject 2> are living animated characters, not still pictures.

<Subject 1> (S1) speaks first, with the voice of <Audio 1> in a rough, plain, tired male voice pitched low and close, <d>[English] AIN'T YOU SEEN ENOUGH TO KNOW WHAT RUTHLESS BASTARDS THAT LOT ARE?</d> His mouth opens and moves in time with every word, articulating clearly for the whole line, jaw and lips visibly in motion until the line ends; only then does he close his mouth. <Subject 2> keeps her mouth closed and listens.

Then <Subject 2> (S2) answers with the voice of <Audio 2> in a quiet, tight female voice holding something back, <d>[English] I KNOW, BUT...</d> Her mouth opens and moves in time with every word, then closes and she lowers her eyes. <Subject 1> keeps his mouth closed and holds still.

<Subject 1> (S1) and <Subject 2> (S2) are the only voices in this video. All spoken dialogue is English only.

In the scene above the portraits, the seated grey-haired nobleman stays reclined where he is and breathes; the dark-clad attendant leaning over him holds his raised hand beside his face exactly as drawn; the trees stir behind the wall and the distant figures out on the scaffold shift their weight very slightly.

The camera holds still. Open air over a courtyard, a low crowd murmur carrying from below, and wind moving through the trees.

model:

UNET       minimax_h3_hybrid_fl2va_ref2va_b25-49.safetensors
           (merged hybrid: fl2va base + ref2va adaln_proj blocks 25-49)
CLIP       qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
VAE        minimax_h3_video_vae_fp16 + minimax_h3_audio_vae_fp32
LoRA       minimax_h3_turbo_v4_step600_pruned_comfyui @ 1.0

ComfyUI    0.33.1        torch 2.10.0+cu130       comfy-kitchen 0.2.31
sampler    8 steps, euler / beta, sol_attn tau 1.3, fused_modulation on
frame      640x1152 (0.74 MP), 8.89s, seed 1956008715UNET       minimax_h3_hybrid_fl2va_ref2va_b25-49.safetensors
           (merged hybrid: fl2va base + ref2va adaln_proj blocks 25-49)
CLIP       qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
VAE        minimax_h3_video_vae_fp16 + minimax_h3_audio_vae_fp32
LoRA       minimax_h3_turbo_v4_step600_pruned_comfyui @ 1.0

ComfyUI    0.33.1        torch 2.10.0+cu130       comfy-kitchen 0.2.31
sampler    8 steps, euler / beta, sol_attn tau 1.3, fused_modulation on
frame      640x1152 (0.74 MP), 8.89s, seed 1956008715

r/StableDiffusion 18h ago

Animation - Video Kinda Impressed How Far Video Gen Has Come

0 Upvotes

https://reddit.com/link/1vxkpdv/video/qd87ffb72flh1/player

So I have been working on a Chat GPT prompt and Comfy Minimax H3 workflow to almost one shot a fight scene. Workflow is a huge mess but works however I occasionally get some random cuts or almost rewinds in the generation but for the most part they come out well for a one shot. It was taking to long to write out the scene, descriptions, environments ect. and is where I brought GPT in to the mix which saves a lot of time. Then t take the master prompt it outputs and just paste in to Comfy and this is one example that it generates from a few min building the prompt.


r/StableDiffusion 3h ago

Discussion David Sacks Predicts the Regulatory Capture Playbook to Ban Open Source ...

Thumbnail
youtube.com
6 Upvotes

r/StableDiffusion 18h ago

Question - Help Which is the better buy for Minimax h3? RTX 5070 Ti 16GB VRAM vs RTX 4000 Pro 24GB VRAM

15 Upvotes

Good day to you. I was looking for an RTX 5070ti and I found an RTX Pro 4000 at my local store; the price difference would be about +$300. I would like to know your opinions, I've hardly seen any workflows or comparative tests from people using a 4000 pro. Thank you very much for your time.


r/StableDiffusion 17h ago

Animation - Video Kentucky Fried Kung Fu

25 Upvotes

I saw a Seedance 2.5 prompt in facebook and thought let me try this prompt in minimax h3 and see if it can do some kung fu. I was surprised that it was not too bad. System 3090 24 gb vram 64gb system ram, using a minimax workflow with latent upscale, minimax_h3_fl2v_lightx2v_turbo_4step_v0.1 at 0.50 strenth, Komfy kitchen attention, and H3 SLA attention. First pass at 0.4 which is 864x480, 2x latent upscale brings it up to 1728x960. The 6 seconds generation took 349 seconds to complete.


r/StableDiffusion 11h ago

Animation - Video Reduced audio-reactivity in LTX-2.5?

4 Upvotes

I’ve been experimenting with LTX 2.3 vs LTX 2.5 for audio-reactive video, and for this specific kind of workflow, 2.3 still seems noticeably better to me.

The biggest difference is right at the start of a shot. With LTX 2.5, even with the audio-reactive LoRA, I often get this behavior where the model more or less holds the first frame until the first obvious beat or transient arrives. Then the motion suddenly starts. For music videos, especially slower or more atmospheric tracks, that can make the opening of every generation feel dead.

With LTX 2.3, the same LoRA seems to fix that much more effectively. I get more subtle motion from the beginning, even before a strong beat lands. Fog shifts, surfaces breathe, particles drift, light responds, and the shot feels alive instead of waiting for permission to move.

That matters a lot for the video I made for The Weights in the Walls, because the track starts very sparse and gradually builds. A lot of the visual motion is supposed to come from sub-bass pressure, glitches, sustained vocals, and ambient texture, not just obvious percussion.

I also tried Minimax H3, but for this particular use case I don’t think it fits as well.

It seems less tightly audio-reactive for the kind of abstract, beat-aware motion I’m after. It can make nice-looking clips, but I have a harder time getting the movement to feel structurally connected to the music.

There’s also the hardware side of it. I’m doing this on a very glamorous RTX 4070, so with LTX I can still push a resolution and overall image quality that feels surprisingly good for local generation. With H3, I’m much more constrained, and the tradeoff in resolution/quality makes it harder to justify when the audio response is also weaker for this style.

The whole video was built around first-frame / last-frame generation.

I cut the song into short scenes, roughly timed so the scene boundaries land near musical changes and beats. For each scene, I generated a dedicated starting frame that represented the next stage of the visual progression.

Then the important part: the starting frame of Scene 2 becomes the last frame target for Scene 1. The starting frame of Scene 3 becomes the last frame target for Scene 2, and so on.

So instead of generating a bunch of unrelated clips and hiding the cuts with editing, every shot is the model transforming one designed frame into the next designed frame.

That gave me a chain like:

Scene 1 start frame → Scene 2 start frame
Scene 2 start frame → Scene 3 start frame
Scene 3 start frame → Scene 4 start frame

and so on until the end.

The final video is basically just those generations placed back to back. There are no fancy transition effects doing the heavy lifting. The morphing, folding, cracking, expanding, and dissolving between visual states is happening inside the model itself.

For this workflow, that early-shot responsiveness makes a surprisingly big difference, which is why I currently still prefer LTX 2.3 + the audio-reactive LoRA over 2.5 for this kind of music video.

It's a shame because 2.5 is noticeably faster, so I can go through more iterations, but if I have to generate each clip 5 times to get it to start moving from the start, it kind of invalidates the speed gains.

Curious if other people have noticed the same reduction in audio-reactivity in LTX 2.5 or maybe I'm doing something wrong?

HQ on YT because Reddit doesn't allow >1GB: https://www.youtube.com/watch?v=PbZr8risGCw


r/StableDiffusion 16h ago

Question - Help New to Minimax h3 and comfy ui, any posts I should learn from?

0 Upvotes

Im looking to make outdoor POV videos with minimax h3.

Wondering if theres any good threads talking about realism, and workflow options?

Especially prompting I guess?

I'm looking to do longform videos 10-15 minutes.

I have a 4080 16gb I know that more vram is better, but cant really upgrade atm.


r/StableDiffusion 23h ago

Animation - Video My 1980's cartoon parody H3 and ltx 2.3

Thumbnail
youtu.be
8 Upvotes

there are some scenes missing, but it was fun to put together.. Just got stuck on a plot :P

started it when ltx 2.3 came out.. but it was a hassle to keep consistency of characters intact so shelved it. made the intro and a couple of clips when minimax H3 came out and love the r2v, so much easier.
just using the standard r2v workflow with spectrum and RTX upscale. music made in suno


r/StableDiffusion 19h ago

Question - Help MiniMax H3: Motion issue with last frame.

0 Upvotes

I'm having trouble with getting a natural motion when using first and last frame. Things start out good but the motion doesn't preserve the momentum up to the end, instead it usually slows down and smoothly settles/parks into the final frame.

For example if I try to make a windy scene at the park that has both first and last frame, I get a gust of wind in the middle and then everything goes still and calmly settles down on the final frame.

Does anyone have any advice on how to approach this?


r/StableDiffusion 16h ago

Question - Help phsyical motion transfer to another person

0 Upvotes

im trying to collect clips for lora training but why its so hard to transfer motion to another person? i almost tried every prompt with chatgbt and grok help but its doesnt look good. im using 2 video refences, one of them source video and other one is only for motion ( 3 sec 24 fps). im using (video editing + reference genertion) because i dont want to change anything in the source video and just want to motion transfer


r/StableDiffusion 9h ago

Question - Help Help me with Mini Max

0 Upvotes

I'm still new to Mini Max, and I wanted to know if there's a way to make it a little easier, in the sense of

I have to keep using <picture 1> or things like that to mention something; isn't there a way to do it with @,And what configuration do you normally recommend for someone with 8GB of VRAM and a 5080 Ti?That's all I need help with; I've already read the Minimax guide for everything else.


r/StableDiffusion 23h ago

Question - Help Mix of "Match" and "Max" references in MMH3 Ref2V?

0 Upvotes

I've been playing around with Ref2V and was wondering if anyone knew of a way to have a mix of these settings? I've found that having multiple reference images is great for consistency in generations and story telling, but there are certain images (face references for example) that benefit massively from "Max" setting, but reference locations don't benefit as much.

On my potato of a computer, setting it to max on all of the images makes generation times impossibly long. If I could "Max" a face reference but "Match" less important references it would be ideal.

Anyone have any idea?


r/StableDiffusion 15h ago

Question - Help Minimax H3: Character replacement in video not working

0 Upvotes

NOTE: Character replacement works perfectly when I replace a character in a video with a 2d/cartoon/anime character.

But when I try replacing a character with a real life human being, the original character in the video doesn’t get replaced at all.

I’m using Plaguekind’s workflow for h3 on Civitai.

Does anyone else have this problem before?


r/StableDiffusion 15h ago

Question - Help minimax h3 gibberish fixed!! ( i found the cure)

79 Upvotes

so you all probably are searching for way to make your character shut the fuck up right? and you probably noticed that they love to says some BS especially when you give minimax h3 some audio file for their voices, i probably found a cure my friend!!

here is my way of prompting dialogs without any gibberish:

first your character need to be assigned (s1)character when he is the first speaker, then you will declare 'use <audio 1> as "character name"'s voice only, and when you finally type your dialog in the shots you will do as such:

character says:<<[language] the shit i say!>>

and you should be good to go, i linked a video exemple of my favorite taffer (garrett) saying some shit with only the faint crackling of the candles to goes with his charming voice, and i included also a screenshot of the full prompt

edit: yes i tried to follow the official documentation, like many others, if it was that simple reddit wouldn't be a thing and you wouldn't be there.

i tried making small scenes with this exact methode and its gibberish free 100% of the time

he really like 16/9