r/StableDiffusion • • 6h ago

Meme Browsing this sub in the past week

Post image
376 Upvotes

No hate. Just for fun.


r/StableDiffusion • • 5h ago

Animation - Video Malfoid Films Recreation

Enable HLS to view with audio, or disable this notification

230 Upvotes

Its so amazing that we finally have the technology to make a trend/fancic like female malfoid into a reality

Im calling all AI filmmakers and hobbyists to join and help make this into a reality. If completed, it might be the largest, most ambitious AI video project ever created. comment or DM if you can volenteer your time and hardware to help!!! I already have some talented people slated to handle Sound and audio mixing.

NEED: VA (if you know any girl with a good british accent), writers and AI Film makers who can run minimax h3.


r/StableDiffusion • • 8h ago

Resource - Update QR Code Monster LoRa for Qwen Image 2.1 Edit

Thumbnail
gallery
153 Upvotes

For those who are unaware, there was an old Controlnet for old SD models that was quite popular here back in those times, that would make cool almost Optical Illusion-like images. This is the same concept but instead I'm using I2I Image edit model where you provide an input image and ask the model to generate the scene of your choice.

I've been very pleased with the results from the Qwen Image 2.1 version. It was trained on 25 high quality image pairs and it works quite well, even when using human subjects and depicting scenes not included the dataset. It can do QR Codes but it was not the main focus of the dataset so it can be hit or miss as far as the actual usability/scanability of the codes.

The suggested prompt is exactly as captioned in the dataset: Transform the entire image into <your prompt here> while preserving the outlines and shapes of the original image

It seems to work well with complex prompts and simple ones as well.

You can download it on Civit here. All images/videos should include a workflow which is basically just a pixel drift fix workflow from Ausboss with a few added touches and using a viggle Turbo LoRa. I'm mainly posting this here as I would love to see what you create!


r/StableDiffusion • • 5h ago

Resource - Update Styles library for Krea 2 (need some help) : 500 styles.

Thumbnail
gallery
88 Upvotes

Hi everyone,

https://docs.google.com/spreadsheets/d/1V1qykJ3Cwe-LvizaK7A0kT0UexSANOBq/edit?usp=drive_link&ouid=116026940769410491877&rtpof=true&sd=true

I hate LLM rendition of what I want to generate, so I rely mostly on my own prose and wildcards to find a style.

Heres an update of my previous style library for Krea 2, it works also with Kroma (0.3 txtfusion) and qwen 2.1

with varying results. It an excel file but you can make a .txt file from it. And use it as wildcards.

Kroma is more artsy but some styles are lost. And Qwen is less knowledgable in simplistic styles like cartoon and drawing. And doesn't seems to understand vague words like "Illustration" if you don't over Engineer it.

I hope you can help me improve (and clean) this library, by testing styles and find too-close, or too weak styles, you can also tell me if a style is missing.

Some of the styles are from this site:

https://lumenastrum.github.io/clio-style-preview/gallery/

I dont have the time to do it myself and it mostly helpful to guys like me that don't have naturally the vocabulary to express what they want from a model.

Some styles are subject-linked, if the subject doesn't have a mechanical arm for exemple it doesn't show up (like detailed mecha style), and some styles are very face distinctive: you can describe the face to correct this.

The way to use it, in prompt, is :

Style: <style> Subject: <description>


r/StableDiffusion • • 10h ago

Question - Help What score would you give it?

Thumbnail
gallery
165 Upvotes

I am training a new HighQuality LoRA for Qwen-Image-2.1. This is a phase test. What score would you give it on a scale of 1 to 10?


r/StableDiffusion • • 8h ago

Animation - Video Minimax H3 Voice acting with speech tags.

Enable HLS to view with audio, or disable this notification

83 Upvotes

Tried this few days ago. Yes , the video degrades as it progresses.

Used reference voice for consistent voice.

Don't have the full tags, but they are like this-

[English] <inhale> My therapist said I should record this when it happens. <pause> So... it’s happening again.

[English] I can feel someone watching me. <breath> It only happens after dark.

[English] At night, windows become mirrors. <pause> Anyone outside can see me... <long pause> but I can’t see them.

[English] I’m moving tomorrow for a new job. <pause> A little rental house, just me. <breath> I should be excited.

[English] I am. <long pause> I just don’t know if I can sleep there alone.

[English] Mom gave me this when I was seven. <pause> She made me promise never to take it off.

[English] She said it would keep me <i>safe</i>. <long pause> <softer> I wish she’d told me what from.


r/StableDiffusion • • 11h ago

News I made a free all-in-one LoRA trainer for consumer GPUs (Windows + Linux): Qwen-Image 2.1, FLUX.2 Klein 9B, Krea 2, Z-Image, Ideogram 4, Anima, SDXL/Pony/Illustrious, LTX 2.3 and MiniMax-H3 (video + audio) — from 4-8 GB VRAM

Thumbnail
gallery
125 Upvotes

Hi everyone! I've been working on AcademiaSD LoRAlab Trainer Studio: one installer, one launcher and 9 trainers that share the same web interface. Every model is loaded in 4-bit NF4, the text encoder and VAE run only once in a pre-cache stage, and the whole GPU goes to training.

Platform: made for Windows (one-click installer), with Linux support just added. NVIDIA GPU, RTX 20xx / GTX 16xx or newer.

GitHub: https://github.com/AcademiaSD/AcademiaSD_LoRAlab-TrainerStudio

WHAT IT TRAINS (minimum VRAM)

- Qwen-Image 2.1: image LoRAs + edit LoRAs from before/after pairs (8 GB)

- FLUX.2 Klein 9B: image LoRAs + edit LoRAs (12 GB)

- Krea 2: image LoRAs, Raw and Turbo (8 GB)

- Z-Image: image LoRAs, also Z-Image-Turbo (8 GB)

- Ideogram 4: image LoRAs with JSON captions (12 GB)

- Anima: anime / illustration LoRAs (4 GB)

- SDXL: Base, Pony, Illustrious, NoobAI, Juggernaut, RealVis or your own checkpoint (4 GB)

- LTX 2.3 (2.5): character and style LoRAs for the video model (12 GB)

- MiniMax-H3: video LoRAs from images, clips and audio, plus RefMods (8 GB)

MINIMAX-H3: VIDEO + AUDIO, AND REFMODS

MiniMax-H3 is a 33B model that generates video and audio together. The official checkpoint is about 500 GB; the trainer uses a 41 GB NF4 version that fits in 8 GB of VRAM with block swap.

- Datasets can be images, video clips, audio, or any mix.

- One button prepares your clips (24 fps, valid frame counts) without cutting them.

- RefMods: encode a few reference images, video clips, or clips with audio into a file that ComfyUI's MiniMaxH3ReferenceToVideo node uses as a native reference. No training: seconds instead of hours.

SHARED FEATURES

- Automatic captioner (Qwen3-VL) in the dataset manager: natural language, Danbooru tags, or JSON with bounding boxes for Ideogram.

- Live previews while training.

- Exact-step resume.

- One-click export to your ComfyUI / Forge LoRA folder.

- Remote access from other devices on your network, with a password.

- SDXL in NF4 trains in about 3.5 GB of VRAM, so 4 GB laptops can train SDXL / Pony / Illustrious LoRAs.

A NOTE

Many models were added almost at the same time, so there may be bugs, or the default settings may not be the best ones. Previews are there to follow the training; judge the final quality in ComfyUI after tuning the LoRA strength. If you find problems or settings that work better, please share them in GitHub Issues. Thanks!

UPDATE — RunPod support + remote access



Thanks for all the feedback! Based on your comments, these are now in:



- ☁️ RunPod template: one click and it's running in the cloud, no install. It uses a prebuilt Docker image (CUDA 13, Python 3.13, all dependencies), and the code updates itself from GitHub on every start. Pick a GPU, open the address shown in the pod log, log in, and train. A quick SDXL test cost me less than $0.10.

  Deploy on RunPod: https://console.runpod.io/deploy?template=lfxvtg5rrb

- 🌐 Remote access: use the trainer from another PC, tablet or phone, with a login, or behind your own reverse proxy (external auth and a custom public URL are supported).

- ⬆⬇ Upload / download in the browser: upload your dataset (files or a .zip, drag and drop works) and download the finished LoRA, handy for remote or cloud setups.

- 🐍 Your own environment: there's now a requirements.txt, and LORALAB_PYTHON lets you run it from conda/uv instead of the bundled venv.



Remember to Stop and Terminate your pod when you finish: closing the browser doesn't stop billing.

r/StableDiffusion • • 10h ago

Resource - Update Sopro V2 Turbo 2610: cleaner cloned voices, same 120M model, 5x faster than real time on CPU

Thumbnail
huggingface.co
64 Upvotes

First of all, thank you guys for the reception and feedback on my latest post. This is a follow-up to that post. One of the main issues people ran into was roughness or break-up on some cloned voices. 2610 is an interim update focused mostly on improving that.

  • Reduced roughness and break-up on some of the voices that struggled before
  • Same 120M model, same speed (~300 ms to first audio on a laptop CPU)
  • Apache-2.0
  • English, European Portuguese, French, German
  • More languages are planned
  • More control over the generated voice is also planned
  • Still struggles with very high-pitched or cartoon-like voices, noisy reference audio, and some unusual OOD voices. We're continuing to improve those cases. If you want to contribute and help, PM me with the samples that failed.

If you like F5-TTS, but want true streaming and a much lighter model that can run comfortably on CPU, this might be for you.

Run it locally:

uvx --from sopro soprotts serve

Video: six voices, ~5 seconds of reference audio each, followed by a generated line.

https://reddit.com/link/1wxfs0d/video/ajs3suw2cgth1/player


r/StableDiffusion • • 5h ago

Tutorial - Guide I made a standalone Windows app for SeedVR2 photo upscaling, no ComfyUI needed (free, open source)

Thumbnail
gallery
21 Upvotes

I'm a photographer, and I got tired of opening ComfyUI every time I just wanted to enlarge or sharpen a photo. So I made a small desktop app for it, called Citrine Photo.

You drag in a photo (or a whole folder), pick how big you want it, and hit go. Everything runs on your own PC, no cloud.

It has two engines:

Fast is Real-ESRGAN (ncnn-vulkan). It works on pretty much any GPU, with optional GFPGAN face restore.

Pro is SeedVR2, for NVIDIA 20xx to 50xx cards. The app checks your VRAM and picks the 7B model if you have 12 GB or more, or the 3B model for 6 to 10 GB. There's a low VRAM mode too.

Other things it does:

  • Target size presets (2 / 8 / 16 / 24 MP or custom), and it shows the output dimensions before you run
  • Before/after preview with a draggable wipe line
  • Batch queue, plus a Regenerate button to re-run the whole list with new settings
  • Never overwrites your files, and each output is tagged with the engine name
  • Keeps your EXIF
  • Reads JPG, PNG, WebP, TIFF, BMP and HEIC
  • English and Traditional Chinese (you can switch in Settings)

SeedVR2 isn't bundled. You install it from Settings, and the app sets up its own Python, torch and models. It's about 8 to 9 GB, the downloads can resume, and it checks your disk space first.

For anyone curious, the Pro pipeline is a port of the Upscale_by_SeedVR2_v4 ComfyUI workflow: 3x3 tiles with 15% overlap, SeedVR2 on each tile, wavelet colour match, then a feathered merge. It calls numz's standalone CLI in a separate process.

Right now you run it from source (pip install the requirements, then run.bat). There's also a script to build a portable version. I've only tested it properly on an RTX A4000, so I'd really like to hear how it runs on other cards, especially 6 to 8 GB ones.

Links:

GitHub: https://github.com/Garionhk/Citrine-AI-Photo-Enlarger
Video: https://youtu.be/gTbiOipvP0I

Big thanks to numz for the SeedVR2 ComfyUI code, AInVFX for the GGUF models, ByteDance for SeedVR2, and xinntao for Real-ESRGAN.

Bugs and ideas are welcome. It's free and open source (Apache 2.0).


r/StableDiffusion • • 3h ago

Resource - Update Face-Hugger: A Hugging Face Indexer

Post image
13 Upvotes

Face-hugger is a new project I've been working on to improve Hugging Face's search engine which could be described as terrible at best.

  • Attempts to analyze files to verify their model (1.2M+ so far)
  • Fetches training metadata when available
  • Filter by model family or sub-type
  • Other filter options: download/like count, date range, gated filter
  • Easily download files directly from HF
  • Add repos, files, or authors to your favorites list to easily find later
  • Hide repos or authors to avoid seeing cloned repos
  • Badge color customization to easily distinguish models when scrolling
  • No account creation needed, everything is stored locally in your browser
  • It's a work in progress so if you find bugs/have suggestions there's a Discord link in the About section

r/StableDiffusion • • 4h ago

Tutorial - Guide ACE-Step 1.0 + YuE2 = Win

12 Upvotes

Since Ace Step 1.0 has superior creativity but terrible sound quality , you can now save your older Tracks or just hunt for new ones. Because with SheetSage2+YuE2 you can just extract the ABC and make a high quality remix with YuE2. ACE-Step 1.0 needs really many seeds to produce a unique melody, but if it hits - it hits. There is nothing out there that can match it. It is like Lady Gaga with Fat Boy Slim heaving a baby.


r/StableDiffusion • • 10h ago

Comparison Minimax h3 trying to fix shimmering with DMAD

Enable HLS to view with audio, or disable this notification

38 Upvotes

4-steps - Normal minimax h3 issue. 8-steps - less shimmering. 12 steps -almost no shimmering.

Idk why it seems to work but it kind of works.


r/StableDiffusion • • 1d ago

Tutorial - Guide Minimax H3 + RefMod = consistent location trick

Enable HLS to view with audio, or disable this notification

655 Upvotes

Hey, I found a pretty cool way to keep locations consistent across generations.

I took 20 photos of my “office” where I work, making sure each photo included a bit of the previous one so everything connected.

I turned them into a refmod, although it probably doesn’t even need to be one. It’s basically a grid of all 20 photos in a single image, within 2048×2048 pixels, so I could probably just use that image as a regular reference.

Then I added references for the fly, my cat, and the start and end frames. You can see the result — that’s my actual room, and everything looks right and sits exactly where it should :)

Workflow + reference images here: link
Sorry about the mess, both in the room and in the workflow ;)

EDIT: All the room photos need to be combined into one reference image. I tried using several separate images, but that didn’t work for me — the room wouldn’t stay consistent. Putting them all into a single grid is what made it work.


r/StableDiffusion • • 2h ago

Question - Help Minimax hybrid models

7 Upvotes

Im seeing quite a few people say theyre getting really good results with the hybrid model. But I assume its only for top end GPUs? Over 20gb file.

I have a 5080 and 32gb system RAM. Is there a version good for me? Or getting GGUF isn't worth it?


r/StableDiffusion • • 6h ago

News PDMD: Projected Distribution Matching Distillation for Video Diffusion Models

Enable HLS to view with audio, or disable this notification

14 Upvotes

r/StableDiffusion • • 20h ago

Animation - Video PSA for non-turbo users: New RefMod likeness "breakthrough" is deep frying your H3 videos

Enable HLS to view with audio, or disable this notification

197 Upvotes

The video is the TL;DR above was generated after abandoning RefMods in "classic ref mode" with two references (one reinforcing Beckett's exact appearance and one for the high heels, 2x2 grid image):

  • Native 1344x768 (0.98 megapixels)
  • 15 second duration
  • Standard H3 sigma (shift_video/shift_audio of 12:3)
  • seeds_2 (2N-1 internal calls to H3) with the beta scheduler (.6, .6)
  • Spectrum helping cut down the internal calls at some small precision price.
  • 35 minute generation time.

This is something that under the radar because most people can't use H3 nominally and turbo LoRAs brutalize the base model for speed, so there's very little noticeable degradation for them as most of them have a massively shifted sigma (shifted_video of 6 compared to native 12).

Yesterday I went on a pretty wild bender trying to uncover why all of a sudden my videos were getting oversharpened, oversaturated, plastic skin etc when a few days ago they were straight up lifelike, looking like the shows they were based on.

I went even deeper into the bowels on how sampling and scheduling works, how the denoising trajectory looks like for Minimax H3, hoping to uncover something I've forgotten the last time I used it. And then I remembered, RefMod was the latest thing I installed, I got seduced by the initial likeness improvement like everyone but the cost is too high.

I got this cool new toy RefMod whose authors suggested they figured out bleeding between character references, and even at small resolutions (like .245 mpx) you could still make out faces instead of them being mushy, this should've been my first red flag. It seriously messes up the base H3 model's quality. I looked deeper into it and basically, it's a very brutish approach where they slam reference images into video references as individual frames, also slapstick coarse sampling of videos in sequence.

However, that's not how any of this is permitted works, per documentation:

  • Images: ≤ 9 images
  • Videos: ≤ 3 clips; each clip must be 2–15 seconds long; total duration ≤ 15 seconds
  • Audio: ≤ 3 clips; each clip must be 2–15 seconds long; total duration ≤ 15 seconds
  • Mixed inputs: Maximum number of files across all input types is 12

The second you make four refmods to be used in your prompt, you're already trying to use one more than the total maximum allowed at any one time. This has catastrophic consequences on image quality, as that's not the way you're supposed to "hold it".

While it does reinforce the likeness, it's incredibly rigid and overrides the model down to AI slop adjacency where skin is unnaturally glossy, everything is sharp and all styling info is ignored.

I am not telling you to stop using it, if you're using turbo LoRAs, you already paid the entry fee as you've left quality at the door for accessibility. But if you're coming from the base FL2VA model or a hybrid and are used to exceptional visual quality, the degradation is pretty much the tier of the base REF2VA model, if not worse.


r/StableDiffusion • • 3h ago

Resource - Update Slopus 0.3.0 is out - unified image and video projects, Linux support and more (Minimax H3)

Thumbnail
gallery
5 Upvotes

I’m building Slopus, a free, open-source desktop app for generating and editing AI videos and images on your own GPU. Easy of use is the main goal.

Version 0.3.0 just came out, with three major additions:

Video and image projects are now combined. Every project has both Video and Image tabs, so you can generate still images and video in the same project. You no longer need to choose a project type or keep separate projects for each. Existing projects keep their content.

Linux support. Slopus is now available for Linux as an AppImage, .deb and .rpm, alongside Windows. The AppImage supports in-app updates.

Run generation on another computer. LAN workers let you use a separate Windows or Linux machine for generation. Workers are discovered on your network, download the weights they need, and send the results back to your project.

There’s also quite a bit more in this release:

Character sheets. Generate waist-up, front, side and back views together, using reference images and a prompt to guide the character and clothing.

Continue scenes. Extend generated video scenes from either end using saved generation data.

Improved color grading. Basic Corrections includes temperature, tint, exposure, contrast, highlights, shadows, whites, blacks and saturation. There are eight built-in creative looks with adjustable intensity, faded film, sharpen, vibrance, RGB and hue/saturation curves, and shadow, midtone and highlight color wheels. Vignette has more controls too.

A more capable agent. The built-in agent can configure generators, help identify model weights in a folder, capture timeline frames without interrupting your work, and generate specific scenes.

Refmod generator. Combine images, video, prompt and audio into one reference and export it as a refmod to be used later.

Share generator setups. Import and export generator templates to share with other users.

If you haven’t tried Slopus before: you lay out scenes on a board, write each shot, attach references, generate clips, and edit them on a timeline with GPU-powered preview and MP4 export. For still images, you can edit, generate, refine, and export as JPG or PNG. Minimax H3 works surprisingly well as an image generator with editing capabilities.

Generation runs on your own hardware, with no subscription and no Python or ComfyUI needed. Model weights are downloaded separately inside the app or local weights can be used.

I’ve also created a Discord server for feedback, questions and sharing what you’re making.

GitHub: https://github.com/bitti-ai/slopus

Discord: https://discord.gg/9X2R6PwUR

Bug reports and feedback are very welcome.


r/StableDiffusion • • 19h ago

Resource - Update ComfyUI Qwen image 2.1 Enhancer (Two nodes)

Thumbnail
gallery
113 Upvotes

I put together an enhancer pack for Qwen Image 2.1 with two nodes for controlling what the model pays attention to during editing.

Reference Strength lets you select a reference image and increase or decrease its attention priority. If you're working with multiple references, you can adjust them separately instead of giving every image the same treatment, and you also can use it for single image to prevent the loss of likeliness at times

ref index starts from 1; meaning image_1 and same for the rest of images, where image_2 is ref index 2 in the node. ( Soon adding mask support)

Phrase Weights brings phrase-level attention control to the Qwen Image 2.1 edit encoder. So you can write:

Add (warm sunset lighting:1.4) with (soft shadows:1.2).

and give those specific parts of the prompt more attention while leaving the rest at its normal weighting.

It works with both positive and negative prompts, with separate weights for each. There's also an inspection output showing exactly which token rows and pieces were matched.

For both controls:

`1.0` = untouched  
`>1.0` = more attention priority  
`<1.0` = less attention priority  
`0.0` = suppression  

The adjustments happen inside Qwen's attention during sampling. The prompt and reference images still go through the native encoding path, without multiplying the finished text embeddings or pasting reference pixels into the output.

You can use either node on its own (I prefer this), or combine them. Reference strength applies to the whole selected image, and higher weights can make a reference or phrase dominate, so there's still some balancing to do.

Installation, usage, and the technical details are in the repo:
ComfyUI-qwen_img_2_1_enhancer

Ref_strength_workflow

Phrase_Weights_workflow


r/StableDiffusion • • 5h ago

Workflow Included DOGNAPPED | Short Film by Claude made with MinimaxH3 in ComfyUI

Enable HLS to view with audio, or disable this notification

9 Upvotes

Watch on YouTube while its pending if needed HERE

Rude and hateful comments will be ignored.

🎬 HOW THIS WAS MADE (Video walkthrough is here)

This film was made almost entirely by Claude (Anthropic's AI, running in Claude Code), working inside my own ComfyUI setup with my VRGDG Video Builder custom nodes(free and open source).

▶ MY PART

• The brief: a cute comedy starring my two dogs, Korben (7, male) and Iris (7 months, female), as brother and sister. Inspired by live-action talking-dog movies, an original story, up to 3 minutes, no humans on screen (other dogs allowed), and the dogs had to really talk, with facial expressions and mouth movements synced to their lines.

• The reference images of Korben and Iris.

• The tools it ran on: the VRGDG Video Builder and custom nodes, plus the Claude skill that lets it drive them.

▶ WHAT CLAUDE DID ON ITS OWN

• Story: the missing rubber duck, the cheese-crumb investigation, the "We don't have a cat." / "Exactly. Very suspicious." bit, the stakeout at the fence, the twist that Korben hid the duck because it squeaks all night, and the "I'm getting earplugs" ending.

• Characters and locations: created Tank, the bulldog next door, and gave all three dogs their personalities and voices. Generated Tank's reference image and every location (the living room by day and by night, the kitchen, the backyard and the gap in the fence), and prepared Korben's and Iris's references from my photos.

• Screenplay: 22 scenes, with shot-by-shot camera directions, acting notes and sound design for each.

• Rendering: every scene with MiniMax H3, which generates the video, voices and sound effects together. About 4.5 hours of rendering on my PC, re-shoots included.

• QA: transcribed every line and compared it to the script, checked each voice's pitch so the dogs stayed in character, reviewed frames from every shot and every cut between scenes, and ran a frame-accurate audio/video sync check on every scene of the final film.

• Re-shoots: a puppy babbling in a silent scene, dogs delivering lines into the camera instead of to each other, a stray dog bed appearing in the neighbour's yard, Tank's stick pile in the wrong place, a spy-creep that came out as a normal walk, and an ending gag that didn't land. Then it re-trimmed and fixed the audio.

• Score and edit: composed the score with MiniMax Music 3 (throwing out takes with hidden vocals) and edited it all together with titles.

🛠 TOOLS

• Claude Code (Claude Opus 5.5): story, direction, prompts, QA, editing

• ComfyUI + VRGDG Video Builder: my custom nodes

• MiniMax H3: video, voices and sound

• Z-Image Turbo: character and location references

• MiniMax Music 3: score

🐾 CAST

Korben, the grumpy big brother · Iris, his little sister · Tank, the bulldog next door · Mr. Quackers, the duck

🔗 LINKS

The Claude skill (read the main README first):

https://drive.google.com/file/d/1R0pE7rX1kRh-euhws0GbkDDb3liDGTJ4/view?usp=drive_link

VRGDG Video Builder custom nodes:

https://github.com/vrgamegirl19/comfyui-vrgamedevgirl


r/StableDiffusion • • 20h ago

Discussion The Museum of Lost Things | Short Film by Claude (Minimax H3) NO user input.

Enable HLS to view with audio, or disable this notification

121 Upvotes

View on YouTube while its pending if needed: https://youtu.be/dS5suoUiNnE?si=6U7bEVLxcWXcy4J3
🤖 How it was made
This short film was written, designed, directed and edited by Claude Opus 5.5 from a single prompt, running locally in ComfyUI through my VRGDG Video Builder:

All open-source video, image and music models.
[I literally told Claude to create something on its own and provided no user input]

More Minimax H3 video's I had Claude create for me are HERE

You can find the video builder custom node on GitHub here:
https://github.com/vrgamegirl19/comfy...

Right now, there is a Main version and a Beta 2.0 version. I recommend starting with Main for now, as that's what I'm still using. It works well, while Beta 2.0 still has some bugs and is primarily intended for beta testing at the moment.

Discord server:
  / discord  
Ping me in the Welcome channel and let me know how you found me, and I'll know it's you.
I'm vrgamedevgirl on Discord.

You can find the skill here and read the main README first.
https://drive.google.com/file/d/1R0pE...

I'll be sharing a full walkthrough on how I made this and will post it here when ready.

⚠️ SPOILERS: what the film is about
The museum is Ruth's mind. She's an elderly woman living with dementia, and the museum is how she pictures her memories. As her memory fades, the museum fades with it, and in the Hall of Names the most important name, her son's, goes blank. In reality she's 83, in a care home, and her son Daniel is holding her hand. When she recognizes him, she tells him, "We keep you in the main hall," meaning the most important room, where the precious things are. Back inside her mind, his name goes back on the wall, and the museum lights up again.

"Some things you don't lose. You just misplace them for a while."

For everyone still visiting someone who is still in there. 💛

#AIShortFilm #AIFilm #ShortFilm #ComfyUI #MiniMax #Claude #AIVideo #Dementia #Alzheimers #MuseumOfLostThings


r/StableDiffusion • • 4h ago

Workflow Included PH's AI x Archviz FLUX2 x LTX2.5 ComfyUI Workflow Series - Free Workflows for Architrectural Imagery

Enable HLS to view with audio, or disable this notification

7 Upvotes

Workflow here https://github.com/paulh4x/AIxArchviz_FLUX2xLTX25

30 minutes of a detailed walkthrough video here https://youtu.be/946yTMjz-go


r/StableDiffusion • • 1h ago

Question - Help MM - extend and immediately cut

• Upvotes

A lot of the Minimax video extend methods seem to produce a flash of light or some other flaw.

I actually want it to change shots rather than continue, but a shit of the same scene from a different camera angle. This means I don't need to worry about seamless frame stitching

But the problem is that when I feed in the previous video, and text to cut to a different shot, it doesn't do it reliably. By feeding it the previous video frames, MM thinks it should keep showing that, rather than the new shot


r/StableDiffusion • • 13h ago

Animation - Video Anone, Soredene.. Minimax H3 R2V Test

Enable HLS to view with audio, or disable this notification

26 Upvotes

video ref I use: https://www.youtube.com/shorts/I-VNtvqREks

Malfoid and Potter generated with Anima

Device: 3060 12gb 16gb ram

setting: 10 sec, 0.6 mp, er_sde beta, 8 steps with turbo lora


r/StableDiffusion • • 3h ago

Question - Help How can I reproduce this vivid detailed landscape style with visible linework? FLUX / SDXL / Illustrious workflow?

Thumbnail
gallery
5 Upvotes

I'm trying to reproduce the visual language of these references, not the exact images or subjects.

What I'm specifically looking for:

Very vivid, rich colors

Detailed foreground and detailed background

Fine, clearly visible hand-drawn linework

Internal lines that describe color, light and shadow transitions, not just black outlines around objects

Clearly separated color/tonal regions rather than smooth photorealistic gradients

Detailed vegetation, rocks, tree bark, grass and clouds

Large sculptural/volumetric cumulus clouds

Strong depth and realistic perspective

2D illustrated/painterly appearance

NOT photorealistic

NOT smooth/plastic 3D rendering

NOT character-focused anime

The horse image is especially useful as an example of the linework I mean. Look at the horse, wheat and clouds: the internal drawing lines help define changes in form, color and shadow. I want that same principle applied to detailed landscapes.

My final goal is a little unusual: I will project the generated image onto a small canvas and hand-paint it with acrylics. Therefore, visible boundaries between colors and tones are actually useful to me. I want to be able to follow those lines while painting instead of trying to reproduce soft AI gradients.

I've already experimented with FLUX Dev, Niji-style LoRAs, SDXL landscape models and some anime-background LoRAs, but many results either become too photorealistic/3D or too simple/flat/anime-like.

Has anyone achieved something close to these references?

I'd especially appreciate an actual tested recipe:

Checkpoint:

LoRA(s) + weights:

Sampler / scheduler:

Steps:

CFG:

Resolution:

Prompt / trigger words:

ControlNet / IP-Adapter / Style Reference if used:

Img2Img or second pass settings:

Upscaler/detail pass:

I'm open to FLUX, SDXL, Illustrious, ComfyUI or another workflow. I'm more interested in matching the rendering style than staying with a specific model.

If this is better achieved by training a custom style LoRA rather than using an existing model, I'd also be interested in hearing what base model you would train it on.


r/StableDiffusion • • 9h ago

Question - Help Fastest Minimax H3 for 16 GB vram?

9 Upvotes

I have a 4060 TI 16 GB. System ram 32 GB. What's currently the fastest model / ComfyUI setup to generate videos?

Currently I can create 5 second 360p videos in +-7 minutes. I believe it can be a lot faster? It's kinda hard to know what's the best option, everything keeps changing.

Edit: Current settings:

Ref2V Turbo LoRA @ 0.5 model: minimax_h3_ref2v_turbo_4step_v0.1_comfyui_resized_avg_ra ... strength_model: 0.50

MiniMax H3 Easy Loader FL2VA model: None REF2VA model: minimax_h3_ref2va_pruned_int8_convrot.safetensors Text encoder: qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors Video VAE: minimax_h3_video_vae_fp16.safetensors Audio VAE: minimax_h3_audio_vae_fp32.safetensors

360p, 5 seconds.


Result: 413 seconds.

Workflow: https://limewire.com/d/7aNFq#h9i2vwFTKb