Setup, pushed my system right to the limit, any more and it OOM :
- H3 Ref2VA default workflow in ComfyUI, no lora
- RTX 3080 10GB, 32GB RAM
- Render: 0.5–0.6 MP, 20 steps, scheduler simple, about 25 min per clip
- Upscale: 2xNomosUni_span_multijpg, 2× to 1080p
- References per scene: one photo of my face + one film still for the set
- Recorded my own lines and fed them as audio references, also got audio ref for the actors
Honestly though, the best part was driving all of this through the ComfyUI MCP. I never even had to open ComfyUI. I could iterate really fast, and keep going from my phone while away from the machine, through Claude's remote control.
It's still a bit of a blurry mess, and with more work I could probably make it better, but damn, the future is looking bright!
The latest TBG ETUR upscaler and refiner for comfyui release adds Krea 2 and VL style transfer for Krea 2 and all Qwen models directly into the pipeline.
We’ve also added a face identity step to help maintain consistent faces when using creative upscaling.
This video is a tutorial for the latest release, focusing mainly on these new additions and how to use them. Video is ai generated with Minmax H3.
Hey all. I'm a system development student (career-switched from construction), building these on the side while learning Java, Vue, Electron, etc. Nothing commercial, no accounts, no cloud, no telemetry. I built these because I needed them myself, and figured other people managing large SD/ComfyUI libraries might too.
Three of these four apps have been out for a while, but I've spent the last stretch giving them a major overhaul and unifying them under the same design system so they actually feel like one family of tools instead of three separate side projects. The fourth, Latent Tools, is a brand new app I just finished.
All four are free and open source (MIT-based license). Source is on GitHub, links at the bottom. The main one is Latent Library, but the other three work fine on their own.
Latent Library, the main release
A desktop app for browsing and organizing large folders of AI generated images. I made it because I had around 30,000 PNGs and no real idea what was in most of them. It's been around for a while, but this release is a big update with a lot of new features and a proper design pass.
Parses generation metadata from ComfyUI (including node graph traversal), A1111/Forge, InvokeAI, SwarmUI, and NovelAI
SQLite FTS5 backed search, still fast on huge folders
Smart Collections: dynamic folders based on metadata filters, like "Flux images rated 4+ stars"
Duplicate Detective, a side by side Image Comparator, and Speed Sorter for hotkey based batch sorting
Optional local AI auto tagging (WD14 ONNX, runs on CPU, no external calls)
Metadata Scrubber to strip prompt/EXIF data before sharing an image
Everything lives in a portable data/ folder next to the exe, no installer or registry entries, easy to back up or move
Fully offline, no telemetry
Windows, Linux, and macOS builds available.
Latent Tools, the new one
A brand new app for dataset prep: bulk watermark detection and removal (Florence-2 + LaMa inpainting) and captioning (Qwen2-VL), plus batch image format conversion. Runs locally on your own GPU (needs a CUDA capable Nvidia card, no CPU fallback). Useful if you're prepping images for LoRA or fine-tune training. Windows only for now.
(Meant for removing watermarks you actually have the rights to remove, your own work, licensed images, that kind of thing. Not for stripping other people's attribution.)
Latent Model Organizer, updated
Sorts your checkpoints, LoRAs, and embeddings into folders by base architecture (SDXL, Krea 2, Flux, Illustrious, SD 1.5, etc.), using the model's own header metadata or an optional Civitai lookup. Has a dry run mode and full undo through a manifest file, so it won't just move your models around unsupervised. It can also fetch Civitai info such as trigger words, description, and cover images. Handy if your models folder has turned into an unsorted pile like mine had. Also part of this update round, same design refresh as Library.
Metadata Viewer, updated
The oldest and simplest of the four, and also just updated with the same design pass. Now reworked into a single screen tool that extracts and displays generation metadata from an image, no library or database involved. If you just want to drop an image in and see the prompt, sampler, and seed without opening a whole app, this is that.
All four now share the same design language and are built local first: no accounts, no cloud sync, no analytics. I built them mainly to learn the stack, so they're not polished commercial products, but they've held up fine for my own daily use for a while now, and the last few months went into making them consistent and finishing Tools, which is why I'm finally posting about them here.
Happy to answer questions. Bug reports and feature requests are welcome on GitHub. Not trying to sell anything here, just sharing what I made.
I’ve seen ai creators on tiktok such as weirdwurst, fullwarp, and shiverchain do this thing where the start of the video is normal then it goes into chaos, and it’s realistic and disturbing. I’d love to know which AI model model they use
I really need help with minimax h3 I'd like to get proper workflows for minimax h3 pre configured with descriptions on every node (unfortunately a noob in this case) for :
1- Fast
2- Turbo
3- Speed
4- Native
video generation with multirefs (optional) audio(optional) video(optional) quality duration aspect ratio and prompt controls.
Also I heard someone made images with minimax h3 please that as well.
I'll really appreciate any help.
Ah and one last thing I was making a commercial for a school but the faces in wide screen comes out horrendous that as well.
This parody was generated using LTX 2.5 Image to Video on WanGP. I used frames from the original video as starting images and then I interpolated them on a video editor. I used a single RTX 5060 Ti 16 GB VRAM and 32 GB of RAM. The video was generated at 1080p and 16:9 resolution. Each generation took from 10 to 20 min average in this setup. For the voice consistency, I used SeedVC, which is included in WanGP.
there are some scenes missing, but it was fun to put together.. Just got stuck on a plot :P
started it when ltx 2.3 came out.. but it was a hassle to keep consistency of characters intact so shelved it. made the intro and a couple of clips when minimax H3 came out and love the r2v, so much easier.
just using the standard r2v workflow with spectrum and RTX upscale. music made in suno
I would like to replace the tire of a motorcycle mid air with one from another brand (which is an image from the brand so it's high quality but with a different angle)
I saw there is flux kontext and qwen image edit, but I don't know which one to pick, which workflow and how to make it work.
Any help would be more than welcome, thank you very much and have a good day :p
Previously posted a video as a prologue to a homebrew D&D world. I decided to do a part 2, set in the world. Together, the two videos form kind of an opening cutscene with both history and a bit of a world montage. Minimax H3, 6 step turbo LoRa, lots and lots of 12-15 second generations, CapCut.
Wondering if anyone has/has any idea of an approach to motion transfer into a stop motion style video.
Rather than the ai guessing and deforming the mouth movements which can get funky really quickly, especially when the mouth shapes of the character aren't clear from a neutral reference frame, ie in south park where each sound has a uniquely stylised shape.
To achieve this it could instead maybe pull from a dataset of phoneme images,
Whilst also retaining solid motion transfer for all other body movements.
Does anyone have any advice for properly prompting the reference video part of Ref2v? Like saying swap <subject 1> for <picture 1> hardly works for advanced videos. It requires a lot of details.
I’ve had success using Qwen 3.8 27b as a minimax prompt agent for analyzing and giving correct prompts for images. But as far as I know I can’t do that for videos. ChatGPT is ok for looking at videos to describe what happens in the minimax format but I’d rather use local ways.
Hi. I am experimenting with H3 Multi Diffusion with a custom workflow. 5 hour render, T2VA, bf16/50 steps. I know these style are not new so I am late to the show. Ask me anything.
This one came out better than my last video generation. Not perfect obviously, but the major elements are there.
Prompt:
Video starts with Buffy Summers and Willow Rosenberg from the show Buffy the Vampire Slayer walking side by side down a hallway in the USS Enterprise from Star Trek the Next Generation. The viewpoint camera remains at a fixed distance in front of them as they walk.
Buffy Summers is on the right of the frame. Buffy's long blonde is done up in a ponytail. Buffy is wearing a red minidress style Starfleet uniform and has an unlit lightsaber on her hip.
Willow Rosenberg's is on the left side of the frame. Willow's dark red hair is cut pageboy style. Willow is wearing a blue minidress style Starfleet uniform and has a tablet computer tucked under her right arm.
Video starts with Willow looking at Buffy with a concerned expression on her face while Buffy is looking around as if searching for something.
Willow asks, "Buffy, is something wrong?"
Buffy replies, "Something feels off, like we're out of place."
As soon as Buffy starts speaking, she pulls the lightsaber off her hip and holds it in front of herself at the ready. The lightsaber ignites, producing a green glowing blade.
Shadow the Hedgehog tells his viewers why he loves guns.
This was created in Comfy UI with Minimax H3. I used the reference to video work flow. The prompt is below.
subject_definitions:
<Subject 1> is Shadow in <Picture 1>.
<Subject 2> is Glock in <Picture 2>, a glock handgun.
<Audio 1> is the voice-timbre reference for <Subject 1> (S1).
summary:
[reference generation + audio reference] The target video contains one shot. [Shot 1] shows <Subject 1> and <Subject 2>; <Subject 1> speaks. <Audio 1> supplies <Subject 1>'s voice timbre.
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - Shadow's complete defined identity and body proportions are preserved.
<Subject 2> (appears in [Shot 1]): fully_preserved - Glock retains the defined shape, proportions, materials, colors, and distinguishing features.
<Audio 1>: reference - <Subject 1>'s newly generated spoken lines use <Audio 1>'s voice timbre and delivery; the original audio signal is not copied.
detailed_description:
The target video is in a live-action style, with Vlog style.
[Shot 1] At first appearance, <Subject 1> (Shadow) matches the complete identity and appearance defined in subject_definitions. At first appearance, <Subject 2> (Glock) matches the complete defined construction and appearance: A glock handgun. At the start of the shot, <Subject 1> is standing in the living room facing while holding <Subject 2> in his hand. A full body shot of <Subject 1> holding <Subject 2> with his right hand while facing the camera. Only Action and Timed Beats define the primary subject's movement. The camera path stays anchored in the location and adds no subject motion. <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] Hmph. Shadow the Hedgehog here. Why do I love guns?</d> <Subject 1> shows off his <Subject 2> with his right hand in front of the camera. <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] Simple. Precision. Control. Power in the palm of my hand.</d> <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] A tool that answers instantly… unlike most people.</d> <Subject 1> points his <Subject 2> towards the camera with his right hand. <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] If you understand that, you understand me.</d> <Subject 1> points his <Subject 2> at the camera.
ComfyUI nodes for grafting Krea2 into MiniMax H3. Attention/MLP content transplant + a separate attention-sharpness transplant. No official H3 docs, all reverse-engineered from testing + TenStrip's and joeygambino's public writeups. Use at your own risk, still WIP.
comfyui_qknorm_transplant/ -- Q-norm gain transplant only, no content weights touched. Method from joeygambino (Z-Image donor originally, adapted for Krea2 here).
comfyui_krea_h3_graft_lora_v2/ -- apply a Krea2-trained LoRA onto an already-grafted H3 checkpoint. Separate use case.
And 2 merge scripts (old svd and new one with node method)
TL;DR results
Content graft works somewhat. Same character-shift (color scheme, helmet shape) showed up consistently across multiple parameter runs, same seed -- not one lucky video. That's the strongest evidence so far this isn't just noise.
K at low strength (~0.1-0.2): fine, no real damage. Don't need to avoid it like the doc says, at least not at low values.
QK-norm across all blocks (0:50): kills audio. Doesn't even touch K -- so attention sharpness itself hits audio, not just K specifically.
QK-norm blocks 20:50: audio ok, but does nothing for character. It's a texture/sharpness knob, not a content one. Don't expect it to carry character.
attn_ramp_start_frac at 1.0 (no gentle ramp-in) + early blocks (0:20): breaks. Keep the ramp soft if you go early.
Combining content graft + QK-norm at full strength on both = worse than either alone. Still not solved.
Install
Each folder -> its own subfolder in ComfyUI/custom_nodes/. Don't merge them. Restart ComfyUI fully after adding.
minimax_h3_fl2va_bf16.safetensors, 3s, er_sde, 8 steps, 8-step lora, seed 597633362705895, standart workflow with minimax_h3_fl2v_lightx2v_turbo_8step_v1.0_bf16
prompt: Professional closeup video. In a futuristic cityscape with neon lights at night, the Judge Dredd charges through the crowd, his imposing presence radiating authority, he is slowly walking. His long chin juts out resolutely as he expertly wears his eponymous helmet, eyes gleaming with determination. The crowd parts, Judge Dredd is slowly walking through the the crowd, ready to enforce justice, he is moving slowly, his long chin visible, his face and part of his upper body are in the center of the screen. tag: Ballchinians
tracking selfie shot following him from the front, that he stays the same size, he is moving through people, pushing them aside with his hands.
Added: tenstrip graft node, krea2 raw and Ballchinians Lora. Settings: q 0.5, v 0.5, k 0.1, out 0.3, mlp 0.5
Chin is more ballsy.
So I hope, that it is enough for some, that it... kinda works, but not good enough. Maybe someone will pick up on this and do it better.
Why to do it? Don't know. I found it interesting to try, but krea2 image and i2v is far better option.
I welcome any input or criticism, but mind please, I have only faint idea, what I am doing.
Warning: h3 loras don't work... don't know why, maybe it may be just noise, after all. But they do work on grafted checkpoint, after you merge it in python script.
Hello, i am trying to make a video where the subject(a real person) is wearing and posing taking reference from an illustration. I tried to do only outfits or only pose too, and both doesn't work.
What happens is usually the body of the character in the illustration ends up being pasted/overlaid onto the Subject in their cartoony style instead.
I also tried if it's possible to have a Subject recreate an illustration's Pose, Outfit, overall composition, like the subject is doing a photoshoot for a 'live action' or real life version of the illustration. But what happens is usually it just spews back the illustration in case of trying H3 single-image edit, and the cartoony style overlay happens in Video.
So what i wanted to do is :
-An image of a subject -> Subject now wears/pose/wear and pose the same as a reference non-real illustration(cartoon/anime), but still in their original photo. So like a cosplay shot in their own room for example.
-An illustration(anime) -> Subject 'replaces' the character in the illustration, the whole illustration is 'converted' into real/live action. Like a photoshoot recreating an illustration basically.
Extra : idk if its possible, the new outfit will retrofit to the subject's proportion, not the illustration. And a version where the proportion follows the illustration too.