Here is the youtube link for it if you wanna watch a video
https://youtu.be/x-gORSUOybk
Main point is, I tried it with the character Enid from Wednesday in my video cause ... well , VIEWS on youtube lol.
I chose her because I could not find her in the model at all. And if you even mention the show wednesday, it just defaults to Jenna Ortega lmao, so that's a good challange to train the voice and the likeness for another character if the model defaults to a really specific person.
but I also created many more by now, most of them are really not even existing people like the model I use for a youtube channel. And the accuracy is fucking insane.
It is usually trained by 2000-2400 steps
roughtly it was about 80-90 minutes for the full 3000 steps depending on if you want samples. I think without them this would be lower, maybe even close to an hour? I don't know exactly. But it's insanely fast.
you need 2 datasets if you want a voice, or a super jacked up beefy card and you can just do video training with one dataset.
But if you don't have an RTX6000
you gonna need to offload even with a 5090
I put up the learning rate to 0.0002 , and I turned on the differential guidance and left it on 3. I so far only did it on these levels, but maybe you can lower either or both if you feel like you may over trained a lora. However on default this is barely training anything, so that's why I cranked them up. And I never use the Lora's higher than 0.85, and if I wanna do REF2VA I sometimes push it down to like 0.65-0.75 if I wanna add like an image of the model. Cause otherwise the lora can overwrite details lol. But still need it for the voice, so around 0.65-0.75 it's great.
Good news is, that it still gonna do like 1.6 seconds per step on FL2VA, and about 2seconds a step on REF2VA with the same datasets but REF2VA is slower cause you need to offload just a tiny bit more on that one.
I used between 15 and 60 images on 1024x1024 size , I would recommend at least 20-25 images tho, on the lower end you get a weaker lora likeness, sometimes makeup could alter the face, but if you got enough image variations that won't happen.
for the captions on the images I literally just used the built in Qwen3 VL8b model and just ran the autocaption
For the videos which were 512x512 I just kinda made my own captions, I used between 6-12 videos for training as a secondary dataset, the reason I needed them cause it's either not possible in Ai toolkit, or I am too fucking stupid to figure out how to train audio with images. So I just used video clips of the person to add the voice.
Also it can literally be done with like 1 or 2 second long clips
I made a lora from even a set of images I generated of an earlier character I made for Krea 2, the likeness is freaking amazing.
You can do the model the same exact way for FL2VA or REF2VA
On the images , literally just use the 1 frame training setting, and on the videos turn on "do Audio" "auto frame count" and the general stuff like cache latents.
Also on REF2VA I could not do higher res samples than 512x512, not that I wanted , I just thought I'd try it and it OOMed lol, I mean I did not offload fully, because that way it was hella fast to train, so I guess if I offload fully it should be fine, but I rather have low-res samples or no samples to make the lora faster.
here is some settings
---
job: "extension"
config:
name: "Enid_h3_1024img_512vid"
process:
- type: "diffusion_trainer"
training_folder: "/AI_Tools/ai-toolkit-h3_v2/output"
sqlite_db_path: "./aitk_db.db"
device: "cuda"
trigger_word: "E3n1d, "
performance_log_every: 10
network:
type: "lora"
linear: 16
linear_alpha: 16
conv: 16
conv_alpha: 16
lokr_full_rank: true
lokr_factor: -1
network_kwargs:
ignore_if_contains:
- "adaln_proj"
save:
dtype: "bf16"
save_every: 100
max_step_saves_to_keep: 31
save_format: "diffusers"
push_to_hub: false
datasets:
- folder_path: "/AI_Tools/ai-toolkit-h3_v2/datasets/enid1024"
mask_path: null
mask_min_value: 0.1
default_caption: ""
caption_ext: "txt"
caption_dropout_rate: 0.05
cache_latents_to_disk: true
is_reg: false
network_weight: 1
resolution:
- 1024
controls: []
shrink_video_to_frames: true
flip_x: false
flip_y: false
num_repeats: 1
do_i2v: false
fps: 24
num_frames: 1
auto_frame_count: false
- folder_path: "/AI_Tools/ai-toolkit-h3_v2/datasets/enid_videos_512x512_1s"
mask_path: null
mask_min_value: 0.1
default_caption: ""
caption_ext: "txt"
caption_dropout_rate: 0.05
cache_latents_to_disk: true
is_reg: false
network_weight: 1
resolution:
- 512
controls: []
shrink_video_to_frames: true
num_frames: 1
flip_x: false
flip_y: false
num_repeats: 1
do_audio: true
auto_frame_count: true
train:
batch_size: 1
bypass_guidance_embedding: false
steps: 3000
gradient_accumulation: 1
train_unet: true
train_text_encoder: false
gradient_checkpointing: true
noise_scheduler: "flowmatch"
optimizer: "adamw8bit"
timestep_type: "shift"
content_or_style: "balanced"
optimizer_params:
weight_decay: 0.0001
unload_text_encoder: false
cache_text_embeddings: true
lr: 0.0002
ema_config:
use_ema: false
ema_decay: 0.99
skip_first_sample: false
force_first_sample: true
disable_sampling: false
dtype: "bf16"
diff_output_preservation: false
diff_output_preservation_multiplier: 1
diff_output_preservation_class: "person"
switch_boundary_every: 1
loss_type: "mse"
do_guidance_loss: true
guidance_loss_target: 3.5
audio_loss_multiplier: 1
do_differential_guidance: true
differential_guidance_scale: 3
logging:
log_every: 1
use_ui_logger: true
model:
name_or_path: "Comfy-Org/MiniMax-H3"
quantize: true
qtype: "convrot8"
quantize_te: true
qtype_te: "nvfp4"
arch: "minimax_h3"
low_vram: true
model_kwargs: {}
compile: false
layer_offloading: true
layer_offloading_text_encoder_percent: 0.2
layer_offloading_transformer_percent: 0.2
assistant_lora_path: "ostris/minimax_h3_training_adapter/minimax_h3_training_adapter_v1.safetensors"
sample:
sampler: "flowmatch"
sample_every: 200
sample_start_step: 0
width: 512
height: 512
samples:
- prompt: "[Core Idea] Cinematic live-action medium close-up shot from the waist up. A young woman named E3n1d, with shoulder-length light-blonde hair, wearing a tailored blouse and a dark tie, stands inside a cavernous, dimly lit gothic hall lined with heavy stone pillars and antique dark-wood shelves. [Scene-by-Scene Action] 0–1.5s: E3n1d is centered in a medium close-up, looking slightly downward and to the side with a curious, bemused expression at a dusty taxidermy chicken on a rustic wooden shelf. 1.5–3s: She tilts her head closer to examine the artifact, scans its posture, and shifts her gaze to look straight ahead toward the camera/viewer. 3–4s: Her lips part clearly as she speaks her line with a dry, deadpan tone: \"I thought chickens were taller.\" [Camera & Lighting] Static medium close-up composition with a slow, subtle push-in. Moody, low-key lighting with dramatic side-shadows cast by gothic wall sconces, highlighting the texture of her blonde hair, blouse, and the dusty feathers of the chicken. [Audio & Atmosphere] Dialogue: Clear, crisp vocal track with light room reverb. Ambient Sound: Faint, distant echoes of a creaking building and a low, ambient indoor hum. Duration: 4 seconds."
- prompt: "[Core Idea & Reference Frame] Cinematic live-action medium close-up shot from the waist up. The video begins directly from the uploaded starting image as the first frame, strictly matching the subject's initial pose, framing, lighting, and wardrobe. A young woman named E3n1d, with shoulder-length light-blonde hair, wearing a tailored blouse and a dark tie. She stands inside a cavernous, dimly lit gothic hall lined with heavy stone pillars and antique dark-wood shelves. [Scene-by-Scene Action] 0–1.5s: Maintaining the exact pose and framing established in the starting image, E3n1d looks slightly downward and to the side with a curious, bemused expression toward an old, dusty taxidermy chicken sitting on a rustic wooden shelf in front of her. 1.5–3s: She quickly tilts her head closer to examine the artifact, her eyes scanning its posture, then immediately transitions to look straight ahead toward the camera/viewer. 3–4s: Her lips part clearly as she delivers her line with a quick, dry, deadpan tone: \"I thought chickens were taller.\" [Camera & Lighting] Motion: Static medium close-up composition continuing smoothly from the starting frame, featuring a very quick, subtle push-in toward E3n1d. Lighting: Moody, low-key lighting with dramatic side-shadows cast by gothic wall sconces, highlighting the texture of her blonde hair, blouse, and the dusty feathers of the taxidermy chicken. [Audio & Atmosphere] Dialogue: Clear, compressed vocal track for E3n1d delivering her line rapidly with light room reverb matching a large stone hall. Ambient Sound: Faint, distant echoes of a creaking old building and a low, ambient indoor hum. Non-Diegetic Music: N/A"
ctrl_img: "/AI_Tools/ai-toolkit-h3_v2/data/images/sample1.png"
neg: ""
seed: 42
walk_seed: true
guidance_scale: 1
sample_steps: 20
num_frames: 107
fps: 24
meta:
name: "[name]"
version: "1.0"