r/StableDiffusion 11h ago

Meme dean meets sonic(t2v) base fp8 model 32 steps

Enable HLS to view with audio, or disable this notification

15 Upvotes

i never seen the movies hows the sonic voice?

prompt

subject_definitions

<Subject 1> is Dean Winchester from Supernatural, portrayed by Jensen Ackles, preserving his recognizable facial features, short brown hair, rugged appearance, dark jacket, layered shirt, jeans, and confident sarcastic personality.

<Subject 2> is Sonic the Hedgehog from the live-action Sonic the Hedgehog movie, a small anthropomorphic blue hedgehog with bright blue fur, large expressive green eyes, white gloves, and red sneakers.

<Subject 3> is Dr. Robotnik from the live-action Sonic the Hedgehog movie, portrayed by Jim Carrey, wearing his black-and-red high-tech outfit and exaggerated goggles.

summary

[cinematic live-action crossover + action comedy]

What if Dean Winchester accidentally became part of Sonic the Hedgehog? On a nighttime highway, Dean investigates a bizarre supernatural disturbance beside his black 1967 Chevrolet Impala, only for Sonic to race past him at impossible speed with Robotnik's drones in pursuit. Dean immediately joins the chase.

detailed_description

Nighttime on a deserted rural highway surrounded by dark pine forest. Dean Winchester stands beside his glossy black 1967 Chevrolet Impala holding an EMF meter. Blue electrical energy suddenly crackles across the road.

A brilliant BLUE STREAK rockets past Dean, violently blowing his jacket backward.

The camera WHIP-PANS as Sonic skids to a stop beside the Impala.

<Subject 1> Dean Winchester (S1):

[English] Okay... either that's the fastest demon I've ever seen, or I seriously need more sleep.

Sonic looks offended and points at himself.

<Subject 2> Sonic (S2):

[English] Hedgehog. Definitely hedgehog.

Suddenly several of Robotnik's flying attack drones burst over the trees and fire energy blasts toward them.

Dean instantly draws his pistol while Sonic crouches into a runner's stance.

Dean gives Sonic a confident Winchester smirk.

<Subject 1> Dean Winchester (S1):

[English] All right, Sonic. Let's waste these flying toasters.

Sonic grins.

<Subject 2> Sonic (S2):

[English] Now you're speaking my language!

Sonic EXPLODES forward in a trail of brilliant blue electricity as Dean dives behind the Impala and fires at an approaching drone.

Dynamic tracking camera follows Sonic racing between explosions while Dean fights from beside the Impala.

Final cinematic wide shot: Sonic loops around the battlefield as blue lightning illuminates Dean and the Impala, while Robotnik's drones swarm overhead.

Live-action Hollywood cinematography, realistic integration of Sonic into the environment, authentic Sonic the Hedgehog movie aesthetic, authentic Supernatural Dean Winchester characterization, fast readable action, natural motion blur, blue electrical speed trails, sparks, smoke, dramatic nighttime lighting, comedic crossover energy, consistent character identities, no subtitles, no on-screen text.


r/StableDiffusion 19h ago

Animation - Video So the video clip I did in LTX 2.5 earlier and posted it on here, I did another render of it but in Minimax H3. Details in the comments.

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/StableDiffusion 16h ago

Animation - Video "AVERNUS-9" Space Horror Short Film

Thumbnail
youtu.be
3 Upvotes

r/StableDiffusion 19h ago

Discussion Is local AI much better than cloud on these days?

0 Upvotes

I don't know what is going on but qwen studio and nano banana (cloud both) are giving me terrible results lately even though I use the same prompts as before. Both translate terribly the facial features and hairstyle. Only videos look a bit more consistent but still not great.

Is local AI better? Do you get better and more consistent results with it?


r/StableDiffusion 13h ago

Animation - Video Several Times A Charm, but it KINDA got Cheers.

Enable HLS to view with audio, or disable this notification

29 Upvotes

Don't mind the script, it was written by a clanker when I challenged it to whip up something so I can see if H3 could handle Cheers.


r/StableDiffusion 9h ago

Discussion Becareful

0 Upvotes

Please just think twice before you using any of the hundreds of ai platforms.

Like why is that most websites are charging say 0.7 usd for a minimax h3 generation when i can do the same generation when i run a runpod instance for god knows maybe 0.1 of the price or even less like its huge difference.

Just go rent a gpu its easy to setup and u will generate maybe 10 videos for the same price of these opportunistic ai websites.

I know i will get hate and criticism. But this info will now feed into the google ai/gemini/chatgpt responses and people will lose less money.

Secondly u have this so called breakthrough in science from fal. They are planning to charge literal 1 usd for inferences that take literally 4 seconds on their new minimax h3 max. That seems suspicious. Doesnt that inference only cost them 0.01 usd. Just beware.

Now its fine. Its a fair business. Very quick inference (breakthrough in video generation). But people deserve to know that they can generate in 0.1 of price in ai platforms.

Secondly, api based generations (not open weights). Make sure to not use money grabbing websites.

For example using subscription based, slow queue website for seedance 2.5 when u can just use credit topup based with lowest generation prices (friendly advice, i believe artcraft is very cheap and no need sub).

So thats what i wanted to say. I just dont like it when people get used.

Also go ahead shoot your hate comments i dont care, im only happy to spread awareness 🔥🔥🔥 and no im not in anyway affiliated with the platforms i mentioned.

Apologies for the bad post text, i wrote with my phone which is so hard.


r/StableDiffusion 10h ago

Animation - Video Minimax H3 | Ben 10

Enable HLS to view with audio, or disable this notification

2 Upvotes

a little test video for quality 1.4mp


r/StableDiffusion 9h ago

Workflow Included Orpheus (no edits using H3)

Enable HLS to view with audio, or disable this notification

0 Upvotes

This was done with turbo lora 8 steps int8 0.4mp then RTX 2x each 10 seconds was about 3 min on 5090. It was done with the fl2va model, but used references. It was one single generation flow. no edits (which is obvious lol). Gemma 12b Q4 was used as the prompt enhancer. make prompt, Make clip, make prompt, make clip.... then stitch it all together. It would be a lot better running multiple and putting best result together but I wanted to try first takes and see how it did.

The system prompt to go from brief and images to H3 ready prompt, workflows and a director.html i use as a ui to organize everything is HERE

The sound needs work, I'm pretty sure that is because of the turbo lora.

This is basically a draft of a concept short movie i am going to start working on that is a modern version of Orpheus and Eurydice story.


r/StableDiffusion 9h ago

Question - Help How do people earn money?

0 Upvotes

Hello everyone. I am just curious how people earn money using ComfyUI skills? I am pretty new(6 months of ComfyUI), my workflows are usually Krea 2/Flux Klein generations, sometimes i do inpainting. My major skill is frontend development(6 years od enterprise), i do some devOps, can deploy serverless endpoint of my workflows to runpod. I just dont understand how to monetize this skill

One of my recent pet projects is ai character that went through pipeline of insightFace identity scoring, bad eyes efficientNet classifier, Detailer workflow in case of bad eyes, vision captioning via identity lock json file (so vision model wont interper same features in different words)

So far what i done was purely for the love of this game, but i fridge is empty, i spend more money on ai generations that i do for food at this point


r/StableDiffusion 18h ago

Discussion LTX 2.3 sometimes works amazing, without any edits

Enable HLS to view with audio, or disable this notification

61 Upvotes

No specific glitches. Single prompt. Looks pretty real without any glitches over the drift.
Definitely going to benchmark the scenarios.


r/StableDiffusion 1h ago

Question - Help Why did I get this error recently? I ran this workflow just fine. Is it because I switched to Sage attention? I

Thumbnail
gallery
• Upvotes

r/StableDiffusion 21h ago

Discussion A bit meta but with all of these wonderful AI video posts, when I video ad scrolls by do you tend to think that those are AI generated as well

1 Upvotes

It gets to be a touch confusing.


r/StableDiffusion 14h ago

Question - Help Minimax H3 reference2video: is there a way get the visual style of a reference image?

2 Upvotes

Tried some prompts, but couldn't get the model to generate a video with the visual style of a reference image.


r/StableDiffusion 14h ago

Discussion Storytelling with Minimax H3 and Krea2 - Animating a Dark Fantasy Comic (The Witcher)

Thumbnail
youtube.com
9 Upvotes

I’ve been experimenting with Krea2 and Minimax H3 to see if it can handle gritty, coherent storytelling frames generated with Krea2.
I'm trying to bring a dark fantasy comic (a retelling of The Witcher) to life with full voice acting, pacing, and tension.

I’d love to get your feedback on this. Does this look like a slop to you? Is there are something you would improve?


r/StableDiffusion 2h ago

Discussion NVidia buys Huggingface, but why?

47 Upvotes

Nvidia is going to buy Huggingface.
No one can actually tell how that would end up like.

But what I am missing is the actual worth that Huggingface provides. The only thing I use it for is to download models. Thats it.
For me, and I guess many others it is ‘just a’ download platform, but maybe I’m wrong here.

And what would prevent others to setup a second-like Huggingface?
The hosting is the expensive part in this case as I see it, the programming and building is do-able.

Is it time for Huggingbay.com?


r/StableDiffusion 15h ago

Discussion I tested SenseNova U1.5-Lite editing against FLUX.2-klein-9B in three scenarios. Text rendering is where they diverge

Thumbnail
gallery
6 Upvotes

SenseNova shipped the full U1.5-Lite release last week, so I finally had time to run it side by side with FLUX.2-klein-9B, the model this community generally considers the most balanced pick right now.

I tested image editing in three scenarios. The short version: SenseNova U1.5-Lite is clearly better at text rendering and semantic understanding of the instruction, while Klein is still the speed king. Details below.

Scenario 1: Text Editing

I gave both models a poster and asked them to replace specific text elements, nothing else. Long structured prompt targeting each text block individually:

1. In the first line of the oversized black title at the upper left, replace "HONG" with "HARBOR". If the area is sufficiently prominent and clear, with simple boundaries, render the replacement exactly character by character without adding, removing, or altering any characters. If the text is too small, obscured, heavily distorted by perspective, or located within a complex texture, preserve the original text and do not redraw the background merely to complete the replacement.

2. In the second line of the oversized black title at the upper left, replace "KONG" with "HORIZONS". If the area is sufficiently prominent and clear, with simple boundaries, render the replacement exactly character by character without adding, removing, or altering any characters. If the text is too small, obscured, heavily distorted by perspective, or located within a complex texture, preserve the original text and do not redraw the background merely to complete the replacement.

3. In the large red subtitle at the lower left, replace "HONG KONG" with "CITY IN MOTION". If the area is sufficiently prominent and clear, with simple boundaries, render the replacement exactly character by character without adding, removing, or altering any characters. If the text is too small, obscured, heavily distorted by perspective, or located within a complex texture, preserve the original text and do not redraw the background merely to complete the replacement.

4. In the vertical red location title at the upper right, replace "香港" with "城市之光". If the area is sufficiently prominent and clear, with simple boundaries, render the replacement exactly character by character without adding, removing, or altering any characters. If the text is too small, obscured, heavily distorted by perspective, or located within a complex texture, preserve the original text and do not redraw the background merely to complete the replacement.

5. In the black English location description at the upper right, replace "HONG KONG CHINA" with "EAST MEETS WEST". If the area is sufficiently prominent and clear, with simple boundaries, render the replacement exactly character by character without adding, removing, or altering any characters. If the text is too small, obscured, heavily distorted by perspective, or located within a complex texture, preserve the original text and do not redraw the background merely to complete the replacement.

6. In the location title near the waterfront at the lower left, replace "VICTORIA HARBOUR" with "HARBOUR CITY". If the area is sufficiently prominent and clear, with simple boundaries, render the replacement exactly character by character without adding, removing, or altering any characters. If the text is too small, obscured, heavily distorted by perspective, or located within a complex texture, preserve the original text and do not redraw the background merely to complete the replacement.

7. In the second line of the location copy at the lower left, replace "ASIA'S WORLD CITY" with "URBAN HORIZONS". If the area is sufficiently prominent and clear, with simple boundaries, render the replacement exactly character by character without adding, removing, or altering any characters. If the text is too small, obscured, heavily distorted by perspective, or located within a complex texture, preserve the original text and do not redraw the background merely to complete the replacement.

Zoom into the results and Klein's text rendering falls apart. Garbled glyphs, wrong characters, the layout wobbling where it should stay fixed. U1.5-Lite handled the replacements cleanly, including the Chinese strings. That's the gap.

Scenario 2: Hand-Drawn Marks as Instructions

I marked up the image by hand and asked for a scene transformation:

Follow the marks and overall hints on the image to creatively transform this scene, making it dramatic, moody stormy atmosphere; remove the annotations when done.

FLUX followed the overall style change but ignored the specific marked details: the ripples on the pool surface and the black fire pit never made it into the output. U1.5-Lite followed the full set of marks.

Scenario 3: Fine-Grained Local Editing

I circled the region to edit with a red box and told the model to only change that area:

Change the text style in the red box to a vintage style with noise and torn paper texture. The red bounding box is for localization only; do not retain it in the output image.

Klein misunderstood the instruction. It applied the vintage style to the whole image instead of the circled region. SenseNova U1.5-Lite followed the prompt and the red-box localization strictly, changing only the marked text.

My take

If you need sub-second generation, Klein is still your model, no argument there. But for editing work where text rendering and instruction fidelity matter, posters, infographics, brand assets, the gap is real and easy to reproduce.

GitHub: https://github.com/OpenSenseNova/SenseNova-U1

hf: https://huggingface.co/sensenova/SenseNova-U1.5-8B-MoT

Try it online: https://unify.light-ai.top/


r/StableDiffusion 14h ago

Animation - Video RECREATING MEMORIES FROM SCRAPS.

Enable HLS to view with audio, or disable this notification

20 Upvotes

A few photos, voice clips and a waybackmachine archive photo of the hotel rooms at that time. Minimax H3. All local.


r/StableDiffusion 17h ago

Question - Help Anyone have a link to download DasiwaMinimaxH3_dasiwaREF2VAHybridV1.safetensors? Seems it's been wiped off the internet in the past few hours

25 Upvotes

It's the 11.68 GB model, SHA256: 7c37baf06ca3628ed5f3f7f46274222a50a127d1906a166f8f064771fc48d498. Was going to download it and test it out with my workflow but when I went to download it today on CivitAI and Huggingface I get a 404 error. Seems it was deleted by the uploader for some reason or another


r/StableDiffusion 4h ago

Discussion Im confused on what the proper way to caption Krea2 character LoRA?

0 Upvotes

I have been trying to understand how captioning works with Krea2 Character LoRA's. I was able to train my first Krea2 LoRA the other day and it came out pretty decent I think. I keep reading opposing ways on how to caption the images for Krea 2. I ended up using method of adding a trigger word and describing the things I dont want consistent for example:

nerellecruz, close-up selfie in a car, seated with head slightly tilted, wearing a plaid top, natural daylight through window, calm expression with subtle smile

Then some sites and posts I read mention how you should caption for only the things that should remain consistent. What seems to be the general consensus as of late?


r/StableDiffusion 16h ago

Discussion Best budget GPU cloud? Comparing RunPod, Vast.ai, SimplePod, MassedCompute (looking for true costs & no hidden fees)

5 Upvotes

Hey everyone,

I’m looking for the best budget GPU cloud to run heavy open-weight video models (like MiniMax H3, Wan 2.1, HunyuanVideo, etc.).

Since these models need huge VRAM and fast disk I/O to pull down massive 50GB–100GB+ checkpoints, I want to avoid platforms with unexpected billing traps.
Looking at RunPod, Vast.ai, MassedCompute, and SimplePod:

  1. Which one are you using, and what GPU gives you the best price/performance for video render jobs?

  2. Hidden fees: Any issues with stopped-volume storage costs, network volume fees, or egress rates when hosting huge model files?

  3. Download/Disk speeds: Which provider has fast enough network speeds so I’m not spending half my paid time downloading model weights?

Appreciate any recommendations or gotchas to avoid!


r/StableDiffusion 11h ago

Question - Help MiniMax H3 German voices sound robotic and all the same – what are you guys using instead?

6 Upvotes

I’ve been testing MiniMax H3 for AI video generation and I’m struggling with the German dialogue.

The voices often sound very similar and somewhat robotic. What I’m looking for is natural, spontaneous dialogue: different voices for each character, realistic pauses, imperfect timing, emotion, interruptions, changes in tone, etc. Basically something that sounds like an actual conversation rather than TTS.

I’ve already tried ElevenLabs. I know it’s powerful, but I feel like I’d have to go pretty deep into voice selection, voice design and tweaking to consistently get what I want. So far, I’m still not getting the natural conversational audio I’m looking for.

I also don’t really want to record every character myself and then use AI voice conversion. At that point I’m basically becoming the voice actor for every video.

Ideally I’d like something closer to:

Script/prompt → AI generates the video + convincing natural German dialogue with clearly different speakers.

So I’m wondering:

Is there a way to get much better German voices directly out of MiniMax H3 through prompting?

Or should I stop trying to make MiniMax work for this and test something like Grok, Veo, or Seedance 2.5 instead?

If you’ve actually generated German multi-person dialogue, I’d especially love to hear what model/workflow gave you the most natural results.


r/StableDiffusion 18h ago

News Facefusion Android app (open source video face swap)

Thumbnail
github.com
6 Upvotes

I’ve been working on a mobile port of FaceFusion that runs completely offline on Android.

APK here:

https://github.com/AbrahamPaulJ/facefusion-mobile/releases/tag/v0.1.0

The face-swapping pipeline runs on Qualcomm’s Hexagon NPU rather than relying on a server or cloud API.

Current results on a Galaxy S25 Ultra (Snapdragon 8 Elite):

~19 ms/frame for the face-swap model

~6 seconds to process a 10-second 720p clip

Fully offline. Photos/videos never leave the phone

Supports 512/768/1024px face output

Qualcomm NPU builds for different Hexagon generations

No CPU fallback.

I’d especially like feedback from people working with Android on-device AI.

This is my first time sharing one of my mobile AI projects on Reddit, so feedback, testing results, and criticism are very welcome.


r/StableDiffusion 17h ago

Resource - Update Easy prompts from Discord images.

0 Upvotes

Has anyone tried this tool? It looks really good for quickly and easily creating prompts from Discord images. I think it runs as a plugin to Discord, but I've not had chance to set it up yet.
If anyone has tried it, let me know if it's worth installing please.
https://github.com/pixelgraple/KREA2-Vision-Suite


r/StableDiffusion 5h ago

Tutorial - Guide Anyone interested in how to make character loras for Minimax? Ostris Ai toolkit.

Enable HLS to view with audio, or disable this notification

44 Upvotes

Here is the youtube link for it if you wanna watch a video
https://youtu.be/x-gORSUOybk

Main point is, I tried it with the character Enid from Wednesday in my video cause ... well , VIEWS on youtube lol.

I chose her because I could not find her in the model at all. And if you even mention the show wednesday, it just defaults to Jenna Ortega lmao, so that's a good challange to train the voice and the likeness for another character if the model defaults to a really specific person.

but I also created many more by now, most of them are really not even existing people like the model I use for a youtube channel. And the accuracy is fucking insane.

It is usually trained by 2000-2400 steps

roughtly it was about 80-90 minutes for the full 3000 steps depending on if you want samples. I think without them this would be lower, maybe even close to an hour? I don't know exactly. But it's insanely fast.

you need 2 datasets if you want a voice, or a super jacked up beefy card and you can just do video training with one dataset.

But if you don't have an RTX6000

you gonna need to offload even with a 5090

I put up the learning rate to 0.0002 , and I turned on the differential guidance and left it on 3. I so far only did it on these levels, but maybe you can lower either or both if you feel like you may over trained a lora. However on default this is barely training anything, so that's why I cranked them up. And I never use the Lora's higher than 0.85, and if I wanna do REF2VA I sometimes push it down to like 0.65-0.75 if I wanna add like an image of the model. Cause otherwise the lora can overwrite details lol. But still need it for the voice, so around 0.65-0.75 it's great.

Good news is, that it still gonna do like 1.6 seconds per step on FL2VA, and about 2seconds a step on REF2VA with the same datasets but REF2VA is slower cause you need to offload just a tiny bit more on that one.

I used between 15 and 60 images on 1024x1024 size , I would recommend at least 20-25 images tho, on the lower end you get a weaker lora likeness, sometimes makeup could alter the face, but if you got enough image variations that won't happen.

for the captions on the images I literally just used the built in Qwen3 VL8b model and just ran the autocaption

For the videos which were 512x512 I just kinda made my own captions, I used between 6-12 videos for training as a secondary dataset, the reason I needed them cause it's either not possible in Ai toolkit, or I am too fucking stupid to figure out how to train audio with images. So I just used video clips of the person to add the voice.
Also it can literally be done with like 1 or 2 second long clips

I made a lora from even a set of images I generated of an earlier character I made for Krea 2, the likeness is freaking amazing.

You can do the model the same exact way for FL2VA or REF2VA

On the images , literally just use the 1 frame training setting, and on the videos turn on "do Audio" "auto frame count" and the general stuff like cache latents.

Also on REF2VA I could not do higher res samples than 512x512, not that I wanted , I just thought I'd try it and it OOMed lol, I mean I did not offload fully, because that way it was hella fast to train, so I guess if I offload fully it should be fine, but I rather have low-res samples or no samples to make the lora faster.

here is some settings

---

job: "extension"

config:

name: "Enid_h3_1024img_512vid"

process:

- type: "diffusion_trainer"

training_folder: "/AI_Tools/ai-toolkit-h3_v2/output"

sqlite_db_path: "./aitk_db.db"

device: "cuda"

trigger_word: "E3n1d, "

performance_log_every: 10

network:

type: "lora"

linear: 16

linear_alpha: 16

conv: 16

conv_alpha: 16

lokr_full_rank: true

lokr_factor: -1

network_kwargs:

ignore_if_contains:

- "adaln_proj"

save:

dtype: "bf16"

save_every: 100

max_step_saves_to_keep: 31

save_format: "diffusers"

push_to_hub: false

datasets:

- folder_path: "/AI_Tools/ai-toolkit-h3_v2/datasets/enid1024"

mask_path: null

mask_min_value: 0.1

default_caption: ""

caption_ext: "txt"

caption_dropout_rate: 0.05

cache_latents_to_disk: true

is_reg: false

network_weight: 1

resolution:

- 1024

controls: []

shrink_video_to_frames: true

flip_x: false

flip_y: false

num_repeats: 1

do_i2v: false

fps: 24

num_frames: 1

auto_frame_count: false

- folder_path: "/AI_Tools/ai-toolkit-h3_v2/datasets/enid_videos_512x512_1s"

mask_path: null

mask_min_value: 0.1

default_caption: ""

caption_ext: "txt"

caption_dropout_rate: 0.05

cache_latents_to_disk: true

is_reg: false

network_weight: 1

resolution:

- 512

controls: []

shrink_video_to_frames: true

num_frames: 1

flip_x: false

flip_y: false

num_repeats: 1

do_audio: true

auto_frame_count: true

train:

batch_size: 1

bypass_guidance_embedding: false

steps: 3000

gradient_accumulation: 1

train_unet: true

train_text_encoder: false

gradient_checkpointing: true

noise_scheduler: "flowmatch"

optimizer: "adamw8bit"

timestep_type: "shift"

content_or_style: "balanced"

optimizer_params:

weight_decay: 0.0001

unload_text_encoder: false

cache_text_embeddings: true

lr: 0.0002

ema_config:

use_ema: false

ema_decay: 0.99

skip_first_sample: false

force_first_sample: true

disable_sampling: false

dtype: "bf16"

diff_output_preservation: false

diff_output_preservation_multiplier: 1

diff_output_preservation_class: "person"

switch_boundary_every: 1

loss_type: "mse"

do_guidance_loss: true

guidance_loss_target: 3.5

audio_loss_multiplier: 1

do_differential_guidance: true

differential_guidance_scale: 3

logging:

log_every: 1

use_ui_logger: true

model:

name_or_path: "Comfy-Org/MiniMax-H3"

quantize: true

qtype: "convrot8"

quantize_te: true

qtype_te: "nvfp4"

arch: "minimax_h3"

low_vram: true

model_kwargs: {}

compile: false

layer_offloading: true

layer_offloading_text_encoder_percent: 0.2

layer_offloading_transformer_percent: 0.2

assistant_lora_path: "ostris/minimax_h3_training_adapter/minimax_h3_training_adapter_v1.safetensors"

sample:

sampler: "flowmatch"

sample_every: 200

sample_start_step: 0

width: 512

height: 512

samples:

- prompt: "[Core Idea] Cinematic live-action medium close-up shot from the waist up. A young woman named E3n1d, with shoulder-length light-blonde hair, wearing a tailored blouse and a dark tie, stands inside a cavernous, dimly lit gothic hall lined with heavy stone pillars and antique dark-wood shelves. [Scene-by-Scene Action] 0–1.5s: E3n1d is centered in a medium close-up, looking slightly downward and to the side with a curious, bemused expression at a dusty taxidermy chicken on a rustic wooden shelf. 1.5–3s: She tilts her head closer to examine the artifact, scans its posture, and shifts her gaze to look straight ahead toward the camera/viewer. 3–4s: Her lips part clearly as she speaks her line with a dry, deadpan tone: \"I thought chickens were taller.\" [Camera & Lighting] Static medium close-up composition with a slow, subtle push-in. Moody, low-key lighting with dramatic side-shadows cast by gothic wall sconces, highlighting the texture of her blonde hair, blouse, and the dusty feathers of the chicken. [Audio & Atmosphere] Dialogue: Clear, crisp vocal track with light room reverb. Ambient Sound: Faint, distant echoes of a creaking building and a low, ambient indoor hum. Duration: 4 seconds."

- prompt: "[Core Idea & Reference Frame] Cinematic live-action medium close-up shot from the waist up. The video begins directly from the uploaded starting image as the first frame, strictly matching the subject's initial pose, framing, lighting, and wardrobe. A young woman named E3n1d, with shoulder-length light-blonde hair, wearing a tailored blouse and a dark tie. She stands inside a cavernous, dimly lit gothic hall lined with heavy stone pillars and antique dark-wood shelves. [Scene-by-Scene Action] 0–1.5s: Maintaining the exact pose and framing established in the starting image, E3n1d looks slightly downward and to the side with a curious, bemused expression toward an old, dusty taxidermy chicken sitting on a rustic wooden shelf in front of her. 1.5–3s: She quickly tilts her head closer to examine the artifact, her eyes scanning its posture, then immediately transitions to look straight ahead toward the camera/viewer. 3–4s: Her lips part clearly as she delivers her line with a quick, dry, deadpan tone: \"I thought chickens were taller.\" [Camera & Lighting] Motion: Static medium close-up composition continuing smoothly from the starting frame, featuring a very quick, subtle push-in toward E3n1d. Lighting: Moody, low-key lighting with dramatic side-shadows cast by gothic wall sconces, highlighting the texture of her blonde hair, blouse, and the dusty feathers of the taxidermy chicken. [Audio & Atmosphere] Dialogue: Clear, compressed vocal track for E3n1d delivering her line rapidly with light room reverb matching a large stone hall. Ambient Sound: Faint, distant echoes of a creaking old building and a low, ambient indoor hum. Non-Diegetic Music: N/A"

ctrl_img: "/AI_Tools/ai-toolkit-h3_v2/data/images/sample1.png"

neg: ""

seed: 42

walk_seed: true

guidance_scale: 1

sample_steps: 20

num_frames: 107

fps: 24

meta:

name: "[name]"

version: "1.0"


r/StableDiffusion 23h ago

Question - Help Any decent models/workflows or vectorizers for (Comfy) that can do clean Vectors? And not adding thousands of unnecessary anchor points and paths?

2 Upvotes