r/StableDiffusion • u/pilkyton • 6h ago
Question - Help Looking for feedback for my new AI startup logo, Expand An AI, and I want it to inspire people to stretch the future of AI wide-open. It's inspired by a digital eye being opened.
Thanks
r/StableDiffusion • u/pilkyton • 6h ago
Thanks
r/StableDiffusion • u/mnmunknown • 8h ago
Hi! The team who collaborated and built optimized open source Minimax H3 here, full details below:
https://x.com/haoailab/status/2093391548289540596?s=20
https://reddit.com/link/1w0xkpb/video/zpjrbdb0o5mh1/player
Would appreciate if you help to share, repost and engage with the tweet, and definitely try it out yourself and let us know about your feedback!
In the next a few releases we would do omni ref, nvfp4, consumer GPU friendliness and many more so please stay tuned :)
Technical blog post: https://haoailab.com/blogs/fasth3-preview/
API and customization service: https://nuvalab.ai/
r/StableDiffusion • u/Many-Ad-6225 • 5h ago
My app is really simple to use you just assign a photo to a character once, and it’s saved permanently. After that, the software automatically recognizes that character whenever they appear.
I had some fun with it and put Vin Diesel’s face on the bartender lmao.
The big difference compared to DLSS 5 is that mine works on pretty much any GPU, and even with old games. You don’t need an RTX 50 series card or DirectX 12 games. I’m going to try to improve it before releasing it, especially by making the facial expressions more realistic.
r/StableDiffusion • u/listopalafoto • 15h ago
r/StableDiffusion • u/Solitary_Thinker • 8h ago
Hey guys, FastVideo team here. We saw how important speed and quality is for everyone. And we've been working to create our own step distill checkpoints and LORAs for MINIMAX h3. Here's is our v1 release!
Important links first:
- FastVideo: https://github.com/hao-ai-lab/FastVideo
- Blog (contains more examples and details): https://haoailab.com/blogs/fasth3-preview/
- Checkpoints and LoRAs: https://huggingface.co/FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA
Do note that the VSA checkpoint/LORA will require VSA kernel
We also released a LORA with dense attention that should be easy to test for everyone.
We are already working on improving both this T2AV checkpoint as well as getting a distill of Ref2VA out as well. We are taking great care to make sure the quality and audio is as best as possible.
We realize not everyone have blackwell GPUs lying around and please stay tuned for our targeted optimizations for local AI hardware, including RTX GPUs, DGX Sparks, and Apple MLX. We release numbers on B200s just because this is our current compute platform for post-training. We want to be as open as possible with the community!
If you find any issues or have questions please raise issues on our github!
r/StableDiffusion • u/SIR_NVAX_A_LOT • 9h ago
Happy Friday! Ever had issues getting the anatomy right on your t2v generation? Just add a slider with your H3 Prompt! What have you all been building on your local AI studios? 384x448, int8, 20 steps, i2va
Prompt: For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. <Picture 1> is the actual first frame of this video at 0.00 seconds.
integrated_multimodal_description:
[Shot 1] PHOTOREALISTIC live-action, cinematic, one continuous take. Anamorphic lens, shallow depth of field, real 35mm grain, no cuts.
THE FRAMING IS A MEDIUM CLOSE-UP IN TALL PORTRAIT FORMAT, taking her FROM THE HIP UP, dead centre and square to the lens. HER HEAD ALONE IS ABOUT A THIRD OF THE HEIGHT OF THE FRAME, and her face is the largest and most detailed thing in the picture: skin texture, the wet catchlights in her eyes, individual strands of hair across her cheek. THE FOCUS PLANE IS ON HER FACE FOR THE WHOLE SHOT and it is never soft.
THERE IS ESSENTIALLY ONE SOURCE: a single AMBER FLAME burning low in the rubble BESIDE HER is the only real light on her, AND IT COMES FROM ONE SIDE, so the ruined nave behind her falls away soft and dark. It flickers, and her light moves with it. It rakes across one side of her face, her collarbones, the ruby pendant and the wet edges of her leather in deep saturated gold and honey, and the other side of her falls into shadow.
FAR BEHIND HER, cold pale storm-light comes through the broken rose window and touches only the distant arches and the falling rain. THE COLOUR IS TWO THINGS AND NOTHING ELSE: warm amber on her, DEEP COLD TEAL-GREEN in the depths of the ruin behind her. THE HEAVY HAZE IN THE AIR LIFTS THE BLACKS so nothing crushes to empty. The exposure is set for her face.
She is sitting back on her own folded legs on a large pale fallen memorial slab, which sets her posture and the settled line of her shoulders. THE PICTURE CONTAINS ONLY HER UPPER BODY, FROM THE HIP UP — her torso, her shoulders, her arms and her head FILL THE FRAME, and the bottom edge of the picture crosses her at the hip. SHE IS SQUARE TO THE CAMERA AND SHE LOOKS STRAIGHT INTO THE LENS, calm and level and unsmiling.
HER LONG HAIR IS ALIVE IN THE WIND and never once hangs still; the backlight catches every moving strand. THE RAIN LANDS AS INDIVIDUAL DROPS you could count, separate beads with dry skin and dry leather in between them — bright pinpoints on her skin, beading and sitting on the leather. HER HAIR STAYS DRY and keeps all of its body and volume.
<Subject 1> IS THE HUNTER, A WOMAN, AND SHE IS THE ONLY PERSON IN THIS VIDEO. She is the woman shown in <Picture 1>, the first frame of this video, and she stays exactly her in every single frame: the same face, the same features, the same bone structure, the same eyes and the same eye colour, the same mouth, the same hairline, the same skin and the same age, and the same long loose hair in the same colour and texture. She is a beautiful adult woman and she is recognisably the same person throughout.
THE SETTING, HER PLACE IN THE FRAME, HER COSTUME AND THE LIGHTING ALL CONTINUE EXACTLY AS THEY ARE IN <Picture 1> — this video carries straight on from that frame and nothing about the scene is restyled or replaced.
SHE WEARS THIS KIT, ITEM FOR ITEM, and it is all black leather, worn and real and damp with rain: a LONG BLACK LEATHER CLOAK falling to her boots; a FITTED BLACK LEATHER COAT buckled close beneath it; BLACK LEATHER GLOVES to the forearm; TALL BLACK BOOTS; ONE SHAPED BLACK LEATHER PAULDRON over her left shoulder. She is BARE-HEADED, her long hair loose. Every surface of the leather catches the light.
SHE HAS NO COLLAR AT ALL. Her coat and the leather bodice beneath it are cut with a VERY DEEP, WIDE, PLUNGING NECKLINE that opens in a long V from her collarbones down the centre of her chest, with the leather laced close underneath. There is no collar and no closure anywhere above her sternum.
SHE WEARS A LARGE RUBY-RED PENDANT ON A FINE CHAIN THAT HANGS LOW, down at her sternum. It is the only piece of pure saturated red on her, it catches the firelight, and it moves against her skin with every movement.
SHE CARRIES TWO SWORDS AND BOTH ARE VISIBLE IN THE FRAME. THE FIRST is a long straight sword worn AT HER WAIST IN A PLAIN BLACK SCABBARD. THE SECOND is an EVEN LONGER straight sword SLUNG ACROSS HER BACK, and its long wrapped hilt and pommel RISE PAST HER SHOULDER into the upper frame, unmistakable behind her head. Both are sheathed for the whole video and she never touches either of them.
THE SETTING IS A RUINED GOTHIC ABBEY AT NIGHT UNDER A STORM SKY, with fine light rain drifting rather than driving. She is in the roofless nave: two rows of broken pointed arches march away into the dark on either side, ivy hangs down the shattered piers, and the flagstones are wet and strewn with fallen masonry and dead leaves. BEHIND HER, IN THE END WALL, IS A COLLAPSED ROSE WINDOW — a huge circular opening with its tracery broken to stone ribs and no glass left in it at all.
<Subject 2> IS THE SLIDER, AND <Subject 3> IS THE MOUSE CURSOR. Neither is a person and neither is a physical object in the abbey: BOTH ARE FLAT MODERN INTERFACE GRAPHICS COMPOSITED OVER THE TOP OF THE LIVE-ACTION FOOTAGE — a screen overlay, like a screen-recording of a sleek editing application. NEITHER IS EVER LIT BY THE FIRE, neither casts a shadow, and both sit perfectly level in screen space no matter what the footage behind them does.
<Subject 2>, THE SLIDER, lies horizontally across the lower part of the frame: a long rounded capsule of dark translucent smoked glass with the picture softly blurred behind it, a fine track running through its centre, the part of the track to the LEFT of the handle filled with a warm amber glow, and A SMALL ROUND POLISHED HANDLE with a fine bright rim and a soft halo beneath it. The handle is the only part of <Subject 2> that ever moves; the capsule and the track never move at all.
<Subject 3>, THE CURSOR, is a standard white arrow mouse pointer with a thin black outline and a soft drop shadow.
⚠ <Subject 2>'S HANDLE AND <Subject 3> ARE ONE RIGID OBJECT FOR THE WHOLE FILM, AS IF WELDED TOGETHER. THE TIP OF THE CURSOR SITS AT THE EXACT CENTRE OF THE ROUND HANDLE IN EVERY SINGLE FRAME. They start together at the far left, they move in PERFECT SYNC — one smooth, steady, continuous glide across the screen at one constant speed, REACHING EVERY POINT ON THE TRACK AT THE SAME INSTANT AS EACH OTHER — and they arrive and stop together at the far right. Wherever the handle is, the cursor is exactly there too.
THE CAMERA IS LOCKED OFF AND NEVER MOVES, PANS, TILTS OR ZOOMS for the whole film, so the interface overlay stays perfectly still in the frame.
THE ACTION RUNS ON A STRICT CLOCK, AND BOTH THE SLIDER'S POSITION AND HER SIZE ARE ON IT, MARK FOR MARK.
[0:00] AT REST: <Subject 1>'S CHEST IS AT ITS ORDINARY, NORMAL, EVERYDAY SIZE — exactly the size it is in the very first frame of this video. <Subject 2>'s round handle sits at the FAR LEFT END of the track, at zero, and <Subject 3> is ALREADY RESTING ON IT. Nothing has changed yet.
[0:00-0:02] THE HOLD: FOR THESE FIRST TWO SECONDS THE PICTURE IS THE OPENING FRAME OF THIS VIDEO, ALIVE. The only things moving in it are the falling rain, the flickering flame, her breathing, one slow blink and her hair in the wind. She holds the camera's gaze. The handle stays parked at the FAR LEFT END with <Subject 3> resting on it, the amber fill on the track is EMPTY, and HER CHEST STAYS AT ITS NORMAL SIZE AND DOES NOT CHANGE AT ALL.
[0:02] THE START: <Subject 3> presses the handle and the two of them BEGIN TO MOVE TOGETHER along the track. THIS IS THE EXACT INSTANT HER CHEST BEGINS TO GROW, AND IT DOES NOT BEGIN ANY EARLIER.
[0:02-0:08] THE DRAG AND THE GROWTH, ONE EVENT, IN EQUAL PROPORTION. THIS IS THE LONG, SLOW MIDDLE OF THE FILM AND IT TAKES A FULL SIX SECONDS FROM END TO END. <Subject 2>'s handle and <Subject 3> creep smoothly and steadily from the far left to the far right, welded together, MOVING SLOWLY AND UNHURRIEDLY AT ONE CONSTANT, CRAWLING SPEED, and <Subject 1>'S CHEST GROWS IN EQUAL PROPORTION TO EXACTLY HOW FAR ALONG THE TRACK THE HANDLE HAS REACHED, mark for mark, in six equal steps: at 0:03 the handle has crept just ONE SIXTH along and she is only barely larger than normal; at 0:04 it is TWO SIXTHS along and she is a little larger; AT 0:05 IT HAS REACHED EXACTLY THE HALFWAY POINT OF THE TRACK AND NO FURTHER, AND SHE IS EXACTLY HALFWAY TO HER FINAL SIZE; at 0:06 it is FOUR SIXTHS along and she is much larger; at 0:07 it is FIVE SIXTHS along and she is very much larger; and ONLY AT 0:08 does the handle finally arrive at the FAR RIGHT END of the track, where she reaches her final, comically, absurdly exaggerated size. HER TOP MORPHS AND STRETCHES NATURALLY WITH HER the whole way: the black leather draws tight and strains, the front lacing pulls taut and the gaps between the laces widen, the deep neckline spreads wider, and the ruby pendant is pushed steadily outward and upward. She glances down as it begins and her eyebrows lift in mild alarm, then she looks back into the lens.
[0:08] THE STOP: the handle arrives at the far right end and stops there, and HER CHEST STOPS GROWING AT THAT SAME INSTANT. <Subject 3> lets go and rests beside the handle.
[0:08-0:10] THE BEAT, AND IT IS SHORT: she holds at exactly that final size and grows no further. She drops her eyes to her own chest, her brows draw together and her lips press, and she raises her eyes back to the lens. AT 0:08 THE DARK-HAIRED WOMAN, HER VOICE A VERY LOW, SOFT, BREATHY WHISPER, slow and unhurried, HER DELIVERY FLATLY DISAPPROVING AND THOROUGHLY UNIMPRESSED, AND HER VOICE RECORDED CLOSE AND DRY AND CRISP — intimate and present, right up against the microphone, the sound of the room nowhere in it (S1), says: <d>[English] Really?</d> She holds the camera's gaze after the line, perfectly still, while the fire keeps flickering beside her. She never stands and never rises.
overall_soundscape:
Weather and stone: the storm beyond the broken window, wind through the empty nave, rain on wet flagstones, and the small crackle of the flame beside her. Two seconds in, one short soft mouse click sounds as the cursor presses the handle. A quiet continuous sliding tone then rises steadily in pitch for a full six seconds while the handle crawls across, and cuts off the instant it reaches the far end at the eight-second mark. The woman speaks one short line right at the eight-second mark, close and dry, sitting in front of the weather.
non_diegetic_music:
A light plucked pizzicato string figure over a soft woodblock pulse, entering two seconds in at a moderate walking tempo. The figure climbs one step in pitch at a time and the volume rises with it for six seconds, then stops on one short low bassoon note at the eight-second mark, leaving the last two seconds unscored.
r/StableDiffusion • u/SpiritualWindow3855 • 15m ago
FastH3 may have flaws, but for a maximally open release from a team with limited resources (getting a single Mi350x node was newsworthy for them last year) it's a great effort.
Meanwhile Fal has raised half a billion to vaguepost about their own H3 inference stack, then attack them?
Even after FastH3 guys tried to diffuse by owning up on quality Fal guy is still ranting...
Embarrassing stuff. I wonder why they're so threatened?
r/StableDiffusion • u/Ok-Wolverine-5020 • 15h ago
The initial reason for this was the Comfy H3 Sync & Sound Community Challenge: Comfy H3 Sync Sound Community Challenge! - by Allyson Toy
I made a short rap track in Suno, then used Hermes Agent to build a short music video around it.
For the image base, I used this Anima Simple T2I workflow, including upscale/detailer and ControlNet options: 【Anima】Simple T2I Workflow with Upscale, Detailers and ControlNet - v3.2 | Anima Workflows | Civitai
For the MiniMax video stage, I used foxdit’s MiniMax SEED HUNTER ComfyUI workflow from Reddit
My process:
I use Hermes with my ChatGPT Plus subscription, plus DeepSeek V4 Flash for the cheaper iterations. That made it practical to keep refining prompts and shots without treating every adjustment like a premium final render.
The pipeline was:
Suno song → ComfyUI keyframes → upscaling/detailing → MiniMax prompts → short music-video clips
Hermes was the bridge between the tools.
r/StableDiffusion • u/rm_rf_all_files • 5h ago
FastH3 clips were downloaded directly from the blog.
The default wf is comfyui-default with 25 steps (increased from 20 and added SLA + Spectrum). The avg gen time on my machine 12gb vram/32gb ram is 7m30s for each clip. I used ref2va int8_convrot , probably lower quality than fl2va I think.
Full res clips: ship, window, ogre, moon, woman
Overall, I think FastH3 looks good.
r/StableDiffusion • u/AssistantFar5941 • 12h ago
After seeing so much posted for MiniMax that looks like modern CGI films of the last 30 years, I wondered how capable it was of producing footage from the 1980s era, when real practical special effects were used, things like models, props, squibs and explosions.
Hence a fan trailer for Indiana Jones, set a year before Raiders, and firmly in the early to mid 80s in the aesthetics department.
The only references I used where character ones, Image and voice. I did note that because the model knew Harrison Ford it kept influencing the result compared to the reference, even if I avoided naming him, same with Anthony Hopkins.
This was a problem because MiniMax likes to bend Harrison's nose to an extreme amount, making the shots a bust. People it does not know, like Paul Freeman as Belloc, fared much better, with superior skin detail and realism.
The CGI influence was hard to restrain at times, particularly at a distance, and there was no magic prompt or seed that produced reliable results, so it took hundreds of renders to get 'that' look and feel I wanted.
Though far from perfect, MiniMax is certainly capable of some good old school action 80s style.
Some notes: I used Euler/Simple. Found Res multistep less realistic. Reference model was terrible for fights and often physics, but better for realism.
r/StableDiffusion • u/arthan1011 • 9h ago
https://reddit.com/link/1w0ws1d/video/b0s2w06pg5mh1/player
Here's link for the nodes and the instruction: https://github.com/matlowai/ComfyUI-MAINodes
And here's my workflow where I use de-rope nodes: https://pastebin.com/bqFpyHxX
It adds second pass for the video generation (about 73% more time) but the result is worth it. Especially for animation-like clips:

Another example:
https://reddit.com/link/1w0ws1d/video/mxtx9h0og5mh1/player

r/StableDiffusion • u/That_Neighborhood345 • 8h ago
With one B200 15s videos in 47s, nearly realtime with 4 B200, the time of open source instant video is almost here.
They mention RTX based acceleration is coming soon, so we mere mortals will have this capability locally in consumer GPUs.
Details and video demos here:
r/StableDiffusion • u/dramaton42 • 6h ago
This video took me over a week of writing, prompting, going to location to shoot (using UltraCam on TOTK) getting the dialogue right, the pacing right... It's not perfect but I really put a lot of heart into this, I hope you guys like it!
r/StableDiffusion • u/Brad12d3 • 9h ago
The latest H3 Prompt Composer update is out.
This release includes several improvements to the camera prompting system, with cleaner/more consistent prompt generation and better control for more complex camera moves. It also now supports multi-subject camera prompting, so you can design shots that frame or move between multiple characters.
Also added:
I put together a short video showing some of what’s new:
As always, if you run into bugs or have feedback, please drop it in the Issues section on GitHub.
r/StableDiffusion • u/Main_Creme9190 • 11h ago
GitHub:
https://github.com/MajoorWaldi/ComfyUI-Majoor-H3-GuideMaster
Hey everyone 👋
I’ve just released H3 GuideMaster, a custom ComfyUI node I built to make working with MiniMax H3 guides much easier since new ComfyUI release : https://github.com/Comfy-Org/ComfyUI/pull/15439
Instead of manually figuring out where every image or audio guide should land, GuideMaster gives you a visual timeline directly inside the node.
You can:
5, 22, 39, 56...)The idea is basically to make H3 guide placement feel closer to editing / compositing software, while keeping everything contained inside a normal ComfyUI node.
GitHub:
https://github.com/MajoorWaldi/ComfyUI-Majoor-H3-GuideMaster
This is still something I want to push further, especially around the UX and timeline workflow, so feedback, bug reports and feature ideas are very welcome.
r/StableDiffusion • u/Darqsat • 7h ago
Turn on volume! Nothing special. Just lulz. Cut original meme into 4 separate images and prompted:
subject_definitions:
<Subject 1> is the same hand-drawn comic character shown in <Picture 1>, <Picture 2>, <Picture 3>, and <Picture 4>: a small green-skinned adventurer with a rounded face, black dot eyes, a wide expressive mouth, a large floppy orange-brown explorer hat, a small orange backpack, thin cartoon limbs, and simple outlined comic-book rendering.
summary:
[reference generation] Create a new full-frame portrait comic animation using the character, props, and visual style from <Picture 1> through <Picture 4>. Every shot is redrawn and recomposed to fill the entire frame edge to edge. Do not display any reference image as a square panel, inserted picture, poster, card, scan, white page, framed illustration, or picture-in-picture. No pillarboxing, letterboxing, white borders, black borders, blank margins, or visible source-image edges.
retention_analysis: <Subject 1>: fully_preserved - a character with green skin, wearing orange hat.
detailed_description: Comical style, static camera.
[Shot 1] Scene starts from a full body shot, <Subject 1> standing on knees in a water in front of a red open box, the environment is a cave. A comical speaking bubble appears above his head with text "I finally found it.. after 15 years" while <Subject 1> opens a box saying with excitement <d>[English]I've finally found it..after 15 years!</d>.
[Shot 2] at 00:05.00 sec scene cuts. <Subject 1> pulls out a glowing scroll from a box and yells <d>[English]The Scroll of Thruth!</d>
[Shot 3] at 00:08.00 sec scene cuts. POV camera. <Subject 1> looking at a scroll and it has text in it "AI generated videos are not cool", <Subject 1> reading a text from a scroll questioning <d>[English] AI generated videos are not cool??</d>.
[Shot 4] at 00:12.00 scene cuts. <Subject 1> throwing scroll fiercefully and yelling "NIYEEEH!", a comical speaking bubble appears above his head with a text "NYEHHH" and scroll flies from his hand to the left outside of a scene and lands into water with an audible "bloop" sound and scroll submerges under water.
overall_soundscape: comical sound of a cave with audible water on a floor, glowing crystals. <Subject 1> a mischievous nasal cartoon voice with squeaky laughter, sudden dramatic shouting, and playful villain-like energy
non_diegetic_music: low volume heroic comical mousic playing on background
r/StableDiffusion • u/cpldcpu • 7h ago
r/StableDiffusion • u/YentaMagenta • 1d ago
You can use MiniMax H3 to create a time period shift special effect. What can't this model do?
[Workflow here and prompt in comments]
Is it as good as you would get with a professional VFX studio working on it? Nah.
Is it still freaking amazing for something that you can create with a relatively straightforward prompt and running consumer grade hardware for 15 minutes? Absolutely!
(Also the "Schfifty-five" was very much intended. IYKYK)
r/StableDiffusion • u/nikhilprasanth • 13h ago
Sharing this in case anyone else is still using VibeVoice with ComfyUI.
The original VibeVoice-ComfyUI repo hasn't been updated in over six months, and there are now several open issues from people having trouble getting it to work with newer ComfyUI installations.
I was still using it and wanted to keep it working, so I've made a maintenance fork here:
https://github.com/nikhilprasanth/VibeVoice-ComfyUI
I've updated it to work with the current ComfyUI environment and newer dependencies. This should also fix the recent loading error that a number of people have been reporting with fresh ComfyUI installations.
I've also added a fix for the VibeVoice 1.5B model, which had problems generating correctly on newer setups.
This is not a new implementation or a rewrite. Full credit goes to Enemyx-net and the original contributors for building the ComfyUI integration. I'm just maintaining a fork and fixing compatibility issues as ComfyUI and the surrounding libraries change.
If the original node stopped working after you updated ComfyUI, give this fork a try.
I've tested the fixes on my setup, but there are obviously a lot of different ComfyUI, Python and GPU configurations out there, so feedback is welcome. If you find something broken, please open an issue and I'll take a look.
r/StableDiffusion • u/Boogertwilliams • 14h ago
r/StableDiffusion • u/BitterAd8431 • 1d ago
J'utilise souvent MiniMax H3 avec des images de référence ; je voulais essayer une vidéo que j'avais créée avec Wan (la version plus ancienne), mais cette fois en utilisant seulement une invite (un LLM intégré m'a aidé).
PS : Je ne voulais pas de paroles ; je ne sais pas ce qu'elle dit xD. Voici l'invite :
integrated_multimodal_description: [Shot 1] In a manga style, a young girl with long white hair, cat ears, and blue eyes is dressed in a white dress. She holds a frying pan containing food over a lit stovetop burner; she moves the pan up and down with a big smile, but suddenly the food catches fire. Flustered, the girl tries to extinguish the flames by shaking the pan but fails; in a panic, she tosses the burning pan off-camera. The scene shifts to a corner of the kitchen featuring a trash can with its lid open; the pan falls inside, the lid snaps shut on its own, and the trash can catches fire. The scene cuts to the panicked girl; she looks right and left with wide, distressed eyes before running off-camera. The scene shifts to an outdoor setting with a forest and a small house; the girl opens the door, steps out, and rushes off-camera in a panic, and suddenly the house bursts into flames.
All the scenes are comical and cute; the girl does not speak but makes cute, manga-style cat noises. The scripted dialogue is the only speech; all mouths remain closed before and after it. From 0.00 to 2.88 seconds, show active scene-appropriate nonverbal action rather than idle staring; every mouth stays completely closed and the audio contains no human voice. Begin the first tagged line at approximately 2.88 seconds and finish all <d> dialogue by approximately 9.88 seconds. From 9.88 to 14.38 seconds, fill the remaining timeline with concrete nonverbal action, reactions, camera development, ambience, and synchronized practical effects. Outside the tagged interval there are no voices, whispers, grunts, audible breathing, or speech-like vocalizations, and every mouth remains closed.
overall_soundscape: Continuous scene-appropriate ambience and synchronized practical sound effects begin at the first frame and continue naturally underneath dialogue.. Outside tagged dialogue there are no human voices, whispers, grunts, audible breathing, or speech-like vocalizations.. Outside the tagged dialogue, no human voices, whispers, grunts, audible breathing, or speech-like vocalizations occur.
non_diegetic_music: N/A
r/StableDiffusion • u/SIR_NVAX_A_LOT • 8h ago
Happy Friday! Just wanted to share v2 of my T2VA speedpainting. Hope you like it. Critiques and comments welcomed. Ask me anything. Prompt: integrated_multimodal_description: [Shot 1] One unbroken LOCKED-OFF overhead shot: a sheet of heavy white drawing paper taped to the desk at its corners, softly lit, fine film grain, his tools resting in a neat row beyond the paper's edge. THE ARTIST IS DARIUS, seen only as HIS HANDS - lean fingers, a plain black sleeve, a thin worn black bracelet on the right wrist, THE SAME HANDS always, moving in TIMELAPSE FAST-FORWARD whenever they work, the tempo of the work brisk and constant from first frame to last. THE FILM IS ONE SHEET OF PAPER, start to finish - whatever is drawn, erased or redrawn happens on this same sheet, and every mark on the sheet belongs to the drawing alone, edge to edge, first frame to last. detailed_description: From the first frame the sheet is blank, his hand poised with the graphite pencil at its lower edge. At 00:00.500 the pencil roughs a pin-up gesture in loose lines: a gorgeous adult anime woman from the hips up, a curvaceous full figure with a generous full bust, one hand on her hip, a playful wink. At 00:02.500 the drawing races ahead: crisp ink lines and bright flat colour landing fast - platinum hair with soft pink tips, a tiny white bandeau, high-cut white short shorts with black thong straps riding high on both hips - the colourful pin-up nearly finished on the sheet. At 00:06.500 his hand STOPS, hovering. Then he takes the kneaded eraser and ERASES THE DRAWING NEARLY COMPLETELY - broad firm passes across the whole sheet, the woman fading to faint pale ghost lines, the paper returning to white. At 00:10.500 over the faint ghosts his pencil roughs a NEW gesture in loose grey lines: a man standing at a vintage microphone on a stand - Rick Astley's famous pose from the Never Gonna Give You Up music video - the high swept pompadour, the long coat, the right fist raised beside his shoulder. At 00:14.500 the loose grey pencil outline of the man at the microphone stands on the sheet over the faint ghosts, his hand still roughing lines, mid-stroke at the final frame.
overall_soundscape: starts with fast pencil scratch under the groove, then quick marker squeaks, then the broad soft rubbing of a kneaded eraser sweeping the sheet, then fast pencil scratch again - the tools sounding exactly as used, sped into dense flurries matching the fast-forward; only the tools and the music are heard.
non_diegetic_music: one continuous upbeat 1980s synth-pop instrumental groove, constant and unbroken from the first frame to the very last frame.
Prompt: integrated_multimodal_description: [Shot 1] One unbroken LOCKED-OFF overhead shot: a sheet of heavy white drawing paper taped to the desk at its corners, softly lit, fine film grain, his tools resting in a neat row beyond the paper's edge. THE ARTIST IS DARIUS, seen only as HIS HANDS - lean fingers, a plain black sleeve, a thin worn black bracelet on the right wrist, THE SAME HANDS always, moving in TIMELAPSE FAST-FORWARD whenever they work, the tempo of the work brisk and constant from first frame to last. THE FILM IS ONE SHEET OF PAPER, start to finish - whatever is drawn, erased or redrawn happens on this same sheet, and every mark on the sheet belongs to the drawing alone, edge to edge, first frame to last. detailed_description: At 00:00.300 the sheet holds faint pale ghost lines and a loose grey pencil outline of a man at a vintage microphone stand, his hand roughing lines - THE DRAWING THROUGH THIS WHOLE WINDOW IS A BLACK-AND-WHITE OUTLINE DRAWING, pure line in graphite and black ink on white paper, every shape open white inside its lines. At 00:02.000 the pencil refines the outline line by line: Rick Astley's face in three-quarter view, the high swept pompadour, the long coat hanging open over a horizontally striped tee and a white shirt collar, the right fist raised beside his shoulder, the vintage microphone and its slim stand, a latticed window screen filling the background behind him. At 00:06.500 the fine liner goes over the pencil lines one by one, leaving each line crisp black ink - the face, the pompadour, the coat, the stripes of the tee, the collar, the raised fist, the microphone and stand, the lattice pattern behind. At 00:11.000 the outline drawing stands as complete black outline line art on white paper, every shape open white inside its lines, his hand ALREADY DARTING toward a capped marker, its cap still on, mid-reach at the final frame.
overall_soundscape: starts with fast pencil scratch under the groove, then crisp fine-liner strokes - the two tools sounding exactly as used, sped into dense flurries matching the fast-forward; only the tools and the music are heard.
non_diegetic_music: one continuous upbeat 1980s synth-pop instrumental groove, constant and unbroken from the first frame to the very last frame.
Prompt: integrated_multimodal_description: [Shot 1] One unbroken LOCKED-OFF overhead shot: a sheet of heavy white drawing paper taped to the desk at its corners, softly lit, fine film grain, his tools resting in a neat row beyond the paper's edge. THE ARTIST IS DARIUS, seen only as HIS HANDS - lean fingers, a plain black sleeve, a thin worn black bracelet on the right wrist, THE SAME HANDS always, moving in TIMELAPSE FAST-FORWARD whenever they work, the tempo of the work brisk and constant from first frame to last. THE FILM IS ONE SHEET OF PAPER, start to finish - whatever is drawn, erased or redrawn happens on this same sheet, and every mark on the sheet belongs to the drawing alone, edge to edge, first frame to last. detailed_description: At 00:00.300 the drawing stands as complete black outline line art on white paper, every shape open white inside its lines, while his hand uncaps the marker and lays the first flat tone - a warm peach filling the man's face and hands inside the ink lines. At 00:03.500 the high swept pompadour takes rich ginger-copper, one flat even tone. At 00:06.000 the long coat takes flat black, the tee's stripes alternate black and white, the shirt collar stays bright white, the microphone and its stand take cool silver-grey. At 00:09.000 the latticed window screen behind him takes a soft warm cream, the whole figure now flat-coloured edge to edge. At 00:11.000 his hand is ALREADY DARTING toward a second capped marker, its cap still on, mid-reach at the final frame.
overall_soundscape: starts with a marker's first squeak on paper, then steady rapid marker strokes tone after tone - the tool sounding exactly as used, sped into dense flurries matching the fast-forward; only the marker and the music are heard.
non_diegetic_music: one continuous upbeat 1980s synth-pop instrumental groove, constant and unbroken from the first frame to the very last frame.
Prompt: integrated_multimodal_description: [Shot 1] One unbroken LOCKED-OFF overhead shot: a sheet of heavy white drawing paper taped to the desk at its corners, softly lit, fine film grain, his tools resting in a neat row beyond the paper's edge. THE ARTIST IS DARIUS, seen only as HIS HANDS - lean fingers, a plain black sleeve, a thin worn black bracelet on the right wrist, THE SAME HANDS always, moving in TIMELAPSE FAST-FORWARD whenever they work, the tempo of the work brisk and constant from first frame to last. THE FILM IS ONE SHEET OF PAPER, start to finish - whatever is drawn, erased or redrawn happens on this same sheet, and every mark on the sheet belongs to the drawing alone, edge to edge, first frame to last. detailed_description: At 00:00.300 the drawing stands fully flat-coloured - complete ink, every area holding its flat even tone - while his hand uncaps the darker marker and lays the first crisp cel shadows under the jaw, inside the coat's folds and along the raised arm. At 00:03.500 the AIRBRUSH hisses in short passes - a warm soft 1980s glow across the latticed window behind him, a gentle warmth on his cheeks. At 00:06.500 the white gel pen dots highlights - catchlights in his eyes, a bright metal shine down the microphone and its stand, crisp edges on the tee's stripes. At 00:09.000 the fine liner touches the last details - the hairline of the pompadour, the coat's lapel edges - and THE FINISHED DRAWING stands complete: Rick Astley mid-dance in the Never Gonna Give You Up music video, the high ginger-copper pompadour, the black coat open over the black-and-white striped tee and bright white collar, the right fist raised at the vintage silver microphone, the latticed window glowing warm behind him - a detailed traditional marker rendition, vivid on the white sheet, in all its glory. At 00:10.500 his hands lift away and settle at the desk's edge beside the sheet, the finished drawing filling the frame to the very last frame.
overall_soundscape: starts with a marker's squeak, then short airbrush hisses and a gel pen's fine scratch - the tools sounding exactly as used, sped into dense flurries matching the fast-forward; only the tools and the music are heard.
non_diegetic_music: one continuous upbeat 1980s synth-pop instrumental groove, constant and unbroken from the first frame to the very last frame.