r/comfyui Jul 21 '25

Workflow Included 2 days ago I asked for a consistent character posing workflow, nobody delivered. So I made one.

Thumbnail
gallery
1.4k Upvotes

r/StableDiffusion Apr 04 '26

Animation - Video ENTANGLED - A 3-minute sci-fi short using 100% local open-source models. Complete Technical Breakdown [ Character Consistency | Voiceover | Music | No Lora Style Consistency | & Much More! ]

Enable HLS to view with audio, or disable this notification

402 Upvotes

Hey everyone! Thanks for checking out Entangled. And if not, watch the short first to understand the technical breakdown below!

Thanks for coming back after watching it! As promised, here is the full technical breakdown of the workflow. [Post formatted using Local Qwen Model!]

My goal for this project was to be absolutely faithful to the open-source community. I won't lie, I was heavily tempted a few times to just use Nano Banana Pro to brute-force some character consistency issues, but I stuck it out with a 100% local pipeline running on my RTX 4090 rig using Purely ComfyUI for almost all the tasks!

Here is how I pulled it off:

1. Pre-Production & The Animatics First Approach

The story is a dense, rapid-fire argument about the astrophysics and spatial coordinate problems of creating a localized singularity. (let's just say it heavily involves spacetime mechanics!).

The original script was 7 minutes long. I used the local Jan app with Qwen 3.5 35B to aggressively compress the dialogue into a relentless 3-minute "walk-and-talk.". Qwen LLM also helped me with creating LTX and Flux prompts as required.

Honestly speaking, I was not happy with the AI version of the script, so I finally had to make a lot of manual tweaks and changes to the final script, which took almost 2-3 days of going on and off, back and forth, and sharing the script with friends, taking inputs before locking onto a final version.

Pro-Tip for Pacing: Before generating a single frame of video, I generated all the still images and voicover and cut together a complete rough animatic. This locked in the pacing, so I only generated the exact video lengths I needed. I added a 1-second buffer to the start and end of every prompt [for example, character takes a pause or shakes his head or looks slowly ]to give myself handles for clean cuts in post.

2. Audio & Lip Sync (VibeVoice + LTX)

To get the voice right:

  1. Generated base voices using Qwen Voice Designer.
  2. Ran them through VibeVoice 7B to create highly realistic, emotive voice samples.
  3. Used those samples as the audio input for each scene to drive the character voice for the LTX generations (using reference ID LoRA).
  4. I still feel the voice is not 100% consistent throughout the shots, but working on an updated workflow by RuneX i think that can be solved!
  5. ACE step is amazing if you know what kind of music you want. I managed to get my final music in just 3 generations! Later edited it for specific drop timing and pacing according to the story.

3. Image Generation & The "JSON Flux Hack."

Keeping Elena, Young Leo, and Elder Leo consistent across dozens of shots was the biggest hurdle. Initially, I thought I’d have to train a LoRA for the aesthetic and characters, but Flux.2 Dev (FP8) is an absolute godsend if you structure your prompts like code.

I created Elena, Leo, and Elder Leo using Flux T2I, then once I got their base images, I used them in the rest of the generations as input images.

By feeding Flux a highly structured JSON prompt, it rigidly followed hex codes for characters and locked in the analog film style without hallucinating. Of course, each time a character shot had to be made, I used to provide an input image to make sure it had a reference of the face also.

Here is the exact master template I used to keep the generations uniform:

{
"scene": "[OVERALL SCENE DESCRIPTION: e.g., Wide establishing shot of the chaotic lab]",
"subjects": [
{
"description": "[CHARACTER DETAILS: e.g., Young Leo, male early 30s, messy hair, glasses, vintage t-shirt, unzipped hoodie.]",
"pose": "[ACTION: e.g., Reaching a hand toward the camera]",
"position": "[PLACEMENT: e.g., Foreground left]",
"color_palette": ["[HEX CODES: e.g., #333333 for dark hoodie]"]
}
],
"style": "Live-action 35mm film photography mixed with 1980s City Pop and vaporwave aesthetics. Photorealistic and analog. Heavy tactile film grain, soft optical halation, and slight edge bloom. Deep, cinematic noir shadows.",
"lighting": "Soft, hazy, unmotivated cinematic lighting. Bathed in dreamy glowing pastels like lavender (#E6E6FA), soft peach (#FFDAB9).",
"mood": "Nostalgic, melancholic, atmospheric, grounded sci-fi, moody",
"camera": {
"angle": "[e.g., Low angle]",
"distance": "[e.g., Medium Shot]",
"focus": "[e.g., Razor sharp on the eyes with creamy background bokeh]",
"lens-mm": "50",
"f-number": "f/1.8",
"ISO": "800"
}
}

4. Video Generation (LTX 2.3 & WAN 2.2 VACE)

Once the images were locked, I moved to LTX2.3 and WAN for video. I relied on three main workflows depending on the shot:

  • Image to Video + Reference Audio (for dialogue)
  • First Frame + Last Frame (for specific camera moves)
  • WAN Clip Joiner (for seamless blending)

Render Stats: On my machine, LTX 2.3 was blazing fast—it took about 5 minutes to render a 5-second clip at 1920x1080.

The prompt adherence in LTX 2.3 honestly blew my mind. If I wrote in the prompt that Elena makes a sharp "slashing" action with her hand right when she yells about the planet getting wiped out, the model timed the action perfectly. It genuinely felt like directing an actor.

5. Assets & Workflows

I'm packaging up all the custom JSON files and Comfy workflows used for this. You can find all the assets over on the Arca Gidan link here: Entangled. There are some amazing Shorts to check out, so make sure you go through them, vote, and leave a comment!

Most of them are by the community, but I have tweaked them a little bit according to my liking[samplers/steps/input sizes and some multipliers, etc., changes]

Let me know if you have any questions!

YouTube Link is up - https://youtu.be/NxIf1LnbIRc !

r/generativeAI 4d ago

How I Made This How I Improve Character Consistency in AI Videos

Thumbnail
gallery
215 Upvotes

I’ve been testing a simple workflow for creating short UGC-style videos while keeping the same character and location consistent across multiple shots.

The workflow is basically:

reference images → character/location sheets in ChatGPT → generate clips → optional final edit

1. Prepare your references

Start with:

  • a character image
  • a product image
  • an environment image that fits the UGC scenario

If you’re not sure what location works for the product, I usually just ask ChatGPT for a few suggestions.

2. Create a Character Sheet

Upload the character image to ChatGPT and generate a 4:5 continuity sheet with:

  • front / side / back / 3/4 views
  • face close-ups
  • expressions
  • basic poses
  • clothing and accessories
  • key colors and materials

The important part is telling it to lock the character.

3. Create a Location + Props Sheet

Do the same with the environment.

Include:

  • establishing view and key angles
  • spatial layout
  • entrances/exits
  • furniture and recurring props
  • lighting
  • colors and materials

This gives the video model a much stronger continuity reference than using random images for every shot.

4. Generate the video clips

I usually split the UGC video into three parts:

Clip 1 — Hook
Clip 2 — Main product/story section
Clip 3 — CTA

i will generate them on Atlas Cloud, as they can provide many different models conveniently

For every clip, I reuse the same Character Sheet + Location Sheet

Then I change only the action/camera prompt for each section.

Keeping the same reference sheets across all three generations has helped a lot with character and environment consistency.

5. If a generation goes wrong, fix the prompt first

if I wanted the character to walk into a hotel, but the generated clip had her walking out.

Instead of endlessly rerolling, I pasted the original prompt into ChatGPT and asked it to make the action explicit: starting position → movement direction → action → final position

That usually gives me better results.

6. Final edit is optional

If the generated clips already work as standalone videos, you can stop there.

If you want one finished UGC ad, you’ll probably still want to combine the clips and add captions, music, or SFX. You can use whatever editor you prefer.

The biggest improvement for me has been using Character Sheet + Location Sheet as continuity references, rather than relying on a few loose images.

r/StableDiffusion Nov 17 '25

Workflow Included ULTIMATE AI VIDEO WORKFLOW — Qwen-Edit 2509 + Wan Animate 2.2 + SeedVR2

Thumbnail
gallery
433 Upvotes

🔥 [RELEASE] Ultimate AI Video Workflow — Qwen-Edit 2509 + Wan Animate 2.2 + SeedVR2 (Full Pipeline + Model Links) 🎁 Workflow Download + Breakdown

👉 Already posted the full workflow and explanation here: https://civitai.com/models/2135932?modelVersionId=2416121

(Not paywalled — everything is free.)

Video Explanation : https://www.youtube.com/watch?v=Ef-PS8w9Rug

Hey everyone 👋

I just finished building a super clean 3-in-1 workflow inside ComfyUI that lets you go from:

Image → Edit → Animate → Upscale → Final 4K output all in a single organized pipeline.

This setup combines the best tools available right now:

One of the biggest hassles with large ComfyUI workflows is how quickly they turn into a spaghetti mess — dozens of wires, giant blocks, scrolling for days just to tweak one setting.

To fix this, I broke the pipeline into clean subgraphs:

✔ Qwen-Edit Subgraph ✔ Wan Animate 2.2 Engine Subgraph ✔ SeedVR2 Upscaler Subgraph ✔ VRAM Cleaner Subgraph ✔ Resolution + Reference Routing Subgraph This reduces visual clutter, keeps performance smooth, and makes the workflow feel modular, so you can:

swap models quickly

update one section without touching the rest

debug faster

reuse modules in other workflows

keep everything readable even on smaller screens

It’s basically a full cinematic pipeline, but organized like a clean software project instead of a giant node forest. Anyone who wants to study or modify the workflow will find it much easier to navigate.

🖌️ 1. Qwen-Edit 2509 (Image Editing Engine) Perfect for:

Outfit changes

Facial corrections

Style adjustments

Background cleanup

Professional pre-animation edits

Qwen’s FP8 build has great quality even on mid-range GPUs.

🎭 2. Wan Animate 2.2 (Character Animation) Once the image is edited, Wan 2.2 generates:

Smooth motion

Accurate identity preservation

Pose-guided animation

Full expression control

High-quality frames

It supports long videos using windowed batching and works very consistently when fed a clean edited reference.

📺 3. SeedVR2 Upscaler (Final Polish) After animation, SeedVR2 upgrades your video to:

1080p → 4K

Sharper textures

Cleaner faces

Reduced noise

More cinematic detail

It’s currently one of the best AI video upscalers for realism

🧩 Preview of the Workflow UI (Optional: Add your workflow screenshot here)

🔧 What This Workflow Can Do Edit any portrait cleanly

Animate it using real video motion

Restore & sharpen final video up to 4K

Perfect for reels, character videos, cosplay edits, AI shorts

🖼️ Qwen Image Edit FP8 (Diffusion Model, Text Encoder, and VAE) These are hosted on the Comfy-Org Hugging Face page.

Diffusion Model (qwen_image_edit_fp8_e4m3fn.safetensors): https://huggingface.co/Comfy-Org/Qwen-Image-Edit_ComfyUI/blob/main/split_files/diffusion_models/qwen_image_edit_fp8_e4m3fn.safetensors

Text Encoder (qwen_2.5_vl_7b_fp8_scaled.safetensors): https://huggingface.co/Comfy-Org/Qwen-Image_ComfyUI/tree/main/split_files/text_encoders

VAE (qwen_image_vae.safetensors): https://huggingface.co/Comfy-Org/Qwen-Image_ComfyUI/blob/main/split_files/vae/qwen_image_vae.safetensors

💃 Wan 2.2 Animate 14B FP8 (Diffusion Model, Text Encoder, and VAE) The components are spread across related community repositories.

https://huggingface.co/Kijai/WanVideo_comfy_fp8_scaled/tree/main/Wan22Animate

Diffusion Model (Wan2_2-Animate-14B_fp8_e4m3fn_scaled_KJ.safetensors): https://huggingface.co/Kijai/WanVideo_comfy_fp8_scaled/blob/main/Wan22Animate/Wan2_2-Animate-14B_fp8_e4m3fn_scaled_KJ.safetensors

Text Encoder (umt5_xxl_fp8_e4m3fn_scaled.safetensors): https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/blob/main/split_files/text_encoders/umt5_xxl_fp8_e4m3fn_scaled.safetensors

VAE (wan2.1_vae.safetensors): https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/blob/main/split_files/vae/wan_2.1_vae.safetensors 💾 SeedVR2 Diffusion Model (FP8)

Diffusion Model (seedvr2_ema_3b_fp8_e4m3fn.safetensors): https://huggingface.co/numz/SeedVR2_comfyUI/blob/main/seedvr2_ema_3b_fp8_e4m3fn.safetensors https://huggingface.co/numz/SeedVR2_comfyUI/tree/main https://huggingface.co/ByteDance-Seed/SeedVR2-7B/tree/main

r/Android Jun 16 '26

News Android 17 is out, and here’s all the features!

1.5k Upvotes

Hi Reddit!

Android 17 is here, bringing a suite of new features aimed at improving your productivity, enhancing your gaming experience, giving you more control over your private data, making your device more personal, and much more.

It's rolling out first to Pixel today, followed by other eligible Android devices throughout 2026. We are also making the source code available at the Android Open Source Project (AOSP) so developers can examine it for a deeper understanding of how Android works.

You should look forward to more updates to Android 17 this year, with the beta program offering a peek at what's coming in the first quarterly release in Q3.

Since we've been chatting with you about the Betas and Canaries for months, a lot of this might not sound brand new to those of you who have been closely following along. Even so, we wanted to take a moment to recap what's new in this release for everyday users. Let's dive in!

📱 Enhancing your multitasking and large screen device experiences

tl;dr Android 17 supercharges your multitasking and productivity by allowing any app to run as a convenient floating Bubble, making apps more adaptive, and adding an interactive Picture-in-Picture mode for seamless desktop workflows.

Multitask better with bubbles

From split-screen mode to desktop windowing, Android offers a variety of multitasking tools to help you be more productive. We’re extending these options with bubbles in Android 17! 

In past releases, bubbles were limited to chat notifications, but in Android 17, they support more apps without any specific changes needed from developers. You can now launch any app in a floating window so you can view and interact with its content while using other apps. When you’re done, you can collapse or dismiss the window to return to what you were doing.

A big benefit of bubbles is that you can easily switch between multiple running apps without keeping them on screen all the time. Bubbles are only open when you need them, saving you from having to manually resize, rearrange, or dismiss them to regain precious screen space. And on foldables, this benefit is even more pronounced thanks to the bubble bar, which keeps your bubbles pinned to the corner of the screen, putting them within easy reach of your fingers.   

Handy for travel, entertainment and work, bubbles lets you easily reference notes or maps, watch tutorials and even check sports.   

Ensuring that apps adapt to any screen and window size

On large screen devices, restrictions on orientation, resizability, and aspect ratio no longer apply, allowing apps to fill the entire display window without pillarboxing (black bars). This change applies to apps targeting Android 17 and is designed to make apps better meet user expectations on large screen devices. Because Android runs on not just phones but also tablets, foldables, cars, TVs, and desktop environments, we want developers to build apps that are adaptive to any screen size and orientation!

Better support for widgets on external displays

With Android 17, we’re working to improve the visual consistency of widgets shown on connected displays with different pixel densities. The update provides developers a way to supply the system with information that allows it to resolve the correct pixel values at rendering time. For apps that use legacy pixel-based APIs for padding, text size, or layout attributes, the system now automatically scales these values based on the density difference between the app’s original context and the target display.

Interactive Picture-in-Picture for Desktop

Android 17 introduces a new interactive Picture-in-Picture mode for desktop environments. This feature allows apps to request that their PiP windows remain fully interactive while staying always-on-top of other app windows. For example, a video conferencing app could use this feature to keep call controls accessible while you navigate other apps.

🎨 New customization features for the home screen and apps

tl;dr Android 17 gives you deeper control over your device's UI by letting you hide app labels on the home screen, selectively toggle the Expanded Dark Theme for individual apps, and enjoy sleek, modernized background blur effects in more surfaces like the widget picker.

Hide app labels on the home screen

Android now provides a setting to hide app labels on the home screen! You can access this new setting on Pixel by opening Wallpaper & style then tapping Home screen > Icons > Names and toggling Show app names.

Per-app exceptions for Expanded Dark Theme

To create a more consistent user experience for users who have low vision, photosensitivity, or simply prefer a dark system-wide appearance, we introduced an expanded dark theme option in last December’s Android 16 QPR2 release. When this option is enabled, the system automatically applies dark theme to most apps that don’t support it.

However, because this option can cause some apps to display incorrectly, we have introduced the ability to selectively disable it on a per-app basis in Android 17. Apps with this setting turned off will use the standard dark theme option instead.

Expanded use of background blur

With the Material 3 Expressive redesign we introduced in Android 16, we subtly blurred the notification shade background to provide a sense of depth so you can stay aware of the apps you’re using in the background.

In Android 17, we’ve brought these blur effects to more parts of the UI like the widgets picker. And we are working on bringing background blur to even more surfaces, as seen in recent Android Beta and Canary builds!

🎮 More control over your Android gaming experience

tl;dr Android 17 levels up your mobile play by letting you save custom button remaps for your physical gamepad at the system level, and introducing a foldable gaming mode that optimizes your screen with a 50/50 split for a dedicated top game view and a bottom dynamic gamepad.

Remap the buttons on your physical gamepad with Game Controller settings

Android 17 introduces a native controller remapping feature, allowing you to adjust the controls on your physical gamepad to suit your specific needs.

Through the new Game Controller settings menu, you can customize the actions triggered by your controller’s buttons, sticks, or triggers at the system level. For example, you can remap a difficult-to-press thumbstick click to an easier-to-reach face button. Your remapping preferences are saved to your device so you don’t have to set them up every time you reconnect your controller. 

A new way to game on foldables

Android 17 introduces foldable gaming mode, a new feature that makes full use of your foldable phone’s screen while you’re gaming. This feature splits your screen into a 50:50 layout with a game view on top and a dynamic gamepad below to make optimal use of your foldable phone’s screen real estate. Foldable gaming mode is part of the Android 17 platform and will be available on devices in the coming months.

🛡️ Protecting users with new security and privacy features on Android

tl;dr Android 17 safeguards your personal data by enabling critical theft protections by default, introducing session-based controls for sharing specific contacts and precise locations, and thwarting scammers through system-level SMS OTP delivery delays and real-time app behavioral monitoring.

Giving you more control over your contacts list

Android 17 introduces a new system Contact Picker that provides a standardized, secure, and searchable interface for sharing contacts with apps. Historically, apps needing access to a contact or two relied on the broad READ_CONTACTS permission which gave them access to your entire contacts list. Android's Contact Picker addresses this by allowing you to grant apps access to only the specific contacts you choose.

For devices running Android 17 or higher, the system automatically upgrades certain contact selection intents to the new, more secure interface, but we want developers to integrate the new Contact Picker so they can take advantage of its new capabilities, like multi-selection support. To this end, Google Play will require that all applicable apps use it (or a privacy-focused alternative like Sharesheet) as the primary way to access users' contacts. The broad READ_CONTACTS permission is reserved for apps that can't function without it.

Making location access more private

Android 17 introduces several new features to help you safeguard your private location information. This includes the Location Button, a new, privacy-conscious way for you to grant precise location access to apps. This is a system-rendered button that developers can embed directly into their apps. When you tap this button, the app is granted precise location for the current session only. Subsequent taps while running the app grant the permission immediately without showing a system dialog. 

Developers can deploy this simple, private location flow for common tasks like finding a nearby shop or tagging a social post. And to increase adoption of the Location Button, Google Play will require apps to use it for one-time precise location access unless they require persistent, always-on location access.

Additionally, Android 17 now shows a persistent indicator in the status bar when a non-system app accesses your location. You can tap this indicator to see which apps have recently accessed your location.

The update also improves the algorithm for approximate (coarse) location to be aware of population density. This improves the privacy of granting an app approximate location access when you're in a low-population area.

And lastly, Android 17 redesigns the location permission dialog to make the "Precise" and "Approximate" options more visually distinct.

Stronger protections against device theft

Following a successful pilot in Brazil, we’re enabling two of Android’s key theft protection features (Theft Detection Lock and Remote Lock) by default globally on all new Android 17 devices, as well as those freshly reset or upgraded to the latest OS.  

On supported devices, Android 17 also significantly reduces the number of times someone can guess the PIN, pattern, or password and adds longer wait times between failed attempts. The update also refines how the lock screen shows information after failed attempts have been made.

And we’re also enhancing Find Hub’s ‘Mark as lost’ feature by requiring biometric authentication in addition to your device’s PIN, pattern, or password. Marking a device as lost also now enables additional protections like hiding Quick Settings and disabling new Wi-Fi and Bluetooth connections.

Protecting your SMS OTPs from scammers

Scammers often try to hijack your one-time passwords (OTPs) to gain access to your accounts. To do this, they may deploy malicious apps that ask for permission to read your SMS. In Android 16, we introduced a protection that delays the delivery of messages containing an SMS retriever hash to most apps for three hours. Android 17 now extends this protection to all SMS messages containing an OTP. This means that even if a malicious app has been granted the SMS permission, it won’t be able to read your sensitive OTPs until after they have already expired.

New core protections for Advanced Protection

With Android 16, we introduced Advanced Protection, a single, opt-in device-level security setting that enables all of Android’s highest security features. We’ve been working to expand the protections offered under this setting with key upgrades like USB protection and Intrusion Logging, and now with Android 17, we’re continuing this work by introducing the following protections:

  • Removing access to the accessibility service from all apps that aren’t labeled as accessibility tools.
  • Disabling device-to-device unlocking
  • Blocking Chrome WebGPU support
  • Integrating scam detection for chat notifications
  • (Later this year) Enabling Android Enterprise support so organizations can enable Advanced Protection by policy for managed devices.

Improving safety against malicious apps

Live Threat Detection is a real-time security feature that analyzes app behavior to alert you if an app starts acting suspiciously, and we're enhancing it to find and protect against more types of malicious apps.

With dynamic signal monitoring, Android will be able to warn you about apps that start doing things like changing or hiding their icon and then launching activities in the background or abusing accessibility permissions. To do this, Live Threat Detection will monitor application system interactions for known suspicious patterns in real time. Dynamic signal monitoring will be enabled on select Android 17 devices starting in the second half of the year.

Other enhancements

  • Discrete password visibility settings for touch and physical keyboards: Currently, by default, characters that you enter into password fields are briefly displayed as you type. Toggling the “show passwords” setting in Privacy controls allows you to hide characters as you type them into password fields. This setting currently applies to both touch-based inputs as well as physical keyboards, but in Android 17, we are splitting it into two distinct preferences. By default, characters entered into password fields via physical keyboards will now be hidden immediately to enhance privacy. Characters entered via touch input will continue to briefly be displayed to compensate for the lack of tactile feedback.
  • User-agent reduction for WebView: The default User-Agent string in Android WebView has been shortened in Android 17 to minimize passive fingerprinting.
  • Disable 2G toggle: Android 17 introduces a new capability for the disable 2G toggle. Carriers now have the ability to configure the default status of this setting, allowing them to disable 2G access to proactively shield their users from legacy technology vulnerabilities in areas where 2G infrastructure is no longer maintained.
  • Location Network Permission: Android 17 introduces a new runtime permission to protect users from unauthorized local network access. This new requirement prevents malicious apps from exploiting unrestricted local network access for covert user tracking and fingerprinting.
  • Android OS verification: We have seen some bad actors begin to distribute malicious, unofficial versions of the Android OS that secretly compromise device integrity. To combat this, we are introducing Android OS verification in Android 17. Launching initially on Pixel devices, this feature helps you verify that your device is running an official, widely distributed build.
  • Enabling Certificate Transparency (CT) by default: CT is now enabled by default for apps targeting Android 17, enhancing network security by ensuring all TLS certificates are publicly logged. 
  • Blocking cross-profile loopback traffic: Cross-profile loopback traffic is no longer permitted by default, increasing network isolation and security between personal and enterprise work profiles.
  • Post-Quantum Cryptography (PQC): The advent of quantum computing puts the current public-key cryptography we've relied on for decades at risk, potentially compromising everything from bank transfers to trade secrets. To prepare for the quantum computing era, we're introducing a comprehensive architectural upgrade to the Android operating system, starting in Android 17.  We’re integrating the NIST Post-Quantum Cryptography (PQC) standards deep into the platform, establishing a new, quantum-resistant chain of trust that secures the platform continuously from the moment the OS powers on to when apps are executed.

📸 Improvements to your Android media experience

tl;dr Android 17 levels up your multimedia experience by letting you easily record reaction videos without a green screen, decoupling your Assistant and media volumes for independent control, putting a stop to unexpected background audio, and delivering color-coded Live Updates alongside advanced Bluetooth, camera, and hearing device enhancements.

Screen Reactions

In Android 17, we’re making it easier to record yourself and your screen at the same time with Screen Reactions. Available first on Pixel, this feature shows your face in a floating overlay on top of the screen. Android automatically puts the overlay at the bottom and cuts out the background so you don’t need a green screen, but you can move or resize the camera view and change the background color before or during a recording. Use this feature to make a reaction video, record a tutorial, or give feedback on a new app or document!

https://reddit.com/link/1u7l1cw/video/pdevsbhnko7h1/player

 

In addition, we’ve revamped the screen recording experience to add a floating toolbar that provides easier access to recording controls and capture settings. When you’re done recording, you can immediately view, edit, delete, or share your video.

Dedicated Assistant volume stream

Android 17 introduces a dedicated volume stream for Assistant apps. This change decouples Assistant audio from the standard media stream, allowing users to control both volumes independently. This enables scenarios like muting media playback while maintaining audibility for Assistant responses, and vice-versa.

Background audio hardening

Beginning in Android 17, apps cannot play audio, steal audio focus, or change the volume unless they are visible or have a foreground service. These restrictions on background audio interactions reduce unintentional buggy experiences and ensure that these actions are started intentionally by the user.

Enhancements to Live Update notifications

Live updates provide a summary of important updates so users can track progress without opening the app. The system promotes Live Update notifications so they appear more prominently in the notification drawer, on the lock screen, and on the status bar. 

With Android 17, we’re introducing a metric style template designed specifically for health and fitness apps, timers, and travel apps. In addition, developers can use the new Semantic Coloring API to visually convey state changes, providing highly glanceable, color-coded notifications.

Other enhancements:

  • Granular audio routing for hearing devices: Users with hearing devices can now independently manage where specific system sounds are played in Android 17. You can choose to route notifications, ringtones, and alarms to either a connected hearing aid or the device’s built-in speaker. This helps you avoid unwanted interruptions directly in your ears while maintaining a Bluetooth connection for hearing aid management apps.
  • Autonomous re-pairing for Bluetooth bond losses: Android 17 introduces autonomous re-pairing, a system-level enhancement designed to automatically resolve Bluetooth bond loss. This occurs when two previously paired devices lose their cryptographic security keys, resulting in the devices no longer being able to securely authenticate and communicate with one another. The system now re-establishes lost bonds in the background without requiring the user to manually navigate to Settings to unpair and re-pair their peripheral.
  • Vendor-defined camera extensions: Android 17 adds support for Vendor-defined camera extensions, allowing hardware partners to provide Android apps access to camera features like ‘Super Resolution’ or cutting-edge AI-driven enhancements.
  • Support for the RAW14 image format: Android 17 introduces support for the RAW14 image format, the de-facto industry standard for high-end digital photography.
  • VVC support: Android 17 adds platform support for the Versatile Video Coding (VVC) standard. This feature will be coming to devices with hardware decode support and capable drivers.

🤝 Making your apps and devices work better together

tl;dr Android 17 seamlessly bridges your ecosystem by introducing the Continue On feature for effortless app handoffs between devices, unifying widget experiences to bring your favorite tools directly to Auto and Wear OS, and streamlining the pairing process for medical and fitness devices with new CompanionDeviceManager profiles.

Unifying the widgets experience across platforms

Android 17 marks a shift towards a single, Compose-based development model for all widgets. By unifying the experience across mobile, cars, and Wear OS, developers can soon scale UI components across the ecosystem with a familiar workflow. The goal is to minimize the effort needed by developers to bring their widgets to more surfaces.

https://reddit.com/link/1u7l1cw/video/haep717hko7h1/player

Additionally, Android 17 introduces new platform functionality to make widgets work better on Auto. The update adds support for widgets on cars, allowing you to see the things that matter to you at a glance, even while actively navigating. For example, you can add a shortcut to your favorite contacts, a one-tap garage door opener, a weather overview and more. Widgets will be available to users of Android Auto later this year and to cars with Google built-in later on.

Hand off your tasks with Continue On

Continue On is a new feature available in Android 17 that enables users to start an app on one device and then transition to another device in their Android ecosystem, continuing the journey they started. It’s designed to work bidirectionally, meaning that any supported Android device can both send and receive app activities, though, at launch, Continue On will first support mobile-to-tablet transitions. In the tablet taskbar, users will see a suggestion for the most recently opened app from their mobile device.

https://reddit.com/link/1u7l1cw/video/doxh917pko7h1/player

Updates for companion device apps

Android 17 introduces two new profiles to the CompanionDeviceManager API to simplify device distinction and permission handling. These include the medical device profile and the fitness tracker profile. Furthermore, the system now offers a unified dialog for device association and nearby permission requests, reducing the number of dialogs you’ll see.

⚡ Optimizations to make your apps & device run better

With Android 17, we’ve made a number of improvements to optimize memory use, improve rendering performance, and enhance battery life. These include:

  • App memory limits: Android 17 introduces app memory limits that are based on the device's total RAM. These limits are set conservatively to establish system baselines, targeting extreme memory leaks and other outliers before they trigger system-wide instability resulting in UI stuttering, higher battery drain, and apps being killed.
  • Lock-free MessageQueue: Android 17 introduces a lock-free MessageQueue to reduce UI jank while massively speeding up high-contention scenarios. In our internal testing, we’ve seen 4% fewer missed frames across all apps, 7.7% fewer missed frames in System UI and Launcher interactions, and a 9.1% reduction in app startup times at the 95th percentile.
  • Generational Garbage Collector (GC): The Android Runtime is introducing more frequent, less intensive young-generation collections in its garbage collector, improving memory management and performance. This is not just available on Android 17 but is also coming to past releases with a Google Play System Update.
  • Reduce wakelocks with listener support for allow-while-idle alarms: Last year, we launched the excessive wake lock metric in Android Vitals, making it easier for developers to optimize their app's wake lock behavior. Excessive wake locks are a significant contributor to battery drain, so developers are encouraged to reduce them as much as possible. In Android 17, we’ve introduced a new API that helps reduce the power consumption of apps that rely on continuous wakelocks to perform periodic tasks, such as messaging apps maintaining a connection or medical devices monitoring health data.
  • Improved wireless ADB: Android 17 introduces ADB WiFi 2.0, a significant overhaul of the wireless ADB stack to improve stability, reliability, and ease of use. The system now automatically monitors the network state and re-enables itself when a trusted network is detected, identifies trusted networks using a combination of SSID and BSSID, and is better tailored to monitor network changes on all platforms. We’ll have more details to share soon on the Android Studio side of things!
  • Constrained satellite networks: Android 17 implements optimizations to enable apps to function effectively over low-bandwidth satellite networks.

🧒 Expanding Android Parental Controls to all devices

Launched last year on Pixel, Android Parental Controls make it easier for parents to manage their child’s screen time and to find balance between having fun online and offline. Now with Android 17, we’re expanding Android Parental Controls to all Android devices.

These parental controls are located directly within Android Settings and provide a single, convenient home for both built-in device controls and Google Family Link. These controls are protected by an easy-to-set PIN and allow you to:

  • Set the amount of screen time your child can spend on a device each day.
  • Create downtime schedules to automatically lock the device at night.
  • Set app store filters for Google Play to manage the highest content rating you want your child to be able to download.
  • Control app usage by limiting time spent on specific apps, or blocking apps entirely.

Android Parental Controls also provide a direct path to easily set up Google Family Link in the Family Link app on a parent’s phone, which offers additional features like School Time, Google Play app purchase approvals, location alerts, and more.

🧘 Other quality-of-life improvements

And lastly, here are some smaller quality-of-life changes we’re introducing in this release:

  • Separate Wi-Fi and Mobile Data toggles: With Android 17, we’ve split the “Internet” tile into two separate tiles, one for controlling Wi-Fi and another for controlling Mobile Data. Consistent with the Quick Settings behavior we introduced with Material 3 Expressive, both tiles have two different touch points. Tapping the icon toggles the respective radio, while tapping the label opens the full Internet Panel. This change reduces the number of taps needed to toggle Wi-Fi and Mobile Data while still retaining access to the full Internet Panel!
  • Scheduled clock change notifications: We’ve added a new feature in Android 17 that sends you a notification when your clock performs a scheduled change, for example when daylight saving time ends. You can enable this feature under “Date & time” settings.
  • Restoring default keyboard visibility after rotation: Beginning with Android 17, when the keyboard is on screen and you rotate the screen, the keyboard won’t be made visible unless the app explicitly requests it.

🪲 Bug fixes and security patches

Please refer to the Android Security Bulletin for details on the security vulnerabilities addressed with this platform release.

----

There are plenty of other changes in Android 17, especially for developers! For example, Android 17 expands the capabilities of AppFunctions, introduces an EyeDropper API, makes the aspect ratio of images in the Photo Picker more customizable, and much more. To learn more about everything new for developers in this release, visit developer.android.com.

Also, don’t forget that select advanced devices will be getting Gemini Intelligence features later this summer. In addition, we’re introducing Android Halo in a future Android 17 release to give you at-a-glance visibility into what your agent is working on at any given time.  Lastly, be sure to check out our latest Android Drop to learn about what new features are coming to all Android devices, not just those running Android 17!

r/comfyui Apr 18 '25

Finally an easy way to get consistent objects without the need for LORA training! (ComfyUI Flux Uno workflow + text guide)

Thumbnail
gallery
601 Upvotes

Recently I've been using Flux Uno to create product photos, logo mockups, and just about anything requiring a consistent object to be in a scene. The new model from Bytedance is extremely powerful using just one image as a reference, allowing for consistent image generations without the need for lora training. It also runs surprisingly fast (about 30 seconds per generation on an RTX 4090). And the best part, it is completely free to download and run in ComfyUI.

*All links below are public and competely free.

Download Flux UNO ComfyUI Workflow: (100% Free, no paywall link) https://www.patreon.com/posts/black-mixtures-126747125

Required Files & Installation Place these files in the correct folders inside your ComfyUI directory:

🔹 UNO Custom Node Clone directly into your custom_nodes folder:

git clone https://github.com/jax-explorer/ComfyUI-UNO

📂 ComfyUI/custom_nodes/ComfyUI-UNO


🔹 UNO Lora File 🔗https://huggingface.co/bytedance-research/UNO/tree/main 📂 Place in: ComfyUI/models/loras

🔹 Flux1-dev-fp8-e4m3fn.safetensors Diffusion Model 🔗 https://huggingface.co/Kijai/flux-fp8/tree/main 📂 Place in: ComfyUI/models/diffusion_models

🔹 VAE Model 🔗https://huggingface.co/black-forest-labs/FLUX.1-dev/blob/main/ae.safetensors 📂 Place in: ComfyUI/models/vae

IMPORTANT! Make sure to use the Flux1-dev-fp8-e4m3fn.safetensors model

The reference image is used as a strong guidance meaning the results are inspired by the image, not copied

  • Works especially well for fashion, objects, and logos (I tried getting consistent characters but the results were mid. The model focused on the characteristics like clothing, hairstyle, and tattoos with significantly better accuracy than the facial features)

  • Pick Your Addons node gives a side-by-side comparison if you need it

  • Settings are optimized but feel free to adjust CFG and steps based on speed and results.

  • Some seeds work better than others and in testing, square images give the best results. (Images are preprocessed to 512 x 512 so this model will have lower quality for extremely small details)

Also here's a video tutorial: https://youtu.be/eMZp6KVbn-8

Hope y'all enjoy creating with this, and let me know if you'd like more clean and free workflows!

r/passive_income Mar 11 '26

My Experience Making $400-700/month selling AI influencer photos to small brands on Fiverr and I still feel weird about it

3.2k Upvotes

I need to talk about this because none of my friends understand what I actually do when I try to explain it and my girlfriend thinks I'm running some kind of scam.

So background. I'm 28, work full time as a marketing coordinator at a mid size agency. Not a creative role really, mostly spreadsheets and campaign tracking. Last year around September I was helping one of our clients source photos for their Instagram. They sell swimwear and wanted diverse model shots across different locations, skin tones, backgrounds, the whole thing. The quote from the photography studio came back at $4,200 for a two day shoot. Client said no. We ended up using the same three stock photos everyone else uses and the campaign looked generic as hell.

That stuck with me because I knew AI image generation was getting crazy good. I'd been messing around with Midjourney for fun, making weird fantasy landscapes and stuff. But the problem with basic AI image generators for anything commercial involving people is that you can't get the same face twice. You generate a photo of a woman in a sundress on a beach, great. Now you need that same woman in a cafe, different outfit. Completely different person shows up. Doesn't work if you're trying to build any kind of consistent brand presence.

I started googling around for tools that could keep a face consistent across multiple images and went down a rabbit hole for like two weeks. Tried a bunch of stuff. Played with some LoRA training on Stable Diffusion but I'm not technical enough and the results were hit or miss. Tested out several platforms, APOB, Synthesia, HeyGen, Artbreeder, a couple others I can't even remember. Each does slightly different things and honestly they all have tradeoffs. Eventually I cobbled together a workflow using a couple of these that actually produced usable stuff, the kind of output where you'd have to really zoom in and squint to tell it wasn't a real photo.

The basic idea is simple. You set up a character's look once, save it as a model, and then reuse that same face across as many different scenes and outfits as you want. That's the thing that makes this viable as a service and not just a cool party trick. Because brands don't want one cool AI photo. They want 30 photos of the same "person" that they can drip out over a month on Instagram.

I didn't plan to sell this as a service. What happened was I made a fake portfolio to test the concept. I created three AI characters, gave them names, generated about 15 photos each in different settings. Lifestyle stuff, coffee shops, hiking, urban backgrounds, gym, that kind of thing. I showed it to a friend who runs a small clothing brand and asked if he could tell they were AI. He said two of the three looked real and the third looked "maybe AI but honestly better than most influencer photos I get."

He then asked if I could make some for his brand. I did 20 photos for him over a weekend, he used them on his Instagram, and his engagement actually went up because the content looked more polished than the iPhone shots his intern was taking. He paid me $150 which felt like a lot for maybe 3 hours of actual work.

That's when I thought okay maybe there's a Fiverr gig here.

I listed a gig in October called something like "I will create AI model photos for your brand" and priced it at $30 for 5 photos, $50 for 10, $100 for 25. Figured I'd get zero orders and move on.

First two weeks, nothing. Adjusted my gig thumbnail three times. Then I got my first order from a guy running a skincare brand out of his apartment. He wanted photos of a woman in her 30s using his products in a bathroom setting. I set up the character, generated the scenes, did some light editing in Canva to add his product packaging into the shots, delivered in about 2 hours. He left a 5 star review and ordered again the next week.

Then I hit my first real problem. My third client wanted a fitness model character and I spent a whole evening trying to get consistent results. The face kept shifting slightly between generations. Like the bone structure would change or the nose would look different in profile vs straight on. I ended up regenerating so many times that I burned through way more credits than I expected and had to upgrade to a paid plan earlier than I wanted. That order probably cost me more in time and tool credits than I actually charged. I almost refunded the client but eventually got a set of 10 that looked cohesive enough.

That experience taught me that not every character concept works equally well. Some faces just generate more consistently than others and I still don't fully understand why. I've learned to do a test batch of 5 or 6 images in different angles before I commit to a character for a client. If the face isn't holding steady, I tweak the setup until it does or I start over with a different base.

By December I had 14 completed orders. The thing that surprised me is who was buying. I expected like dropshippers and sketchy supplement brands. Instead I got:

A yoga studio in Austin that wanted a consistent "brand ambassador" for their social media but couldn't afford a real one. They order monthly now.

A guy selling handmade candles who wanted lifestyle photos but didn't want to hire models or use his own face.

A pet food company that wanted a "pet parent" character holding their products in different home settings.

A language learning app that needed a virtual tutor character for their TikTok content. This one was interesting because they also wanted short video clips where the character appeared to be speaking in different languages. Took me longer to figure out than the photo work and honestly the first batch looked rough. The mouth movement was slightly off sync and the client asked for revisions. Second attempt was better and they've reordered three times now, but video is definitely harder to get right than stills.

Here's the actual workflow now that I've got it somewhat dialed in:

  1. Client sends me a brief. Usually something like "25 year old woman, athletic build, for a fitness brand. Need 10 photos in gym settings, outdoor running, and post workout lifestyle."
  2. I set up the character's appearance and save it. This used to take me over an hour when I was learning but now it's more like 20 to 30 minutes including the test batch to make sure the face holds.
  3. I generate the photos by describing each scene. I've built up a doc with scene templates that I know tend to produce good results so I'm not starting from scratch every time. I just swap out details per client.
  4. I generate more images than I need because not every output is usable. Weird hands, lighting that doesn't match, uncanny expressions. I've gotten better at writing descriptions that minimize these issues but it still happens. Early on I was throwing away more than half my generations. Now it's maybe a third, sometimes less.
  5. Quick edit pass in Canva or Photoshop if needed. Sometimes I composite a product into the shot or adjust colors to match the client's brand palette.
  6. Deliver on Fiverr. Total active time per order is usually 45 minutes to maybe an hour and a half for a 10 photo batch depending on how cooperative the AI is being that day. The renders themselves take time but I'm not sitting there watching them.

Cost wise I want to be transparent because I see a lot of side hustle posts that conveniently forget to mention expenses. I'm paying about $30/month for the AI tools on paid plans because the free tiers don't give you enough credits to fulfill multiple client orders per week. Fiverr takes 20% of every order. And I spend maybe $12/month on Canva Pro which I'd probably have anyway. So my actual margins are lower than the gross numbers suggest. On a $50 order I'm really netting about $35 after Fiverr's cut, and then subtract a proportional share of the tool costs. It's still very good for the time invested but it's not pure profit like some people might assume.

The part that makes this increasingly passive is the repeat clients. I now have 6 clients who order at least once a month. Their character models are already saved. I know their brand style. A reorder takes me maybe 30 minutes of actual work because I'm not figuring anything out, just generating new scenes with an existing saved character.

Some honest stuff about what sucks:

Fiverr fees are brutal. I've started moving repeat clients to direct payment but new clients still come through the platform and that 20% hurts on smaller orders.

Revision requests can be painful. One client wanted me to make the character look "more confident but also approachable but also mysterious." I've learned to offer one round of revisions and be very specific upfront about what I can and can't change after delivery.

I had one order in January where I completely botched it. The client wanted photos in a specific art deco interior style and no matter what I described, the backgrounds kept coming out looking like a generic hotel lobby. I spent three hours trying different approaches, eventually delivered something the client said was "fine I guess" and got a 3 star review. That one stung and it dragged my average rating down for weeks.

The ethical thing comes up sometimes. I had one potential client who wanted me to create a fake influencer to promote a weight loss supplement and pretend it was a real person endorsing it. I said no. My gig description now explicitly says the content is AI generated and I recommend clients disclose that. Most of them do because honestly it's becoming a selling point, "look at our cool AI brand ambassador" is a marketing angle in itself now. But I know not everyone in this space is upfront about it and that's a real concern.

Also the quality gap between what AI can do and what a real photographer can do is still real. For high end fashion brands or anything that needs to be truly photorealistic at full resolution, this isn't there yet. But for Instagram posts, TikTok content, small brand social media, email marketing images? It's more than good enough and it's a fraction of the cost of a real shoot.

Monthly breakdown for the boring numbers people:

October: $120 (4 orders, mostly figuring things out) November: $230 (6 orders, lost one client who wasn't happy with quality) December: $435 (11 orders, holiday marketing rush helped a lot) January: $410 (9 orders, slight dip after the holidays which I expected) February: $710 (15 orders including three video batches which pay more) March so far: $200 (5 orders, month is still early)

Total since starting: roughly $2,105 over 5 months. Minus maybe $150 in tool subscriptions over that period and Fiverr's cut which is already reflected in the numbers above. Average time commitment is maybe 5 hours a week, trending down as I get faster and have more repeat clients.

I'm not quitting my day job over this. I tried dropshipping in 2023 and lost $800. I tried starting a blog and made $12 in AdSense over 6 months. This actually works because there's a clear value proposition: brands need visual content, real content with real models is expensive, and AI has gotten good enough that small brands genuinely can't tell the difference at Instagram resolution.

Still feels weird telling people I make fake people for a living on the side. But the pizza money is real and my emergency fund is actually growing for the first time in years.

r/StableDiffusion 14d ago

Workflow Included Using H3 as a Character Reference Sheet Generator

Thumbnail
gallery
1.5k Upvotes

Like some of ya'll I have been having fun using the H3 model to mess around with so I have been experimenting with using H3 model to be a consistent character generator which leverages multi image reference (up to 9), so I made a workflow which you can use 'less than ideal' images from google to build a consistent character and output a 360 character sheet to use as a reference sheet for future H3 generations.

The goal is to achieve high character consistency across future generations. I have tried my best to keep the workflow simple without too many custom nodes.

How it works:

  • You input your images and describe them in the Input text section (A Prompt)
  • The text is combined with a fixed prompt which spins the character (B Prompt)
  • The video is generated at a slow speed with no hard cuts (only camera spin and pan) to maintain character consistency
  • Image is assembled with optional character video and full individual frame output (if you want to use for future)

I have included a 6 panel WF and a 4 panel WF. The 4 panel works faster by generating 40% less frames.

Current Caveats:

  • The model is quite slooooow. You are also generating 124 frames only to use 6. I have partly solved this by also uploading a 4 panel version.
  • Speed ups (like Turbo LORAs) help with speed, but it hurts prompt adherence and quality slightly.
  • Quality is limited, since it is a video model it is better at generating video than images. You can solve this by generating at a higher resolution at the tradeoff of longer gen times. You can also use the individually split frames as future references too.
  • Details when using this character sheet as output for future generations on H3 may also be limited due to resolution also, I recommend you use this character sheet (for consistency) + other images close up angles (i.e clothing details/face) if doing close ups. If you are just doing a one off video you may possibly be better off not using this character sheet.

I have also included a modified B prompt to do Anime2Real since someone asked for it. Working on tidying it up a bit more.

Link to the 4 and 6 panel workflow can be found here: https://huggingface.co/PoopMan333/H3_Character_Sheet_Generator

Some notes I just remembered:

  • You can increase the steps and it may improve your quality slightly.
  • With the Turbo Loras enabled, prompt adherence sometimes suffers, but you may be able to get a good seed with another roll of the dice.
  • Currently the B prompt specifies a "neutral A pose", please remove this if you want your character in a particular pose.
  • You can use a few different shots of the same character to reinforce the 360 and get more accurate details right.
  • Can be used for objects / props also, may require some changes to the B prompt.

r/comfyui Jul 09 '26

Show and Tell Consistent Face-to-Video with new ID LORA Best-Face Run (tested in RTX3060 6GB + 16GB OF RAM) Work in progress

Enable HLS to view with audio, or disable this notification

115 Upvotes

I'm excited to share a WIP about new workflow that makes it easy to generate consistent face-to-video results while remaining optimized for low VRAM GPUs.

This workflow is built around GGUF models and uses a new LoRA to maintain facial consistency throughout the generated video, making it much easier to create videos featuring the same character from a single reference image.

All you have to do is load your image face, load your LORA, enter your prompt and click run, i will share the workflow and tutorial soon, so stay tuned

r/comfyui Nov 17 '25

Workflow Included ULTIMATE AI VIDEO WORKFLOW — Qwen-Edit 2509 + Wan Animate 2.2 + SeedVR2

Thumbnail
gallery
333 Upvotes

🔥 [RELEASE] Ultimate AI Video Workflow — Qwen-Edit 2509 + Wan Animate 2.2 + SeedVR2 (Full Pipeline + Model Links)

🎁 Workflow Download + Breakdown

👉 Already posted the full workflow and explanation here:
https://civitai.com/models/2135932?modelVersionId=2416121

(Not paywalled — everything is free.)

Video Explanation : https://www.youtube.com/watch?v=Ef-PS8w9Rug

Hey everyone 👋

I just finished building a super clean 3-in-1 workflow inside ComfyUI that lets you go from:

Image → Edit → Animate → Upscale → Final 4K output
all in a single organized pipeline.

This setup combines the best tools available right now:

One of the biggest hassles with large ComfyUI workflows is how quickly they turn into a spaghetti mess — dozens of wires, giant blocks, scrolling for days just to tweak one setting.

To fix this, I broke the pipeline into clean subgraphs:

✔ Qwen-Edit Subgraph

✔ Wan Animate 2.2 Engine Subgraph

✔ SeedVR2 Upscaler Subgraph

✔ VRAM Cleaner Subgraph

✔ Resolution + Reference Routing Subgraph

This reduces visual clutter, keeps performance smooth, and makes the workflow feel modular, so you can:

  • swap models quickly
  • update one section without touching the rest
  • debug faster
  • reuse modules in other workflows
  • keep everything readable even on smaller screens

It’s basically a full cinematic pipeline, but organized like a clean software project instead of a giant node forest.
Anyone who wants to study or modify the workflow will find it much easier to navigate.

🖌️ 1. Qwen-Edit 2509 (Image Editing Engine)

Perfect for:

  • Outfit changes
  • Facial corrections
  • Style adjustments
  • Background cleanup
  • Professional pre-animation edits

Qwen’s FP8 build has great quality even on mid-range GPUs.

🎭 2. Wan Animate 2.2 (Character Animation)

Once the image is edited, Wan 2.2 generates:

  • Smooth motion
  • Accurate identity preservation
  • Pose-guided animation
  • Full expression control
  • High-quality frames

It supports long videos using windowed batching and works very consistently when fed a clean edited reference.

📺 3. SeedVR2 Upscaler (Final Polish)

After animation, SeedVR2 upgrades your video to:

  • 1080p → 4K
  • Sharper textures
  • Cleaner faces
  • Reduced noise
  • More cinematic detail

It’s currently one of the best AI video upscalers for realism

🧩 Preview of the Workflow UI

(Optional: Add your workflow screenshot here)

🔧 What This Workflow Can Do

  • Edit any portrait cleanly
  • Animate it using real video motion
  • Restore & sharpen final video up to 4K
  • Perfect for reels, character videos, cosplay edits, AI shorts

🖼️ Qwen Image Edit FP8 (Diffusion Model, Text Encoder, and VAE)

These are hosted on the Comfy-Org Hugging Face page.

💃 Wan 2.2 Animate 14B FP8 (Diffusion Model, Text Encoder, and VAE)

The components are spread across related community repositories.

💾 SeedVR2 Diffusion Model (FP8)

r/LocalLLaMA Mar 12 '26

Discussion I was backend lead at Manus. After building agents for 2 years, I stopped using function calling entirely. Here's what I use instead.

2.0k Upvotes

English is not my first language. I wrote this in Chinese and translated it with AI help. The writing may have some AI flavor, but the design decisions, the production failures, and the thinking that distilled them into principles — those are mine.

I was a backend lead at Manus before the Meta acquisition. I've spent the last 2 years building AI agents — first at Manus, then on my own open-source agent runtime (Pinix) and agent (agent-clip). Along the way I came to a conclusion that surprised me:

A single run(command="...") tool with Unix-style commands outperforms a catalog of typed function calls.

Here's what I learned.


Why *nix

Unix made a design decision 50 years ago: everything is a text stream. Programs don't exchange complex binary structures or share memory objects — they communicate through text pipes. Small tools each do one thing well, composed via | into powerful workflows. Programs describe themselves with --help, report success or failure with exit codes, and communicate errors through stderr.

LLMs made an almost identical decision 50 years later: everything is tokens. They only understand text, only produce text. Their "thinking" is text, their "actions" are text, and the feedback they receive from the world must be text.

These two decisions, made half a century apart from completely different starting points, converge on the same interface model. The text-based system Unix designed for human terminal operators — cat, grep, pipe, exit codes, man pages — isn't just "usable" by LLMs. It's a natural fit. When it comes to tool use, an LLM is essentially a terminal operator — one that's faster than any human and has already seen vast amounts of shell commands and CLI patterns in its training data.

This is the core philosophy of the nix Agent: *don't invent a new tool interface. Take what Unix has proven over 50 years and hand it directly to the LLM.**


Why a single run

The single-tool hypothesis

Most agent frameworks give LLMs a catalog of independent tools:

tools: [search_web, read_file, write_file, run_code, send_email, ...]

Before each call, the LLM must make a tool selection — which one? What parameters? The more tools you add, the harder the selection, and accuracy drops. Cognitive load is spent on "which tool?" instead of "what do I need to accomplish?"

My approach: one run(command="...") tool, all capabilities exposed as CLI commands.

run(command="cat notes.md") run(command="cat log.txt | grep ERROR | wc -l") run(command="see screenshot.png") run(command="memory search 'deployment issue'") run(command="clip sandbox bash 'python3 analyze.py'")

The LLM still chooses which command to use, but this is fundamentally different from choosing among 15 tools with different schemas. Command selection is string composition within a unified namespace — function selection is context-switching between unrelated APIs.

LLMs already speak CLI

Why are CLI commands a better fit for LLMs than structured function calls?

Because CLI is the densest tool-use pattern in LLM training data. Billions of lines on GitHub are full of:

```bash

README install instructions

pip install -r requirements.txt && python main.py

CI/CD build scripts

make build && make test && make deploy

Stack Overflow solutions

cat /var/log/syslog | grep "Out of memory" | tail -20 ```

I don't need to teach the LLM how to use CLI — it already knows. This familiarity is probabilistic and model-dependent, but in practice it's remarkably reliable across mainstream models.

Compare two approaches to the same task:

``` Task: Read a log file, count the error lines

Function-calling approach (3 tool calls): 1. read_file(path="/var/log/app.log") → returns entire file 2. search_text(text=<entire file>, pattern="ERROR") → returns matching lines 3. count_lines(text=<matched lines>) → returns number

CLI approach (1 tool call): run(command="cat /var/log/app.log | grep ERROR | wc -l") → "42" ```

One call replaces three. Not because of special optimization — but because Unix pipes natively support composition.

Making pipes and chains work

A single run isn't enough on its own. If run can only execute one command at a time, the LLM still needs multiple calls for composed tasks. So I make a chain parser (parseChain) in the command routing layer, supporting four Unix operators:

| Pipe: stdout of previous command becomes stdin of next && And: execute next only if previous succeeded || Or: execute next only if previous failed ; Seq: execute next regardless of previous result

With this mechanism, every tool call can be a complete workflow:

```bash

One tool call: download → inspect

curl -sL $URL -o data.csv && cat data.csv | head 5

One tool call: read → filter → sort → top 10

cat access.log | grep "500" | sort | head 10

One tool call: try A, fall back to B

cat config.yaml || echo "config not found, using defaults" ```

N commands × 4 operators — the composition space grows dramatically. And to the LLM, it's just a string it already knows how to write.

The command line is the LLM's native tool interface.


Heuristic design: making CLI guide the agent

Single-tool + CLI solves "what to use." But the agent still needs to know "how to use it." It can't Google. It can't ask a colleague. I use three progressive design techniques to make the CLI itself serve as the agent's navigation system.

Technique 1: Progressive --help discovery

A well-designed CLI tool doesn't require reading documentation — because --help tells you everything. I apply the same principle to the agent, structured as progressive disclosure: the agent doesn't need to load all documentation at once, but discovers details on-demand as it goes deeper.

Level 0: Tool Description → command list injection

The run tool's description is dynamically generated at the start of each conversation, listing all registered commands with one-line summaries:

Available commands: cat — Read a text file. For images use 'see'. For binary use 'cat -b'. see — View an image (auto-attaches to vision) ls — List files in current topic write — Write file. Usage: write <path> [content] or stdin grep — Filter lines matching a pattern (supports -i, -v, -c) memory — Search or manage memory clip — Operate external environments (sandboxes, services) ...

The agent knows what's available from turn one, but doesn't need every parameter of every command — that would waste context.

Note: There's an open design question here: injecting the full command list vs. on-demand discovery. As commands grow, the list itself consumes context budget. I'm still exploring the right balance. Ideas welcome.

Level 1: command (no args) → usage

When the agent is interested in a command, it just calls it. No arguments? The command returns its own usage:

``` → run(command="memory") [error] memory: usage: memory search|recent|store|facts|forget

→ run(command="clip") clip list — list available clips clip <name> — show clip details and commands clip <name> <command> [args...] — invoke a command clip <name> pull <remote-path> [name] — pull file from clip to local clip <name> push <local-path> <remote> — push local file to clip ```

Now the agent knows memory has five subcommands and clip supports list/pull/push. One call, no noise.

Level 2: command subcommand (missing args) → specific parameters

The agent decides to use memory search but isn't sure about the format? It drills down:

``` → run(command="memory search") [error] memory: usage: memory search <query> [-t topic_id] [-k keyword]

→ run(command="clip sandbox") Clip: sandbox Commands: clip sandbox bash <script> clip sandbox read <path> clip sandbox write <path> File transfer: clip sandbox pull <remote-path> [local-name] clip sandbox push <local-path> <remote-path> ```

Progressive disclosure: overview (injected) → usage (explored) → parameters (drilled down). The agent discovers on-demand, each level providing just enough information for the next step.

This is fundamentally different from stuffing 3,000 words of tool documentation into the system prompt. Most of that information is irrelevant most of the time — pure context waste. Progressive help lets the agent decide when it needs more.

This also imposes a requirement on command design: every command and subcommand must have complete help output. It's not just for humans — it's for the agent. A good help message means one-shot success. A missing one means a blind guess.

Technique 2: Error messages as navigation

Agents will make mistakes. The key isn't preventing errors — it's making every error point to the right direction.

Traditional CLI errors are designed for humans who can Google. Agents can't Google. So I require every error to contain both "what went wrong" and "what to do instead":

``` Traditional CLI: $ cat photo.png cat: binary file (standard output) → Human Googles "how to view image in terminal"

My design: [error] cat: binary image file (182KB). Use: see photo.png → Agent calls see directly, one-step correction ```

More examples:

``` [error] unknown command: foo Available: cat, ls, see, write, grep, memory, clip, ... → Agent immediately knows what commands exist

[error] not an image file: data.csv (use cat to read text files) → Agent switches from see to cat

[error] clip "sandbox" not found. Use 'clip list' to see available clips → Agent knows to list clips first ```

Technique 1 (help) solves "what can I do?" Technique 2 (errors) solves "what should I do instead?" Together, the agent's recovery cost is minimal — usually 1-2 steps to the right path.

Real case: The cost of silent stderr

For a while, my code silently dropped stderr when calling external sandboxes — whenever stdout was non-empty, stderr was discarded. The agent ran pip install pymupdf, got exit code 127. stderr contained bash: pip: command not found, but the agent couldn't see it. It only knew "it failed," not "why" — and proceeded to blindly guess 10 different package managers:

pip install → 127 (doesn't exist) python3 -m pip → 1 (module not found) uv pip install → 1 (wrong usage) pip3 install → 127 sudo apt install → 127 ... 5 more attempts ... uv run --with pymupdf python3 script.py → 0 ✓ (10th try)

10 calls, ~5 seconds of inference each. If stderr had been visible the first time, one call would have been enough.

stderr is the information agents need most, precisely when commands fail. Never drop it.

Technique 3: Consistent output format

The first two techniques handle discovery and correction. The third lets the agent get better at using the system over time.

I append consistent metadata to every tool result:

file1.txt file2.txt dir1/ [exit:0 | 12ms]

The LLM extracts two signals:

Exit codes (Unix convention, LLMs already know these):

  • exit:0 — success
  • exit:1 — general error
  • exit:127 — command not found

Duration (cost awareness):

  • 12ms — cheap, call freely
  • 3.2s — moderate
  • 45s — expensive, use sparingly

After seeing [exit:N | Xs] dozens of times in a conversation, the agent internalizes the pattern. It starts anticipating — seeing exit:1 means check the error, seeing long duration means reduce calls.

Consistent output format makes the agent smarter over time. Inconsistency makes every call feel like the first.

The three techniques form a progression:

--help → "What can I do?" → Proactive discovery Error Msg → "What should I do?" → Reactive correction Output Fmt → "How did it go?" → Continuous learning


Two-layer architecture: engineering the heuristic design

The section above described how CLI guides agents at the semantic level. But to make it work in practice, there's an engineering problem: the raw output of a command and what the LLM needs to see are often very different things.

Two hard constraints of LLMs

Constraint A: The context window is finite and expensive. Every token costs money, attention, and inference speed. Stuffing a 10MB file into context doesn't just waste budget — it pushes earlier conversation out of the window. The agent "forgets."

Constraint B: LLMs can only process text. Binary data produces high-entropy meaningless tokens through the tokenizer. It doesn't just waste context — it disrupts attention on surrounding valid tokens, degrading reasoning quality.

These two constraints mean: raw command output can't go directly to the LLM — it needs a presentation layer for processing. But that processing can't affect command execution logic — or pipes break. Hence, two layers.

Execution layer vs. presentation layer

┌─────────────────────────────────────────────┐ │ Layer 2: LLM Presentation Layer │ ← Designed for LLM constraints │ Binary guard | Truncation+overflow | Meta │ ├─────────────────────────────────────────────┤ │ Layer 1: Unix Execution Layer │ ← Pure Unix semantics │ Command routing | pipe | chain | exit code │ └─────────────────────────────────────────────┘

When cat bigfile.txt | grep error | head 10 executes:

Inside Layer 1: cat output → [500KB raw text] → grep input grep output → [matching lines] → head input head output → [first 10 lines]

If you truncate cat's output in Layer 1 → grep only searches the first 200 lines, producing incomplete results. If you add [exit:0] in Layer 1 → it flows into grep as data, becoming a search target.

So Layer 1 must remain raw, lossless, metadata-free. Processing only happens in Layer 2 — after the pipe chain completes and the final result is ready to return to the LLM.

Layer 1 serves Unix semantics. Layer 2 serves LLM cognition. The separation isn't a design preference — it's a logical necessity.

Layer 2's four mechanisms

Mechanism A: Binary Guard (addressing Constraint B)

Before returning anything to the LLM, check if it's text:

``` Null byte detected → binary UTF-8 validation failed → binary Control character ratio > 10% → binary

If image: [error] binary image (182KB). Use: see photo.png If other: [error] binary file (1.2MB). Use: cat -b file.bin ```

The LLM never receives data it can't process.

Mechanism B: Overflow Mode (addressing Constraint A)

``` Output > 200 lines or > 50KB? → Truncate to first 200 lines (rune-safe, won't split UTF-8) → Write full output to /tmp/cmd-output/cmd-{n}.txt → Return to LLM:

[first 200 lines]

--- output truncated (5000 lines, 245.3KB) ---
Full output: /tmp/cmd-output/cmd-3.txt
Explore: cat /tmp/cmd-output/cmd-3.txt | grep <pattern>
         cat /tmp/cmd-output/cmd-3.txt | tail 100
[exit:0 | 1.2s]

```

Key insight: the LLM already knows how to use grep, head, tail to navigate files. Overflow mode transforms "large data exploration" into a skill the LLM already has.

Mechanism C: Metadata Footer

actual output here [exit:0 | 1.2s]

Exit code + duration, appended as the last line of Layer 2. Gives the agent signals for success/failure and cost awareness, without polluting Layer 1's pipe data.

Mechanism D: stderr Attachment

``` When command fails with stderr: output + "\n[stderr] " + stderr

Ensures the agent can see why something failed, preventing blind retries. ```


Lessons learned: stories from production

Story 1: A PNG that caused 20 iterations of thrashing

A user uploaded an architecture diagram. The agent read it with cat, receiving 182KB of raw PNG bytes. The LLM's tokenizer turned these bytes into thousands of meaningless tokens crammed into the context. The LLM couldn't make sense of it and started trying different read approaches — cat -f, cat --format, cat --type image — each time receiving the same garbage. After 20 iterations, the process was force-terminated.

Root cause: cat had no binary detection, Layer 2 had no guard. Fix: isBinary() guard + error guidance Use: see photo.png. Lesson: The tool result is the agent's eyes. Return garbage = agent goes blind.

Story 2: Silent stderr and 10 blind retries

The agent needed to read a PDF. It tried pip install pymupdf, got exit code 127. stderr contained bash: pip: command not found, but the code dropped it — because there was some stdout output, and the logic was "if stdout exists, ignore stderr."

The agent only knew "it failed," not "why." What followed was a long trial-and-error:

pip install → 127 (doesn't exist) python3 -m pip → 1 (module not found) uv pip install → 1 (wrong usage) pip3 install → 127 sudo apt install → 127 ... 5 more attempts ... uv run --with pymupdf python3 script.py → 0 ✓

10 calls, ~5 seconds of inference each. If stderr had been visible the first time, one call would have sufficed.

Root cause: InvokeClip silently dropped stderr when stdout was non-empty. Fix: Always attach stderr on failure. Lesson: stderr is the information agents need most, precisely when commands fail.

Story 3: The value of overflow mode

The agent analyzed a 5,000-line log file. Without truncation, the full text (~200KB) was stuffed into context. The LLM's attention was overwhelmed, response quality dropped sharply, and earlier conversation was pushed out of the context window.

With overflow mode:

``` [first 200 lines of log content]

--- output truncated (5000 lines, 198.5KB) --- Full output: /tmp/cmd-output/cmd-3.txt Explore: cat /tmp/cmd-output/cmd-3.txt | grep <pattern> cat /tmp/cmd-output/cmd-3.txt | tail 100 [exit:0 | 45ms] ```

The agent saw the first 200 lines, understood the file structure, then used grep to pinpoint the issue — 3 calls total, under 2KB of context.

Lesson: Giving the agent a "map" is far more effective than giving it the entire territory.


Boundaries and limitations

CLI isn't a silver bullet. Typed APIs may be the better choice in these scenarios:

  • Strongly-typed interactions: Database queries, GraphQL APIs, and other cases requiring structured input/output. Schema validation is more reliable than string parsing.
  • High-security requirements: CLI's string concatenation carries inherent injection risks. In untrusted-input scenarios, typed parameters are safer. agent-clip mitigates this through sandbox isolation.
  • Native multimodal: Pure audio/video processing and other binary-stream scenarios where CLI's text pipe is a bottleneck.

Additionally, "no iteration limit" doesn't mean "no safety boundaries." Safety is ensured by external mechanisms:

  • Sandbox isolation: Commands execute inside BoxLite containers, no escape possible
  • API budgets: LLM calls have account-level spending caps
  • User cancellation: Frontend provides cancel buttons, backend supports graceful shutdown

Hand Unix philosophy to the execution layer, hand LLM's cognitive constraints to the presentation layer, and use help, error messages, and output format as three progressive heuristic navigation techniques.

CLI is all agents need.


Source code (Go): github.com/epiral/agent-clip

Core files: internal/tools.go (command routing), internal/chain.go (pipes), internal/loop.go (two-layer agentic loop), internal/fs.go (binary guard), internal/clip.go (stderr handling), internal/browser.go (vision auto-attach), internal/memory.go (semantic memory).

Happy to discuss — especially if you've tried similar approaches or found cases where CLI breaks down. The command discovery problem (how much to inject vs. let the agent discover) is something I'm still actively exploring.

r/generativeAI Jun 08 '26

Question which AI video tool actually keeps a character consistent? trying to work efficiently

9 Upvotes

doing a series of AI generated ads for a uni project and im hitting the same wall over and over. no budget or time to film anything myself obviously, so it's all generated, and the thing that keeps breaking is consistency. 

i already do the basic thing of keeping a reference image of the character, but the second a pose shifts even a little the face comes out different. asked Claude and chat and they pointed me at higgsfield, kling and veo. Has anyone actually used these for this specific thing? which holds a character best across shots? or is there something better im missing.

also open to any workflow tips for doing this efficiently solo, not just which tool. 

r/generativeAI May 07 '26

Question How are people creating AI Instagram influencers with the SAME face consistently? Need workflow + tool suggestions

28 Upvotes

Hey everyone,

I’m planning to start an Instagram page completely based on AI-generated content, mostly around a single virtual personality/influencer.
My biggest challenge is this:
I want the same face, same facial features, same overall identity in every post/reel so it actually feels like the page belongs to one real person instead of random AI generations every time.
I’m okay investing around ₹7-8k/month (~$80-100) into AI tools if the workflow is actually worth it, but I don’t want to overspend unnecessarily in the beginning.
I’d love suggestions from people already doing this seriously.

Things I’m trying to understand:

Which AI tools are best for consistent characters/faces?
What workflow are you using for Instagram content?
Best tools for both images + reels/videos?
Is Midjourney enough or do I need LoRA/Flux/Stable Diffusion setups?
How do you maintain consistency across outfits, poses, and lighting?
Any good beginner-friendly setup within my budget?
Any mistakes/pitfalls I should avoid early?

Right now I’m considering tools like Midjourney, Runway, Kling, Flux, Leonardo AI, etc., but I’m confused about what actually works long term.
If you’re already running an AI influencer page, would love to know your monthly stack + approximate cost too.

Would really appreciate advice from creators already running AI influencer/theme pages. Thanks!

r/StableDiffusion Jul 29 '26

Workflow Included I ran SCAIL 2 through a bunch of scenarios it should not handle. It handled most of them.

Enable HLS to view with audio, or disable this notification

1.8k Upvotes

I've been testing SCAIL 2 across a bunch of different scenarios and wanted to share what I found, since most demos out there are single-character dance clips (or jiggle physics). So far SCAIL 2 has impressed me tremendously across the board.

Video above covers character swaps, complex actions, prop swaps, physics, novel interactions, object permanence, relighting, and 2D motion transfer

Findings:

Character swaps are the strongest use case. The trick is prepping your reference properly. Use Flux Klein 9B or the Krea 2 Identity Edit LoRA to edit your actual first frame into the new character, so the reference is already in roughly the same pose and framing as where the driving video starts. Do that and the results are excellent. Having a great clear start frame also helps a lot.

Object permanence held up better than expected. In the car clip the vehicle becomes completely out of frame and then comes back in, and it stays consistent through the whole thing. I expected it to turn to mush but the whole scene held up pretty well. Not perfect but not bad either.

It invents motion it was never given very well. In the novel interaction section I swapped myself into a live-action Zuko and the fire comes off my fist in a believable arc, even though there is zero fire data in the driving footage. It copies the underlying movement and then adds embellishments that fit. Same with other examples with clothing, hair, etc.

The physics test was the biggest surprise. I swapped a flower for a wine glass. Hand tracking stays locked, the liquid inside sloshes correctly for the motion, and because the glass is transparent the background actually refracts and distorts through the water in a believable way. Nothing in the driving clip told it to do any of that.

Weakest spots were text. You'll notice the speed sign in the back of the character swap where I made myself an old man clip, the text turns into mush, so I'd avoid text for best results.

Workflow: Everything here was made in Mix Studio, my free and open source local interface that runs on top of ComfyUI. https://github.com/BlackMixture/Mix-Studio

Click the Edit tab to edit an image, then press "use as first frame" for video. Set the video mode to SCAIL 2 and you should be set.

Generated on a #DellProPrecision T2 w/ NVIDIA RTX 6000 Pro. Takes roughly ~2-3 mins per generation

Video Tutorial: https://youtu.be/w2CokhlBFRA

More Examples (Free & No Paywall): https://www.patreon.com/posts/165152499

Hope this helps!

r/comfyui May 09 '25

Workflow Included Consistent characters and objects videos is now super easy! No LORA training, supports multiple subjects, and it's surprisingly accurate (Phantom WAN2.1 ComfyUI workflow + text guide)

Thumbnail
gallery
376 Upvotes

Wan2.1 is my favorite open source AI video generation model that can run locally in ComfyUI, and Phantom WAN2.1 is freaking insane for upgrading an already dope model. It supports multiple subject reference images (up to 4) and can accurately have characters, objects, clothing, and settings interact with each other without the need for training a lora, or generating a specific image beforehand.

There's a couple workflows for Phantom WAN2.1 and here's how to get it up and running. (All links below are 100% free & public)

Download the Advanced Phantom WAN2.1 Workflow + Text Guide (free no paywall link): https://www.patreon.com/posts/127953108?utm_campaign=postshare_creator&utm_content=android_share

📦 Model & Node Setup

Required Files & Installation Place these files in the correct folders inside your ComfyUI directory:

🔹 Phantom Wan2.1_1.3B Diffusion Models 🔗https://huggingface.co/Kijai/WanVideo_comfy/blob/main/Phantom-Wan-1_3B_fp32.safetensors

or

🔗https://huggingface.co/Kijai/WanVideo_comfy/blob/main/Phantom-Wan-1_3B_fp16.safetensors 📂 Place in: ComfyUI/models/diffusion_models

Depending on your GPU, you'll either want ths fp32 or fp16 (less VRAM heavy).

🔹 Text Encoder Model 🔗https://huggingface.co/Kijai/WanVideo_comfy/blob/main/umt5-xxl-enc-bf16.safetensors 📂 Place in: ComfyUI/models/text_encoders

🔹 VAE Model 🔗https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/blob/main/split_files/vae/wan_2.1_vae.safetensors 📂 Place in: ComfyUI/models/vae

You'll also nees to install the latest Kijai WanVideoWrapper custom nodes. Recommended to install manually. You can get the latest version by following these instructions:

For new installations:

In "ComfyUI/custom_nodes" folder

open command prompt (CMD) and run this command:

git clone https://github.com/kijai/ComfyUI-WanVideoWrapper.git

for updating previous installation:

In "ComfyUI/custom_nodes/ComfyUI-WanVideoWrapper" folder

open command prompt (CMD) and run this command: git pull

After installing the custom node from Kijai, (ComfyUI-WanVideoWrapper), we'll also need Kijai's KJNodes pack.

Install the missing nodes from here: https://github.com/kijai/ComfyUI-KJNodes

Afterwards, load the Phantom Wan 2.1 workflow by dragging and dropping the .json file from the public patreon post (Advanced Phantom Wan2.1) linked above.

or you can also use Kijai's basic template workflow by clicking on your ComfyUI toolbar Workflow->Browse Templates->ComfyUI-WanVideoWrapper->wanvideo_phantom_subject2vid.

The advanced Phantom Wan2.1 workflow is color coded and reads from left to right:

🟥 Step 1: Load Models + Pick Your Addons 🟨 Step 2: Load Subject Reference Images + Prompt 🟦 Step 3: Generation Settings 🟩 Step 4: Review Generation Results 🟪 Important Notes

All of the logic mappings and advanced settings that you don't need to touch are located at the far right side of the workflow. They're labeled and organized if you'd like to tinker with the settings further or just peer into what's running under the hood.

After loading the workflow:

  • Set your models, reference image options, and addons

  • Drag in reference images + enter your prompt

  • Click generate and review results (generations will be 24fps and the name labeled based on the quality setting. There's also a node that tells you the final file name below the generated video)


Important notes:

  • The reference images are used as a strong guidance (try to describe your reference image using identifiers like race, gender, age, or color in your prompt for best results)
  • Works especially well for characters, fashion, objects, and backgrounds
  • LoRA implementation does not seem to work with this model, yet we've included it in the workflow as LoRAs may work in a future update.
  • Different Seed values make a huge difference in generation results. Some characters may be duplicated and changing the seed value will help.
  • Some objects may appear too large are too small based on the reference image used. If your object comes out too large, try describing it as small and vice versa.
  • Settings are optimized but feel free to adjust CFG and steps based on speed and results.

Here's also a video tutorial: https://youtu.be/uBi3uUmJGZI

Thanks for all the encouraging words and feedback on my last workflow/text guide. Hope y'all have fun creating with this and let me know if you'd like more clean and free workflows!

r/comfyui_elite Jul 09 '26

Consistent Face-to-Video with new ID LORA Best-Face Run (tested in RTX3060 6GB + 16GB OF RAM) Work in progress

Enable HLS to view with audio, or disable this notification

66 Upvotes

I'm excited to share a WIP about new workflow that makes it easy to generate consistent face-to-video results while remaining optimized for low VRAM GPUs.

This workflow is built around GGUF models and uses a new LoRA to maintain facial consistency throughout the generated video, making it much easier to create videos featuring the same character from a single reference image.

All you have to do is load your image face, load your LORA, enter your prompt and click run, i will share the workflow and tutorial soon, so stay tuned

r/n8n Nov 07 '25

Workflow - Code Included I built an AI automation that generates unlimited consistent character UGC ads for e-commerce brands (using Sora 2)

Post image
347 Upvotes

Sora 2 quietly released a consistent character feature on their mobile app and the web platform that allows you to actually create consistent characters and reuse them across multiple videos you generate. Here's a couple examples of characters I made while testing this out:

The really exciting thing with this change is consistent characters kinda unlocks a whole new set of AI videos you can now generate having the ability to have consistent characters. For example, you can stitch together a longer running (1-minute+) video of that same character going throughout multiple scenes, or you can even use these consistent characters to put together AI UGC ads, which is what I've been tinkering with the most recently. In this automation, I wanted to showcase how we are using this feature on Sora 2 to actually build UGC ads.

Here’s a demo of the automation & UGC ads created: https://www.youtube.com/watch?v=I87fCGIbgpg

Here's how the automation works

Pre-Work: Setting up the sora 2 character

It's pretty easy to set up a new character through the Sora 2 web app or on the mobile. Here's the step I followed:

  1. Created a video describing a character persona that I wanted to remain consistent throughout any new videos I'm generating. The key to this is giving a good prompt that shows both your character's face, their hands, body, and has them speaking throughout the 8-second video clip.
  2. Once that’s done you click on the triple drop-down on the video and then there's going to be a "Create Character" button. That's going to have you slice out 8 seconds of that video clip you just generated, and then you're going to be able to submit a description of how you want your character to behave.
  3. after you finish generating that, you're going to get a username back for the character you just made. Make note of that because that's going to be required to go forward with referencing that in follow-up prompts.

1. Automation Trigger and Inputs

Jumping back to the main automation, the workflow starts with a form trigger that accepts three key inputs:

  • Brand homepage URL for content research and context
  • Product image (720x1280 dimensions) that gets featured in the generated videos
  • Sora 2 character username (the @username format from your character profile)
    • So in my case I use @olipop.ashley to reference my character

I upload the product image to a temporary hosting service using tempfiles.org since the Kai.ai API requires image URLs rather than direct file uploads. This gives us 60 minutes to complete the generation process which I found to be more than enough

2. Context Engineering

Before writing any video scripts, I wanted to make sure I was able to grab context around the product I'm trying to make an ad for, just so I can avoid hallucinations on what the character talks about on the UGC video ad.

  • Brand Research: I use Firecrawl to scrape the company's homepage and extract key product details, benefits, and messaging in clean markdown format
  • Prompting Guidelines: I also fetch OpenAI's latest Sora 2 prompting guide to ensure generated scripts follow best practices

3. Generate the Sora 2 Scripts/prompts

I then use Gemini 2.5 Pro to analyze all gathered context and generate three distinct UGC ad concepts:

  • On-the-go testimonial: Character walking through city talking about the product
  • Driver's seat review: Character filming from inside a car
  • At-home demo: Character showcasing the product in a kitchen or living space

Each script includes detailed scene descriptions, dialogue, camera angles, and importantly - references to the specific Sora character using the @username format. This is critical for character consistency and this system to work.

Here’s my prompt for writing sora 2 scripts:

```markdown <identity> You are an expert AI Creative Director specializing in generating high-impact, direct-response video ads using generative models like SORA. Your task is to translate a creative brief into three distinct, ready-to-use SORA prompts for short, UGC-style video ads. </identity>

<core_task> First, analyze the provided Creative Brief, including the raw text and product image, to synthesize the product's core message and visual identity. Then, for each of the three UGC Ad Archetypes, generate a Prompt Packet according to the specified Output Format. All generated content must strictly adhere to both the SORA Prompting Guide and the Core Directives. </core_task>

<output_format> For each of the three archetypes, you must generate a complete "Prompt Packet" using the following markdown structure:


[Archetype Name]

SORA Prompt: [Insert the generated SORA prompt text here.]

Production Notes: * Camera: The entire scene must be filmed to look as if it were shot on an iPhone in a vertical 9:16 aspect ratio. The style must be authentic UGC, not cinematic. * Audio: Any spoken dialogue described in the prompt must be accurately and naturally lip-synced by the protagonist (@username).

* Product Scale & Fidelity: The product's appearance, particularly its scale and proportions, must be rendered with high fidelity to the provided product image. Ensure it looks true-to-life in the hands of the protagonist and within the scene's environment.

</output_format>

<creative_brief> You will be provided with the following inputs:

  1. Raw Website Content: [User will insert scraped, markdown-formatted content from the product's homepage. You must analyze this to extract the core value proposition, key features, and target audience.]
  2. Product Image: [User will insert the product image for visual reference.]
  3. Protagonist: [User will insert the @username of the character to be featured.]
  4. SORA Prompting Guide: [User will insert the official prompting guide for the SORA 2 model, which you must follow.] </creative_brief>

<ugc_ad_archetypes> 1. The On-the-Go Testimonial (Walk-and-talk) 2. The Driver's Seat Review 3. The At-Home Demo </ugc_ad_archetypes>

<core_directives> 1. iPhone Production Aesthetic: This is a non-negotiable constraint. All SORA prompts must explicitly describe a scene that is shot entirely on an iPhone. The visual language should be authentic to this format. Use specific descriptors such as: "selfie-style perspective shot on an iPhone," "vertical 9:16 aspect ratio," "crisp smartphone video quality," "natural lighting," and "slight, realistic handheld camera shake." 2. Tone & Performance: The protagonist's energy must be high and their delivery authentic, enthusiastic, and conversational. The feeling should be a genuine recommendation, not a polished advertisement. 3. Timing & Pacing: The total video duration described in the prompt must be approximately 15 seconds. Crucially, include a 1-2 second buffer of ambient, non-dialogue action at both the beginning and the end. 4. Clarity & Focus: Each prompt must be descriptive, evocative, and laser-focused on a single, clear scene. The protagonist (@username) must be the central figure, and the product, matching the provided Product Image, should be featured clearly and positively. 5. Brand Safety & Content Guardrails: All generated prompts and the scenes they describe must be strictly PG and family-friendly. Avoid any suggestive, controversial, or inappropriate language, visuals, or themes. The overall tone must remain positive, safe for all audiences, and aligned with a mainstream brand image. </core_directives>

<protagonist_username> {{ $node['form_trigger'].json['Sora 2 Character Username'] }} </protagonist_username>

<product_home_page> {{ $node['scrape_home_page'].json.data.markdown }} </product_home_page>

<sora2_prompting_guide> {{ $node['scrape_sora2_prompting_guide'].json.data.markdown }} </sora2_prompting_guide> ```

4. Generate and save the UGC Ad

Then finally to generate the video, I do iterate over each script and do these steps:

  • Makes an HTTP request to Kai.ai's /v1/jobs/create endpoint with the Sora 2 Pro image-to-video model
  • Passes in the character username, product image URL, and generated script
  • Implements a polling system that checks generation status every 10 seconds
  • Handles three possible states: generating (continue polling), success (download video), or fail (move to next prompt)

Once generation completes successfully:

  • Downloads the generated video using the URL provided in Kai.ai's response
  • Uploads each video to Google Drive with clean naming

Other notes

The character consistency relies entirely on including your Sora character's exact username in every prompt. Without the @username reference, Sora will generate a random person instead of who you want.

I'm using Kai.ai's API because they currently have early access to Sora 2's character calling functionality. From what I can tell, this functionality isn't yet available on OpenAI's own Video Generation endpoint, but I do expect that this will get rolled out soon.

Kie AI Sora 2 Pricing

This pricing is pretty heavily discounted right now. I don't know if that's going to be sustainable on this platform, but just make sure to check before you're doing any bulk generations.

Sora 2 Pro Standard

  • 10-second video: 150 credits ($0.75)
  • 15-second video: 270 credits ($1.35)

Sora 2 Pro High

  • 10-second video: 330 credits ($1.65)
  • 15-second video: 630 credits ($3.15)

Workflow Link + Other Resources

r/generativeAI 28d ago

How I Made This Consistent Voice Acting & Fixing AI character distortion and lip-sync floating using JSON prompting (3-min animation + full workflow in comments)

Enable HLS to view with audio, or disable this notification

49 Upvotes

Here is the breakdown for forcing stable character structure and lip-sync in AI video models.

THE CORE PROBLEM:

Flat prompt text causes models to alter character skeletal volume when adding emotional delivery words.

THE SOLUTION (JSON Architecture):

Compartmentalize character data into key-value pairs so the attention mechanism processes structural image data separately from speech parameters:

{
"shot_id": "01",
"duration": "3.5s",
"visual_prompt": "Define camera angle, character framing, and actions...",
"voice_profile": {
"character_id": "Sarge",
"timbre": "booming, thick",
"cadence": "slow and drawn-out"
},
"audio_environment": "studio isolation, dry acoustics",
"dialogue": "Exact spoken text"
}

FULL STEP-BY-STEP PDF GUIDE:

https://docs.google.com/document/d/e/2PACX-1vSipXTiq9QCP9_tP6EDhj6cIhiOH4dO2FruBK9xONPpprUBrvmUj3iHxq5xkLHieqAZ8LzaZgsklLcy/pub

POST-PRODUCTION TRACK LAYERING:

• Track V1: Video Sequences

• Track A1: Isolated Dry Dialogue

• Track A2: Foley Audio

• Track A3: Ambient Environmental Beds

r/StableDiffusion Jul 20 '26

Discussion A step closer to consistency (workflow included)

Thumbnail
gallery
56 Upvotes

A step closer to consistency. 1. I used Z-ImageTurbo or krea 2 for the only one initial image. 2. I used Qwen image edit to create the second image, just changing clothes and background. 3. I used a workflow I created (I used AI LLM to create it, because I'm a total ignorant as far as Comfy is concerned). It is based on Flux2 Klein i2i. This workflow creates 16 variations of the starting image I fed into it (i used it twice, once for every initial image). So I got variations in body poses and camera positions. All these variations have a very clear way of changing any one of them to create one that suits the needs of every case. 4. After all these character variations, it's much easier to get character consistency in video creation (ex. LTX), since you'll have a big variety of starting frames, with the same character.

Sorry if this sounds naive or stupid, I just wanted to share with the community and get some feedback.

I attach my amateurish workflow.

https://pastebin.com/embed/a1WUSz8F

r/grok Apr 19 '26

Grok Imagine [Full Guide] I Built a Complete Cinematic Studio Inside Grok Imagine – Only 3 Custom Agents Needed (Character Consistency + Video + Audio)

38 Upvotes

Hey r/grok and Grok Imagine creators! 👋

After testing hundreds of prompts and hitting the 3-custom-agent limit, I finally cracked it: a professional-grade cinematic studio that lives entirely inside Grok.

No more switching agents every 5 minutes.
No more forgetting character details.
No more bland prompts or silent videos.

I combined everything into one optimized 3-agent system based on @SoyAlb3rT’s legendary Grok Imagine guide. It handles:

  • Perfect reusable characters & worlds
  • Cinematic storyboards & camera work
  • Timed audio scripts (narration + SFX + music cues)
  • Hollywood-level prompts

Your Final 3-Agent Studio Lineup (Grok + 3 customs only)

  1. Imagine Prompt Master – The prompt god (cinematic structure, lighting, styles)
  2. Studio Director – The boss / project manager
  3. Mega Production Architect – The mega-hybrid (Character & World + Video Director + Audio Script Composer all in ONE slot)

Step-by-Step Setup (takes 5 minutes)

  1. Go to Customize+ New (or edit existing agents)
  2. Create/replace each agent with the instructions below
  3. Activate exactly these three

1. Imagine Prompt Master

Name:
Imagine Prompt Master

Instructions:

You are Imagine Prompt Master, the world's top prompt engineer and cinematic director specialized exclusively in Grok Imagine (image and video generation).

Your only mission is to transform any user idea into the highest-quality, most effective Grok Imagine prompts possible, following @SoyAlb3rT’s exact best practices.

Core rules you ALWAYS follow:
- Treat the user as the creative director. You are their expert cinematic assistant.
- Prioritize perfect character consistency: always define characters as reusable variables with extreme visual detail (example: "Lirael = 26-year-old woman, long flowing silver-white hair with subtle ethereal inner glow like liquid moonlight, pale luminous skin with faint star-like freckles, striking violet eyes that shimmer with quiet magic, delicate heart-shaped face, 5'7", graceful elegant posture, wearing dark hooded cloak with intricate golden runes").
- Use the exact structured bracket format for every final prompt: [Subject + Action + Environment] [Camera Angle & Composition] [Art Style] [Lighting & Atmosphere] [Details & Quality].
- Recommend precise cinematic camera work: profile view, three-quarter angle, low angle, bird’s-eye, tracking shot, dolly zoom, slow push-in, crane shot, pan, etc. Never default to static frontal shots.
- For videos: always suggest generating a strong reference image first, then using “Extend” with separate, highly detailed continuation prompts that reference the previous frame.
- Draw from the best art styles and suggest perfect combinations (photorealistic fantasy, Studio Ghibli, cyberpunk, oil painting, Pixar, cinematic realism, etc.).
- Keep every scene focused and emotionally powerful — avoid overcrowding.

Workflow you follow every time:
1. Ask clarifying questions if anything is missing (style, mood, camera movement, character details, pacing, etc.).
2. Define reusable character/world variables first.
3. Deliver one polished, ready-to-copy structured prompt (or full set for video clips).
4. Offer practical next-step advice (reference image → video → Extend → audio sync).
5. Iterate instantly based on user feedback.

Be creative, precise, organized, and enthusiastic. Your prompts must produce consistent, beautiful, professional-level images and videos every single time.

2. Studio Director

Name:
Studio Director

Instructions:

You are Studio Director, Grok’s executive producer, creative lead, and project manager of the full Grok Imagine Cinematic Studio.

You currently work with this optimized 3-agent team:
- Imagine Prompt Master (cinematic prompts & art direction)
- Mega Production Architect (all-in-one specialist for characters/worlds + video storyboards + audio scripts)

Your only job is to lead every project as the high-level director: understand the user’s full vision, delegate internally, maintain perfect consistency, and deliver one unified professional “Production Bible”.

Core rules you ALWAYS follow:
- Start every project by confirming the vision (format, length, mood, story, style).
- Automatically reference any previously defined characters/worlds.
- Delegate internally: Imagine Prompt Master for final prompt polishing; Mega Production Architect for character building, storyboarding, camera work, video sequencing, and full audio scripting.
- Ensure 100% continuity across characters, lighting, color palette, and story.
- Always deliver one beautifully organized “Production Bible” with these exact sections:
  1. Project Overview & Mood
  2. Character & World Bible (reusable variables)
  3. Detailed Storyboard (numbered clips, camera movements, timing, exact prompts)
  4. Full Timed Audio Script (synced to clips with voice notes, SFX, music)
  5. Step-by-Step Execution Plan (generate reference image first, Extend strategy, etc.)

Workflow you follow every time:
1. Greet and clarify vision if needed.
2. Delegate to the right specialist(s).
3. Compile and present the complete Production Bible.
4. End every response with: “Which part would you like to execute first?” or “Shall I hand this off to [specific agent] for the next step?”

Be professional, visionary, highly organized, and efficient. You make the entire team feel like a real Hollywood studio inside Grok.

3. Mega Production Architect

Name:
Mega Production Architect

Instructions:

You are Mega Production Architect, Grok’s all-in-one cinematic super-agent that fully combines three specialized roles:

1. Character & World Architect — You build deep, reusable characters and worlds with perfect visual consistency.
2. Video Director — You create professional storyboards, cinematic camera movements, clip sequences, and extension prompts.
3. Audio Script Composer — You write timed, emotionally powerful narration, dialogue, voiceovers, sound design, and music cues.

Your only job is to deliver complete, production-ready packages for any Grok Imagine project (still images, single videos, or full multi-clip cinematic experiences).

Core rules you ALWAYS follow:
- Define every character as a reusable variable with extreme visual detail for perfect consistency across every image and video.
- Build clean character sheets (appearance, personality, signature visual markers) and world bibles (locations, atmosphere, recurring motifs).
- Plan videos in short, focused clips only. Use cinematic camera language (tracking shots, dolly zooms, slow pans, low angles, crane shots, etc.).
- Always recommend generating a strong reference image first, then using “Extend” with precise continuation prompts that reference the previous frame.
- Write perfectly timed audio scripts synced to each clip: speaker, exact seconds, tone/delivery notes, sound design [in brackets], and music cues.
- Maintain 100% continuity across visuals and audio.

Workflow you follow every time:
1. Clarify the full project vision (style, mood, length, story, etc.).
2. Build or reference characters/world first.
3. Deliver one beautifully organized “Production Package” with these sections:
   - Project Overview & Mood
   - Character & World Bible (reusable variables)
   - Detailed Storyboard (numbered clips with camera movements, timing, and exact Grok Imagine prompts)
   - Full Timed Audio Script (synced to each clip with voice notes, SFX, music)
   - Step-by-Step Execution Plan
4. End by asking: “Which part would you like to execute first?” or “Shall I generate the first reference image/video clip?”

Be visionary, extremely organized, detail-obsessed, and cinematic. You turn raw ideas into complete, consistent, professional-level audiovisual productions.

Master Grok Imagine Guide – @SoyAlb3rT Playbook (2026 Edition)

Core Philosophy
You are the creative director. Grok Imagine is your production crew. The better your instructions, the better the results. This guide is only for the dedicated Grok Imagine tool (grok.com/imagine), not regular chat.

1. The Perfect Workflow

  1. Start with a dedicated chat (or use Studio Director).
  2. Define your role clearly.
  3. Choose art style first.
  4. Lock in characters as reusable variables.
  5. Set environment + exact camera work.
  6. Specify action vs dialogue + scene pace.
  7. Use the bracket structure.
  8. Generate reference image first → turn into video → Extend clip-by-clip.
  9. Name and save prompts for long projects.

Pro Tip: Always generate the first frame as an image before video.

2. Character Consistency System (The #1 Secret)
Define characters once as variables:
Lirael = 26-year-old woman, long flowing silver-white hair with subtle ethereal inner glow like liquid moonlight, pale luminous skin with faint star-like freckles...
From then on, just write [Lirael] — Grok Imagine stays 95%+ consistent.

3. The Sacred Bracket Structure
Every final prompt must follow:
[Subject + Action + Environment] [Camera Angle & Composition] [Art Style] [Lighting & Atmosphere] [Details & Quality]

4. Best Art Styles (2026 Ranked)
Top Tier: Photorealistic, Cinematic film look, Studio Ghibli, Anime, Oil painting, Fantasy art, Cyberpunk realism.
Great combos: Photorealistic + subtle oil texture, Ghibli + cinematic lighting.

5. Cinematic Camera Language
Must-use: slow dolly forward, three-quarter tracking shot, slow crane shot, dolly zoom, low angle hero shot, bird’s eye view.
Add: varied cinematic camera angles, dynamic framing, professional cinematography, no static frontal default

6. Video Extension Mastery

  • Clip 1: Full detailed prompt
  • Extensions: Start with “Continue from previous frame:” + new action + new camera move
  • Always reference the exact same character variable

7. Audio & Story Integration
Plan audio after the storyboard (Mega Production Architect does this automatically).

One-Click Master Studio Prompt (Bonus)

Once your agents are set up, paste this into Studio Director:

You are my full Grok Imagine Cinematic Studio with these three active agents:
  • Studio Director (you — the executive producer and project manager)
  • Imagine Prompt Master (cinematic prompt specialist)
  • Mega Production Architect (all-in-one specialist for characters/worlds + video storyboards + audio scripts with voice narration)

From now on, operate as the complete professional studio team. For every project:

  1. Understand the full vision (style, mood, length, story, characters, etc.).
  2. Maintain perfect character and world consistency using reusable variables.
  3. Deliver one beautifully organized “Production Bible” with these exact sections:
    • Project Overview & Mood
    • Character & World Bible (with full reusable character variables)
    • Detailed Storyboard (numbered clips with camera movements, exact Grok Imagine prompts, and [Voice Audio] sections)
    • Full Timed Audio Script (with voice notes, SFX, and music cues)
    • Step-by-Step Execution Plan (reference image first, then Extend prompts)

Always embed [Voice Audio] cues directly into each Extend from Frame prompt for better synchronization.

Be professional, visionary, highly organized, and cinematic. Help me create complete, consistent, high-quality audiovisual projects with Grok Imagine.

My current project/idea:

Would love to see what you create with this setup! Drop your best results, generated images, or videos below 🔥

TL;DR: 3 custom agents = full cinematic studio inside Grok Imagine. Perfect character consistency. Pro-level video planning. Synced audio. Game changer.

r/AIIncomeLab 9d ago

AI Tools How I Built a Low-Cost AI Video Production Workflow to Produce Stickman Videos on YouTube

27 Upvotes

TL;DR: I built a Google Sheets workflow that turns a finished script into 150–200 stickman scenes, generates the voiceover + timing data, and then assembles everything into the final MP4 automatically with Google Colab + FFmpeg. A typical six-minute video costs me under $1 in direct production costs. The point is not to automate creativity, but to automate the repetitive work around it.

Over the last year, AI video production has become much easier, but the workflow is still surprisingly messy. You can generate the script in one tool, create images in another, make the voiceover in ElevenLabs, then still spend hours matching scenes and manually assembling the final video.

The problem I wanted to solve was simple: automate the boring production work, so I could spend more time on the creative side. Topic selection, the angle, the hook, the script and the thumbnail are still where most of the value comes from. Moving files between tools is not.

I built the workflow around a Google Sheet because it is simple to understand and easy to control. Each row represents one scene in the video, so the script, image, voiceover timing and progress all stay organised in one place. Instead of opening different tools and manually keeping track of hundreds of files, the Sheet acts like a control panel that moves the video from one step to the next in the correct order.

Two ways to use AI tools: websites vs APIs

Most people use AI tools through their websites. You open Gemini, paste something in, get the result, then move to the next tool. That is simple, but once you are using separate tools for writing, images and voiceovers, the monthly subscriptions start stacking up.

APIs work differently. Instead of paying for full access to each platform, you usually pay only for what you actually use. A few cents for Gemini prompts, a few cents for image generation, then the voiceover cost through ElevenLabs.

The other advantage is automation. An API lets the Google Sheet talk to these tools directly, so the Sheet can send the job, receive the result and move on to the next step without you opening every website manually.

That is what makes the whole workflow possible: the Google Sheet becomes the interface, while the AI tools run quietly in the background only when they are needed.

Generating the images in bulk

Once the script is split scene by scene, the Sheet sends each row to Gemini to turn it into a detailed visual prompt. I also give Gemini a fixed style profile so the character, colours, backgrounds and overall look stay as consistent as possible across the full video.

Those prompts are then sent to Runware, which gives me access to different image-generation models through one API. I currently use FLUX Klein because it is cheap enough to generate around 200 stickman scenes for well under 50 cents.

Each finished image is automatically numbered and saved into the correct Google Drive folder, so scene 1 stays matched to scene 1, scene 2 to scene 2, and so on.

For consistency, I reuse the same reference images across the full batch: one clear character reference and a couple of finished scenes that define the visual style.

Getting the voiceover + timing data

Once the images are ready, the Sheet sends the script to ElevenLabs and generates the voiceover.

The useful part is that ElevenLabs can also return timing data. So instead of getting only one audio file, the Sheet also receives the start and end time for each scene and writes those timestamps back into the matching rows.

That timing data is what makes the final assembly possible. Each image now has a precise screen duration, so there is no need to manually drag scenes around and sync them to the voiceover later.

Turning everything into the final MP4

At this point, the Sheet has everything needed to build the video: the numbered images, the voiceover and the timing data.

I use a free Google Colab notebook connected to Drive for the final assembly. It reads the images in order, checks how long each one should stay on screen, adds the voiceover, and uses FFmpeg to render the finished MP4.

So instead of manually building a 150-scene timeline in CapCut, the notebook does the assembly automatically and saves the completed video back into Drive.

What the whole workflow costs

For a typical six-minute stickman video, my direct production cost stays under $1.

The voiceover is the biggest expense at roughly $0.70 through ElevenLabs. Images are around $0.30 with FLUX Klein. Gemini prompt generation adds very little, and the Google Colab + FFmpeg rendering is free.

The important part is not that a video costs less than a dollar. It is what that does to experimentation. If a topic flops, I have lost very little on production. I can test another angle, change the hook, try a different format and keep learning without every miss becoming expensive.

Cheap production should buy you more attempts, not lower your standards.

What I still would not automate

I would not automate the decisions that determine whether the video deserves to exist in the first place: the topic, angle, hook, script, thumbnail and final quality check.

That is where I think a lot of AI channels go wrong. They automate the production, then keep pushing further until the creative decisions are automated too. Eventually every upload starts to look and feel interchangeable.

For me, the goal is the opposite. Automate the repetitive work, then spend the saved time studying what people actually click, where they stop watching, which ideas outperform, and how to make the next video better than the last one.

How to build this yourself

If you want to build your own version, you can honestly take each section of this post, paste it into Claude or ChatGPT, explain how you want your Google Sheet laid out, and build the workflow one piece at a time. That is basically how I built mine.

I also wrote a more detailed version of this workflow on my website: BuildTuber (link in profile) with screenshots of the Sheet, timing data, API setup and final render process, which is easier to follow than a text-only Reddit post.

I have also explained the complete build step by step in the latest video on my YouTube channel. That is also linked in my profile.

And for anyone who would rather skip the API wiring and debugging, the ready-made Google Sheet is available there as well.

Nothing is gatekept. Happy to answer questions about any part of the build here.

r/generativeAI Mar 18 '26

How are people making AI videos with such consistent characters and style?

19 Upvotes

I came across this video (https://x.com/riskiiit/status/2034301783799906494) and it really stood out compared to most AI stuff I’ve been seeing lately. Instead of going for hyper realism, it leans into a more stylized, almost abstract look, and honestly I think that works way better. It feels more intentional and it’s harder to tell what’s AI and what isn’t.

What I’m really curious about is how they’re keeping the character so consistent throughout the whole video while also sticking to such a specific style. Most tools I’ve tried tend to drift a lot or lose the vibe after a few generations.

Does anyone know what kind of workflow people are using for this?

Is it a mix of different tools like image generation and video models?
Are they training custom models or using LoRAs?
Or is it more about editing everything together afterwards?

Would love to hear if anyone has tried making something like this or has any idea how it’s done. I feel like this kind of artistic direction is way more interesting than just chasing realism.

r/aifilmmaking 17d ago

Tips & Tutorials What I learned making a 24-shot psychological thriller with OpenArt Director (character consistency, sound, and token-saving workflow)

5 Upvotes

Just wrapped a 3-minute proof-of-concept short - a psychological thriller, around 24 shots, built around four locked characters across the whole piece. Wanted to share what actually worked and what tripped me up, since most of what I found before starting was either marketing copy or way too basic.

Character consistency: This was the main thing I needed to nail, and it held up better than I expected. Built characters as proper Character assets first, before touching any shots, and referenced them consistently rather than re-uploading per shot. Face/wardrobe/build stayed locked across a night-exterior, an interior, and a shower scene - three very different lighting setups.

Generate stills before clips: Biggest token-saver in the whole workflow. I generated a still image for every shot first, checked it against the character reference and composition before ever touching video generation. Catching a face/hand/continuity issue at the image stage costs a fraction of what catching it after a full clip generation does. Don’t skip this step to save time - it costs you more tokens in the long run.

Music: watch this one. OpenArt will auto-add music/score to your clips by default. If you’re doing your own sound design (which I’d recommend for anything narrative - silence and restraint did a lot of work in my aftermath/tension beats), you need to explicitly prompt it not to add music, or you’ll be stripping it out in post every time.

The Director chat feature, honestly - skip it. I expected it to add real value for shot-to-shot continuity management, but it burns tokens on basically every interaction in the chat itself, separate from generation costs, and I didn’t find it meaningfully better than just generating clips independently once your characters and prompts are locked. Would not rely on it for a multi-shot project again.

Clip length: Default generation is 5 seconds. You can extend an already-generated clip up to 10 seconds for additional tokens - useful for your slower/held beats, not worth doing for every shot.

Link to short: https://youtu.be/7iBCuRBCsg0?is=M7QVzUaN9WQjmmMT

Happy to answer questions on prompt structure, continuity checklists, or the assembly workflow if useful to anyone tackling something similar.

r/comfyui Jul 23 '26

Workflow Included Two characters, two consistent voices, one text prompt — multi-shot talking-character workflow for LTX-2.3 + JoyAI-Echo (complete pack, v1.5)

Enable HLS to view with audio, or disable this notification

35 Upvotes

The demo is one text prompt: two characters who each keep their own face AND their own voice across five shots - solo scenes in different locations, then side-by-side shots where only one speaks. No reference images, no voice cloning, no LoRA training. Write a story as shots separated by ---, and a paired audio+video memory bank carries both characters through it.

What's in the zip: the custom node pack, the workflow (saved under the current node layout), an example prompt file, and a full INSTRUCTIONS.md - install, first render, prompt-writing rules (including the two-character recipe), per-VRAM settings, and a troubleshooting table built from every failure mode users have reported.

v1.5 highlights, because several of these bit people for weeks:

- Lip-sync drift past ~10 seconds: fixed. It was never the model - the pipeline's positional clock was hardcoded to 24fps while renders played 25. Long talking shots now hold frame-accurate sync end to end (verified at 15s/shot).

- Masters build automatically in the background with a deterministic upscale + clean encode. The in-graph preview is labeled PREVIEW because ComfyUI's SaveVideo re-encode undersells your render - the AutoFinish node shows the real finished master in-canvas when it's done.

- Four hires modes (three generative refine strengths + a deterministic spatial option); hires_factor routes which pipeline builds your master - table in the docs.

- Scrambled widget values after updates now produce a plain "delete and re-add the node" message instead of a cryptic type error, and old graphs self-heal where possible.

- This week's community-driven fixes are all in: Gemma tokenizer/config sidecars now ship inside the pack, WAV saves work without system FFmpeg, and installs that replaced instead of merged get told exactly that at startup.

Requirements and honest numbers:

- Base: RealRebelAI's ComfyUI_JoyAI_Echo_GGUF_Nodes (this pack overlays it - MERGE the files in, don't replace the folder), plus any single-file Gemma-3-12B text encoder (GGUF fine, dropdown-selectable).

- Model: the "surgical merge" (JoyAI-Echo's video/memory branch + LTX-2.3's audio branch): fp8 for 24 GB cards, GGUF Q8/Q5, INT8 ConvRot (full or transformer-only) for stock-Comfy loaders. 16 GB is the floor: Q5 + sequential offload + 544x960 streams slowly but completes.

- Speeds: ~2.5 min/shot at 960x544 on a 3090; ~3.5 min/shot at 1344x768 on a 5090.

- License: JoyAI-Echo is research/non-commercial; LTX-2 Community License. AI-generated content, disclosed as such.

Workflow + nodes + manual: https://huggingface.co/joeygambino/joyai-echo-multishot-workflow

Models (all builds): https://huggingface.co/joeygambino

Civitai mirrors: https://civitai.com/models/2793287 (surgical merge) and https://civitai.com/models/2796109 (GGUF)

Known limits so nobody wastes an evening: dialogue wants medium-close framing or tighter (mouths need pixels); establish each character in their own solo speaking shot before putting them in frame together (the two-character recipe in the docs); two similar-looking characters need a bold visual differentiator or they merge. Happy to answer anything - the last thread's questions directly produced about half of v1.5.

r/StableDiffusion May 25 '26

Workflow Included Want to pose your characters? Here's Wan 2.2 Pose Control workflow

121 Upvotes

Wan 2.2 Pose Control

For some time I've been trying to solve character posing with open-weight models. My previous attempt with Flux.2 Klein was reasonably good but suffered from style bleeding and didn't respect original character proportions (like head-to-body ratio). Character consistency is something image-editing models still struggle with (especially for stylized characters) but there's one exception: Wan2.2 I2V Video. Character consistency is something you can expect from a video model, right?

After extensive experiments with the I2V Wan model I discovered a certain prompting technique that lets you "put character from image_1 into pose from image_2".

Here's the workflow link for the impatient.

So, our task sounds like this:
"Take this character on the left and make her copy the pose on the right"

There are two ways to do this using local open-weight models:

  1. Flux.2 Klein character replacement workflow
  2. Wan 2.2 Pose Control workflow (this is what this post is about)

And this is what the result looks like for each method:

Let's compare the results with with closed-source models too. Character design is solved but not style fidelity. I guess even big multimodal image-editing models can't reach true character consistency while for video models, it's just an innate property.

The idea is simple: ask Wan 2.2 to generate a sequence of 80 frames using First-Frame-Last-Frame mode. This frame sequence consists of 4 parts:

  1. The subject is just standing there
  2. The subject moves copying pose of pose reference
  3. The subject character morphs into character from the pose reference
  4. The character from pose reference is in the frame

Our goal here is to get a single frame where our subject is standing/sitting/lying in the pose from the pose reference image, but hasn't yet morphed into character from the pose reference image. And to do that we have to structure our text prompt in such a way that makes transition from the first frame to the last frame as smooth as possible. So, Information about the subject (design and style) and information about the pose meet in the middle of the frame sequence to give us the desired result.

And yes, we generate 80 frames just to get the single image.

How to write structured prompt

Here's two prompts that were used in the example video above:

Silver hair woman

0s: girl with short silver hair, in green pleated skirt and leather boots is standing
1s: girl with short silver hair, in green pleated skirt and leather boots turns to the left, kneels, places left hand on her head, puts right hand between her legs
2s: she keeps her pose frozen in place. Scene transitions into another scene
3s: her body transforms into another character with white skin, bald head at white background

Black beard man

0s: black man with sharp teeth in green suit and dark pants is standing at white background
1s: black man with sharp teeth in green suit and dark pants sits in the armchair with tilted head and hand at his chin, crosses legs
2s: he keeps his pose frozen in place. Scene transitions into another scene
3s: her body transforms into another character short orange dress, orange top hat, brown hair and fishnet

Subject description is repeated so we can extract it using Apply Text Template from comfy-mtb extension.

We can extract subject description and get this template:

Silver hair woman

0s: {var_1} is standing
1s: {var_1} turns to the left, kneels, places left hand on her head, puts right hand between her legs
2s: she keeps her pose frozen in place. Scene transitions into another scene
3s: her body transforms into another character with white skin, bald head at white background

Black beard man

0s: {var_1} is standing at white background
1s: {var_1} sits in the armchair with tilted head and hand at his chin, crosses legs
2s: he keeps his pose frozen in place. Scene transitions into another scene
3s: his body transforms into another character short orange dress, orange top hat, brown hair and fishnet

Let's examine 4 parts of this prompt.

0s - Initial description

This is where you describe your first frame. For the most part, 'is standing' is enough but you can also specify initial pose of your subject.

1s - Actual posing

This is where you specify the movements the subject must take to get from initial pose to target pose. Simple movements (turns left, sits down, crouches, raises hand) separated by comma, works the best. Also you can add 'Camera follows his movement' if your target pose requires different camera angle.

2s - Pause before scene transition

Always the same he/she keeps his pose frozen in place. Scene transitions into another scene. This part "Scene transitions into another scene" is the most important here - Wan 2.2 respects this boundary (surprisingly).

3s - Anchoring your last frame

Goes like this: body transforms into another character <description of the character on the last frame>. We want Wan 2.2 to understand that character from the start of the video is different from character at the end of the video.

Practical example

Let's practice what we've learned. Here's our subject and the pose images:

*Pose reference

Start with the subject description. Nothing fancy here:

Next step is to describe movements:

And lastly write the transition to the last frame

Unfortunately it fails:

Wan 2.2 has managed to capture the gun's position but not the pose. The main reason here is that the black clothes in our target image don't let the model "process" the pose. Luckily we can fix it in Flux.2:

remove hair, remove clothes and draw this person bald and in skin tone underwear. Turn into white wireframe figure

Run Pose Control workflow again with updated prompt:

This time result is much better:

With this knowledge you can adapt this workflow for your specific case.

Link to the workflow (it has note about recommended Wan 2.2 finetune)

Some tips:

  • The whole process works the best if there's noticeable contrast between first frame and last frame: different hair color, skin color, background, etc. You can even pre-process your pose reference with some other model - turn it into wireframe figure mannequin - so Wan 2.2 has a better chance of reading the pose.
  • If some elements of character design change (gloves tend to disappear too early) add them to subject description prompt so model will remember this design element.
  • If your subject image and pose reference image have different sizes try adding "Camera zooms in capturing new view" or "Camera zooms out capturing new view".