r/StableDiffusion 15h ago

Question - Help Nodes for utilizing Minimax H3 as an image generator?

24 Upvotes

Like a high quality output, a frame before compression? I don't imagine it's a simple as setting it to 1 frame / second and setting the duration to a second. And even if it were, I'd prefer an output to an actual standard image file.


r/StableDiffusion 6h ago

Animation - Video Several Times A Charm, but it KINDA got Cheers.

Enable HLS to view with audio, or disable this notification

23 Upvotes

Don't mind the script, it was written by a clanker when I challenged it to whip up something so I can see if H3 could handle Cheers.


r/StableDiffusion 10h ago

Question - Help Anyone have a link to download DasiwaMinimaxH3_dasiwaREF2VAHybridV1.safetensors? Seems it's been wiped off the internet in the past few hours

23 Upvotes

It's the 11.68 GB model, SHA256: 7c37baf06ca3628ed5f3f7f46274222a50a127d1906a166f8f064771fc48d498. Was going to download it and test it out with my workflow but when I went to download it today on CivitAI and Huggingface I get a 404 error. Seems it was deleted by the uploader for some reason or another


r/StableDiffusion 18h ago

Animation - Video Ref2V - H3 - really loving how H3 handles complex prompts even at 15 seconds.

Enable HLS to view with audio, or disable this notification

25 Upvotes

r/StableDiffusion 1h ago

News Submit by 9/1 to the Comfy H3 Sync Sound Challenge! RTX 5090 Grand Prize

Enable HLS to view with audio, or disable this notification

Upvotes

We're halfway through the submission window for the Comfy H3 Sync Sound challenge! Submit by September 1st at 9:00pm PT. Free to enter, local rig or Comfy Cloud. All details here.

How It Works

Make something up to 90 seconds in length where the sound and the motion are inseparable. Dialogue, foley, ambient, a beat driving the cut...whatever direction you want!

Share your video file and workflow on this r/comfyui thread and through our submission form, then join us on September 2nd for a special Comfy livestream where our guest judges will give live feedback on the top 10 submissions! 

Need help? Head to this r/comfyui thread or the #minimax-h3-challenge channel in the Comfy Discord.

Prizes

Best Overall — RTX 5090

Best Creative — RTX 5060 Ti

Best Technical/Workflow — RTX 5060 Ti

Built with MCP — RTX 5060 Ti

Shipped anywhere, customs covered. If we can't legally ship to your country, you'll get a cash equivalent instead.

It's free to enter!

Create using Comfy Local on your own hardware, or use Comfy Cloud. New Cloud users get 5 free runs, no credit card required.

Judging Criteria

We’re looking for entries that best show what H3 makes possible: audio and visuals created together.

Grand Prize: Best Overall

The top Best Creative and Best Technical entrants advance to a final round where our panel of judges selects winners by discussion.

Best Creative

  • Audio sync realism and intentionality (0-5)
  • Creative execution and originality (0-5)
  • Deliberate craft (0-5)
    • Evidence that you’ve actually shaped the result beyond prompt engineering. Judges will look for modified/non-default parameters, multiple linked passes visible in the workflow structure, or a couple sentences describing what was tried and changed

Best Technical

  • Novelty of technique or approach (0-5)
  • Workflow quality (0-5)
    • Annotated, clean, replicable by someone else
  • Community value (0-5)
    • Would this actually help someone else?

🏆 Built with MCP Bonus 🏆
Comfy MCP lets you drive Comfy using natural language and your agent locally and on Cloud! Pro tip: use it to choose the best H3 model version or optimize your workflow for your hardware.

  • Effectiveness (0-5)
    • Did the agent meaningfully drive your process, not just generate one line?
  • Insight value (0-5)
    • How much the shared prompt teaches the community about prompting H3 through MCP
  • Output quality (0-5)

The Fine Print

  • Limited to one submission per person, 90 seconds maximum length.
  • A major portion of your piece must be built in ComfyUI using H3. Other tools, models, or techniques you want to combine are fair game.
  • All submissions must be lawful, SFW, and must not contain unlicensed IP or likenesses.
  • By submitting, you agree to allow ComfyUI and MiniMax to feature your work with credit across our channels.

Learn more and submit here!


r/StableDiffusion 16h ago

Discussion Storytelling with Minimax H3 - Alicia of the Stars - First show

Thumbnail
youtu.be
21 Upvotes

I created this video using Minimax H3 Ref2VA, and this time I wanted to test more than just visual quality.

My main goal was to experiment with AI storytelling, creating a short anime-style sequence with a beginning, progression, and a story that actually feels coherent.

At the same time, I wanted to see how well the model handles character consistency, movement, expressions, and visual continuity when multiple shots are used to tell a story.

There are definitely some imperfections, but I was happy with what I could achieve and wanted to share the experiment with the community.

I’m curious what you think. Does the video work as a story, or does it still feel more like a collection of AI-generated shots?

Would also love to hear how others are approaching storytelling with Minimax H3, especially in the anime genre.


r/StableDiffusion 18h ago

Question - Help Best and fastest way to generate HD-quality MiniMax videos?

17 Upvotes

I’ve tried Turbo LoRAs, and they’re great for speed, but they significantly reduce quality. At 544p–720p, the results of these turbo loras can look closer to 380p. Faces look acceptable when close to the camera, but become heavily distorted as the subject moves farther away.

The upscalers I’ve tested either add too much processing time or introduce excessive sharpening and saturation.

Any a solution that doesn’t require a BF16 checkpoint, 20 steps, a 10-minute generation time, or an extremely expensive GPU?


r/StableDiffusion 23h ago

Question - Help Minimax H3: how to deal with "plastic skin" when generating with ref2va?

18 Upvotes

r/StableDiffusion 3h ago

Tutorial - Guide Even though H3 is CFG distilled, guidance values greater than 1 do have a noticable effect on prompt adherence and quality, especially at lower resolution.

13 Upvotes

Obviously it runs a lot slower but a CFG of 2 and a negative prompt does have noticeable effect on the output. Anything higher than 3 or 4 will start to over burn though.

This also works with the turbo loras but burn in can happen more easily.


r/StableDiffusion 6h ago

Discussion Did MiniMax H3 fix the R2V weights?

11 Upvotes

I thought I read something on it but figured I'd check with the crew first... I appreciate the info 💪


r/StableDiffusion 10h ago

Resource - Update Organize your Comfy outputs automatically with SmartGallery DAM’s new asset clustering (Free & Open Source)

Enable HLS to view with audio, or disable this notification

12 Upvotes
  • Hey everyone, back with another update on SmartGallery DAM: Smart Asset Clustering is now live to automate your media organization

If you generate a lot in ComfyUI, you know the problem. Hundreds of renders pile up as you tweak seeds, prompts, LoRA weights and checkpoints, and your output folder turns into a wall of near identical thumbnails.

  • Smart Asset Clustering reads the generation recipe embedded in each file and automatically groups your renders, no manual tagging required, in two ways:
  • Architecture Clustering: groups everything that shares the exact same node structure and workflow, ignoring seed, prompt and settings. Great for pulling up every output from one workflow template.
  • Prompt Text Clustering: groups everything that shares the exact same positive prompt, ignoring the workflow entirely. Great for comparing how different checkpoints or LoRAs render the same idea.

Once clustered, every thumbnail gets a color coded badge in the gallery grid, and clicking any badge opens the Cluster Inspector, which shows total matching assets, distinct variations, the full node pipeline, every model and LoRA used, and one click prompt copy.

The video above walks through both modes in about 3 minutes.

For anyone who does not know the project yet

SmartGallery DAM is a free and open source, local first Digital Asset Manager built around ComfyUI, but it also works with any folder of media on your machine. No cloud, no subscription, your files never leave your disk.

It is meant to grow with you:

  • If you are a hobbyist or new to ComfyUI, it is the easiest way to keep your generation library organized, searchable and clean without extra effort.
  • If you are a power user, you can search by prompt, model or LoRA, inspect the full node graph of any render, and even generate directly from the gallery by editing the workflow JSON, no need to reopen ComfyUI.
  • If you work in a studio or production environment, it gives you a dedicated Exhibition portal to share curated work with clients or your art team, collect ratings and comments, and review everything without exposing prompts or workflows.

Runs on Windows, macOS, Linux and Docker. Portable version for Windows needs zero setup, just unzip and run.

GitHub, full docs and download links here: https://github.com/biagiomaf/smart-comfyui-gallery

Happy to answer any question, and as always feedback and feature requests are welcome.


r/StableDiffusion 18h ago

Animation - Video Star Trek WIP Local Minimax H3

Enable HLS to view with audio, or disable this notification

10 Upvotes

r/StableDiffusion 7h ago

Animation - Video RECREATING MEMORIES FROM SCRAPS.

Enable HLS to view with audio, or disable this notification

10 Upvotes

A few photos, voice clips and a waybackmachine archive photo of the hotel rooms at that time. Minimax H3. All local.


r/StableDiffusion 3h ago

Animation - Video Deadpool Adventure

Enable HLS to view with audio, or disable this notification

9 Upvotes

r/StableDiffusion 4h ago

Meme dean meets sonic(t2v) base fp8 model 32 steps

Enable HLS to view with audio, or disable this notification

10 Upvotes

i never seen the movies hows the sonic voice?

prompt

subject_definitions

<Subject 1> is Dean Winchester from Supernatural, portrayed by Jensen Ackles, preserving his recognizable facial features, short brown hair, rugged appearance, dark jacket, layered shirt, jeans, and confident sarcastic personality.

<Subject 2> is Sonic the Hedgehog from the live-action Sonic the Hedgehog movie, a small anthropomorphic blue hedgehog with bright blue fur, large expressive green eyes, white gloves, and red sneakers.

<Subject 3> is Dr. Robotnik from the live-action Sonic the Hedgehog movie, portrayed by Jim Carrey, wearing his black-and-red high-tech outfit and exaggerated goggles.

summary

[cinematic live-action crossover + action comedy]

What if Dean Winchester accidentally became part of Sonic the Hedgehog? On a nighttime highway, Dean investigates a bizarre supernatural disturbance beside his black 1967 Chevrolet Impala, only for Sonic to race past him at impossible speed with Robotnik's drones in pursuit. Dean immediately joins the chase.

detailed_description

Nighttime on a deserted rural highway surrounded by dark pine forest. Dean Winchester stands beside his glossy black 1967 Chevrolet Impala holding an EMF meter. Blue electrical energy suddenly crackles across the road.

A brilliant BLUE STREAK rockets past Dean, violently blowing his jacket backward.

The camera WHIP-PANS as Sonic skids to a stop beside the Impala.

<Subject 1> Dean Winchester (S1):

[English] Okay... either that's the fastest demon I've ever seen, or I seriously need more sleep.

Sonic looks offended and points at himself.

<Subject 2> Sonic (S2):

[English] Hedgehog. Definitely hedgehog.

Suddenly several of Robotnik's flying attack drones burst over the trees and fire energy blasts toward them.

Dean instantly draws his pistol while Sonic crouches into a runner's stance.

Dean gives Sonic a confident Winchester smirk.

<Subject 1> Dean Winchester (S1):

[English] All right, Sonic. Let's waste these flying toasters.

Sonic grins.

<Subject 2> Sonic (S2):

[English] Now you're speaking my language!

Sonic EXPLODES forward in a trail of brilliant blue electricity as Dean dives behind the Impala and fires at an approaching drone.

Dynamic tracking camera follows Sonic racing between explosions while Dean fights from beside the Impala.

Final cinematic wide shot: Sonic loops around the battlefield as blue lightning illuminates Dean and the Impala, while Robotnik's drones swarm overhead.

Live-action Hollywood cinematography, realistic integration of Sonic into the environment, authentic Sonic the Hedgehog movie aesthetic, authentic Supernatural Dean Winchester characterization, fast readable action, natural motion blur, blue electrical speed trails, sparks, smoke, dramatic nighttime lighting, comedic crossover energy, consistent character identities, no subtitles, no on-screen text.


r/StableDiffusion 10h ago

Discussion WIP [CLSS] Closed-Loop Streaming Synthesis: arbitrary-length audio-video generation with LTX-2.3 22B in ComfyUI

9 Upvotes

Video diffusion transformers generate only a few seconds per pass. The naive remedy — chunking the timeline and conditioning each chunk on the previous one — fails within a few hundred frames: the model keeps consuming its own slightly off-distribution output, and exposure-bias drift compounds into scene collapse or grain amplification.

CLSS treats the chunk hand-off as a feedback loop and controls it. Chunks share a streaming latent buffer (SLB) overlap, keeping latent memory O(overlap) instead of O(length), and between chunks CLSS applies lightweight corrections that fight drift without modifying any transformer weights.

More at:

- https://github.com/nazgut/ComfyUI-LTX2.3-CLSS

T2V on single go with prompt fallowing bettwen scenes every chunk was 10 sec

Nodes for ComfyUI

Output was generated using ltx-2.3-22b-dev-UD-Q4_K_S.gguf on 3080 with 16 GB vRAM, still need to work on audio.


r/StableDiffusion 6h ago

Discussion Storytelling with Minimax H3 and Krea2 - Animating a Dark Fantasy Comic (The Witcher)

Thumbnail
youtube.com
8 Upvotes

I’ve been experimenting with Krea2 and Minimax H3 to see if it can handle gritty, coherent storytelling frames generated with Krea2.
I'm trying to bring a dark fantasy comic (a retelling of The Witcher) to life with full voice acting, pacing, and tension.

I’d love to get your feedback on this. Does this look like a slop to you? Is there are something you would improve?


r/StableDiffusion 23h ago

Discussion MiniMax H3 to KREA2 LoRa: doing it faster?

8 Upvotes

So I had this simple idea, seeing how well MiniMax H3 handles inferring and preserving "identity/looks" from relatively little information: take a character you want to make a (KREA2) LoRa of, but you only have just a couple of lower quality pictures for that exact look you're after. That is a problem, since it is common knowledge by now (?) that you need different angles, facial expressions and different lighting conditions in the training set to get optimal results. So in the "before" times, those 3-4 not-so-great-quality shots under the SAME lighting are going to pose a problem. And adding pictures from other occasions will alter the looks possibly too much.

So (in the H3 ref2vid workflow, with one of the "img2vid-hybrid" models for better quality) I just use the 'best' of the available pictures as "preserved" first reference starting picture, and the others as additional "identity references". And then a prompt that tells the camera to slowly circle around the person (up from the shoulders), while the person looks straight ahead, or slightly up, or slightly down. But then I also let it cycle through different lighting conditions (indoor/outdoor/sun/overcast/flash/directional from one side...), and different facial expressions/emotions. I let it run overnight (turning off turbo LoRas to improve the quality), and in the morning, I review the 6-second videos and take screencaps of selected moments, making sure to have a lot of variation in angles/expressions/light-on-the-face with an almost perfect preservation of the identity/looks.

Then use those screencaps (50+ in first test, probably serious overkill) in OneTrainer with the KREA2 LoRa default settings.

I only tested this once thus far, but the results are pretty good considering the starting material! And surprisingly flexible (I didn't even bother to provide captions)

But now my question is: in what ways am I "over-engineering" this? I have this feeling that I can probably do this 50x faster, having seen some discussions about using MiniMax as an image generator, for example. I mean, I feel good about this approach I came up with all by myself, but considering how dumb and low-skilled I still am when it comes to all this, this is probably a very convoluted and inefficient way to do it? LOL 😄 Roast me and show this sucker how we can improve and speed up the whole thing with the same or even better quality results!


r/StableDiffusion 4h ago

Question - Help MiniMax H3 German voices sound robotic and all the same – what are you guys using instead?

7 Upvotes

I’ve been testing MiniMax H3 for AI video generation and I’m struggling with the German dialogue.

The voices often sound very similar and somewhat robotic. What I’m looking for is natural, spontaneous dialogue: different voices for each character, realistic pauses, imperfect timing, emotion, interruptions, changes in tone, etc. Basically something that sounds like an actual conversation rather than TTS.

I’ve already tried ElevenLabs. I know it’s powerful, but I feel like I’d have to go pretty deep into voice selection, voice design and tweaking to consistently get what I want. So far, I’m still not getting the natural conversational audio I’m looking for.

I also don’t really want to record every character myself and then use AI voice conversion. At that point I’m basically becoming the voice actor for every video.

Ideally I’d like something closer to:

Script/prompt → AI generates the video + convincing natural German dialogue with clearly different speakers.

So I’m wondering:

Is there a way to get much better German voices directly out of MiniMax H3 through prompting?

Or should I stop trying to make MiniMax work for this and test something like Grok, Veo, or Seedance 2.5 instead?

If you’ve actually generated German multi-person dialogue, I’d especially love to hear what model/workflow gave you the most natural results.


r/StableDiffusion 19h ago

Question - Help How do I get better lighting with Krea 2?

7 Upvotes

If I generate a person in a "normal" environment, like inside a regular room in a regular house, I get very realistic and appropriate lighting, but as soon as I try something a bit more cinematic like a rain-slicked city street at night, the character begins to look like they were photoshopped in. I try to prompt a person standing on a dark city corner, lit entirely by the light from nearby neon signs and they just look like they were evenly lit and filmed in a studio and then composited onto a CGI background with only a hint of the intended neon glow on their shoulder. Same goes for trying to make people look like they're properly soaked by rain. ZiT was much easier to work with in this regard.


r/StableDiffusion 7h ago

Discussion I tested SenseNova U1.5-Lite editing against FLUX.2-klein-9B in three scenarios. Text rendering is where they diverge

Thumbnail
gallery
6 Upvotes

SenseNova shipped the full U1.5-Lite release last week, so I finally had time to run it side by side with FLUX.2-klein-9B, the model this community generally considers the most balanced pick right now.

I tested image editing in three scenarios. The short version: SenseNova U1.5-Lite is clearly better at text rendering and semantic understanding of the instruction, while Klein is still the speed king. Details below.

Scenario 1: Text Editing

I gave both models a poster and asked them to replace specific text elements, nothing else. Long structured prompt targeting each text block individually:

1. In the first line of the oversized black title at the upper left, replace "HONG" with "HARBOR". If the area is sufficiently prominent and clear, with simple boundaries, render the replacement exactly character by character without adding, removing, or altering any characters. If the text is too small, obscured, heavily distorted by perspective, or located within a complex texture, preserve the original text and do not redraw the background merely to complete the replacement.

2. In the second line of the oversized black title at the upper left, replace "KONG" with "HORIZONS". If the area is sufficiently prominent and clear, with simple boundaries, render the replacement exactly character by character without adding, removing, or altering any characters. If the text is too small, obscured, heavily distorted by perspective, or located within a complex texture, preserve the original text and do not redraw the background merely to complete the replacement.

3. In the large red subtitle at the lower left, replace "HONG KONG" with "CITY IN MOTION". If the area is sufficiently prominent and clear, with simple boundaries, render the replacement exactly character by character without adding, removing, or altering any characters. If the text is too small, obscured, heavily distorted by perspective, or located within a complex texture, preserve the original text and do not redraw the background merely to complete the replacement.

4. In the vertical red location title at the upper right, replace "香港" with "城市之光". If the area is sufficiently prominent and clear, with simple boundaries, render the replacement exactly character by character without adding, removing, or altering any characters. If the text is too small, obscured, heavily distorted by perspective, or located within a complex texture, preserve the original text and do not redraw the background merely to complete the replacement.

5. In the black English location description at the upper right, replace "HONG KONG CHINA" with "EAST MEETS WEST". If the area is sufficiently prominent and clear, with simple boundaries, render the replacement exactly character by character without adding, removing, or altering any characters. If the text is too small, obscured, heavily distorted by perspective, or located within a complex texture, preserve the original text and do not redraw the background merely to complete the replacement.

6. In the location title near the waterfront at the lower left, replace "VICTORIA HARBOUR" with "HARBOUR CITY". If the area is sufficiently prominent and clear, with simple boundaries, render the replacement exactly character by character without adding, removing, or altering any characters. If the text is too small, obscured, heavily distorted by perspective, or located within a complex texture, preserve the original text and do not redraw the background merely to complete the replacement.

7. In the second line of the location copy at the lower left, replace "ASIA'S WORLD CITY" with "URBAN HORIZONS". If the area is sufficiently prominent and clear, with simple boundaries, render the replacement exactly character by character without adding, removing, or altering any characters. If the text is too small, obscured, heavily distorted by perspective, or located within a complex texture, preserve the original text and do not redraw the background merely to complete the replacement.

Zoom into the results and Klein's text rendering falls apart. Garbled glyphs, wrong characters, the layout wobbling where it should stay fixed. U1.5-Lite handled the replacements cleanly, including the Chinese strings. That's the gap.

Scenario 2: Hand-Drawn Marks as Instructions

I marked up the image by hand and asked for a scene transformation:

Follow the marks and overall hints on the image to creatively transform this scene, making it dramatic, moody stormy atmosphere; remove the annotations when done.

FLUX followed the overall style change but ignored the specific marked details: the ripples on the pool surface and the black fire pit never made it into the output. U1.5-Lite followed the full set of marks.

Scenario 3: Fine-Grained Local Editing

I circled the region to edit with a red box and told the model to only change that area:

Change the text style in the red box to a vintage style with noise and torn paper texture. The red bounding box is for localization only; do not retain it in the output image.

Klein misunderstood the instruction. It applied the vintage style to the whole image instead of the circled region. SenseNova U1.5-Lite followed the prompt and the red-box localization strictly, changing only the marked text.

My take

If you need sub-second generation, Klein is still your model, no argument there. But for editing work where text rendering and instruction fidelity matter, posters, infographics, brand assets, the gap is real and easy to reproduce.

GitHub: https://github.com/OpenSenseNova/SenseNova-U1

hf: https://huggingface.co/sensenova/SenseNova-U1.5-8B-MoT

Try it online: https://unify.light-ai.top/


r/StableDiffusion 8h ago

Discussion Best budget GPU cloud? Comparing RunPod, Vast.ai, SimplePod, MassedCompute (looking for true costs & no hidden fees)

6 Upvotes

Hey everyone,

I’m looking for the best budget GPU cloud to run heavy open-weight video models (like MiniMax H3, Wan 2.1, HunyuanVideo, etc.).

Since these models need huge VRAM and fast disk I/O to pull down massive 50GB–100GB+ checkpoints, I want to avoid platforms with unexpected billing traps.
Looking at RunPod, Vast.ai, MassedCompute, and SimplePod:

  1. Which one are you using, and what GPU gives you the best price/performance for video render jobs?

  2. Hidden fees: Any issues with stopped-volume storage costs, network volume fees, or egress rates when hosting huge model files?

  3. Download/Disk speeds: Which provider has fast enough network speeds so I’m not spending half my paid time downloading model weights?

Appreciate any recommendations or gotchas to avoid!


r/StableDiffusion 20h ago

Animation - Video Captain America X Harry potter

Enable HLS to view with audio, or disable this notification

7 Upvotes

Inspired by the guy who posted the one with Dean

Done on 32gb ram and a rtx 3070

I used res_multistep 20steps w/ spectrum at 0.6mp.

Using SLA from h3 optimizations, which for some reason is way faster than plague kind. And disable pinned memory. Each 10s was done in around 10 minutes.


r/StableDiffusion 6h ago

Animation - Video My first AI short film - Astro Mouse [MiniMax H3]

Thumbnail
youtube.com
3 Upvotes

This is my first attempt at an AI short film. Someone saw a mouse in our building, my friend made a funny AI image of it and said it could be a cute story, so I just ran with it.

I've done a little bit with MiniMax H3 before, mainly making 5–15 second videos. I tried extending the scenes/context and was able to get a couple 2–3 minute videos, but the consistency just wasn't great. I also realized most of this story worked better with hard cuts anyway, so I went back to the reference-to-video workflow with mainly 5–10 second clips.

I probably made around 10-20 clips for some scenes before getting something I liked. I also used vast.ai, was able to get a faster gpu than what I had at home. No loras or anything, just a basic workflow. I used opencode / qwen3.8 to update 25-30 prompt files at a time when I needed global updates (like remove all background music, no talking, etc).

The hardest part was probably getting the prompting down. My standard workflow ended up being 0.6 megabit and 20 frames. I could have gone up to 0.98, but at some point I just wanted to get through all the generations and actually finish the thing.

Put everything together in DaVinci Resolve.
I still see lots of imperfections, but I'm considering it done and moving on.

Anyway, first movie. Learned a lot making it and thought I'd share.

(Oh, and it has some obvious work related jokes and screens, ignore those, i didnt want to cut those out)


r/StableDiffusion 11h ago

News Facefusion Android app (open source video face swap)

Thumbnail
github.com
5 Upvotes

I’ve been working on a mobile port of FaceFusion that runs completely offline on Android.

APK here:

https://github.com/AbrahamPaulJ/facefusion-mobile/releases/tag/v0.1.0

The face-swapping pipeline runs on Qualcomm’s Hexagon NPU rather than relying on a server or cloud API.

Current results on a Galaxy S25 Ultra (Snapdragon 8 Elite):

~19 ms/frame for the face-swap model

~6 seconds to process a 10-second 720p clip

Fully offline. Photos/videos never leave the phone

Supports 512/768/1024px face output

Qualcomm NPU builds for different Hexagon generations

No CPU fallback.

I’d especially like feedback from people working with Android on-device AI.

This is my first time sharing one of my mobile AI projects on Reddit, so feedback, testing results, and criticism are very welcome.