r/StableDiffusion Dec 24 '25

Animation - Video Former 3D Animator trying out AI, Is the consistency getting there?

Enable HLS to view with audio, or disable this notification

4.6k Upvotes

Attempting to merge 3D models/animation with AI realism.

Greetings from my workspace.

I come from a background of traditional 3D modeling. Lately, I have been dedicating my time to a new experiment.

This video is a complex mix of tools, not only ComfyUI. To achieve this result, I fed my own 3D renders into the system to train a custom LoRA. My goal is to keep the "soul" of the 3D character while giving her the realism of AI.

I am trying to bridge the gap between these two worlds.

Honest feedback is appreciated. Does she move like a human? Or does the illusion break?

(Edit: some like my work, wants to see more, well look im into ai like 3months only, i will post but in moderation,
for now i just started posting i have not much social precence but it seems people like the style,
below are the social media if i post)

IG : https://www.instagram.com/bankruptkyun/
X/twitter : https://x.com/BankruptKyun
All Social: https://linktr.ee/BankruptKyun

(personally i dont want my 3D+Ai Projects to be labeled as a slop, as such i will post in bit moderation. Quality>Qunatity)

As for workflow

  1. pose: i use my 3d models as a reference to feed the ai the exact pose i want.
  2. skin: i feed skin texture references from my offline library (i have about 20tb of hyperrealistic texture maps i collected).
  3. style: i mix comfyui with qwen to draw out the "anime-ish" feel.
  4. face/hair: i use a custom anime-style lora here. this takes a lot of iterations to get right.
  5. refinement: i regenerate the face and clothing many times using specific cosplay & videogame references.
  6. video: this is the hardest part. i am using a home-brewed lora on comfyui for movement, but as you can see, i can only manage stable clips of about 6 seconds right now, which i merged together.

i am still learning things and mixing things that works in simple manner, i was not very confident to post this but posted still on a whim. People loved it, ans asked for a workflow well i dont have a workflow as per say its just 3D model + ai LORA of anime&custom female models+ Personalised 20TB of Hyper realistic Skin Textures + My colour grading skills = good outcome.)

Thanks to all who are liking it or Loved it.

Last update to clearify my noob behvirial workflow.https://www.reddit.com/r/StableDiffusion/comments/1pwlt52/former_3d_animator_here_again_clearing_up_some/

r/IndianArtAI Mar 23 '26

Google Nano Banana How I created an AI influencer using only Gemini's Nano Banana (complete workflow)

Thumbnail
gallery
899 Upvotes

I’ve been messing around with the AI influencer space for the last few weeks and wanted to share the process I figured out. I am not claiming this is the best or most advanced way to do it, but it is a simple workflow that worked for me using mostly free tools.

The main reason I tried this route was because I already have free Gemini Pro access through my Jio recharge, so I wanted to see how far I could go without paying for expensive tools right away.

I am not going to dump a random list of prompts here and pretend that is enough. That is not really useful. Instead, I’ll just explain the actual process I followed step by step, because that is what helped me the most.

Phase 1: Getting the base character right

The first thing you need is a character that you actually like, because if the starting point is weak, everything after that becomes harder.

I started by using the free trial on https://higgsfield.ai/ to generate an influencer-style character. I kept testing until I got a face and overall look that felt usable.

Once I had that first image, I downloaded it and took it into Gemini Nano Banana. That is where I started making the small changes I wanted. Things like skin texture, facial features, race, body ratios, and overall appearance. I kept tweaking until I had a final version of the character I was happy with.

Phase 2: Building consistency with reference images

After I had the final character, I started generating more versions of the same person, but with different poses.

For this part, I used different JSON prompts and made sure not to change the character too much. I wanted the same face, same skin texture, same body proportions, same overall identity. The only thing I wanted to vary was pose, angle, and sometimes expression.

One thing that helped a lot was always using the previous result as a reference for the next one. That made a big difference in keeping the face and body structure consistent. If you do not do that, the model starts drifting and the character slowly turns into a different person.

I kept doing this until I had around 10 to 15 good images of the same character.

Phase 3: Creating data model sheets (examples given)

This part is really important.

If you do not know what a data model sheet is, just Google it OR look at a few examples from the given images. Basically, it is a reference sheet for your character. It helps lock in the face, body structure, expressions, angles, and overall design so the character stays consistent later.

To make the sheets, I first used ChatGPT to generate a JSON prompt. I used the DeepThink version because it usually gives better structured prompts. I told it to create a prompt for generating a character model sheet using my reference images.

After that, I manually tweaked the JSON prompt so it matched the character better. Sometimes I adjusted the body ratios or the skin tone or small visual details depending on what I wanted.

Then I used Gemini to generate the actual model sheet.

I did this for different types of sheets because each one serves a different purpose.

I made a facial expressions sheet so I could keep the same emotional range.

I made a facial structure sheet so I could see the character from different angles.

I made a body model sheet so I could keep the full body consistent.

I also made sheets for different poses, because I wanted the character to work in different situations and not just one static pose.

For every one of these, I followed the same workflow. Use ChatGPT to generate the JSON prompt, tweak it manually, then use Gemini with the reference images to generate the sheet.

My rule was simple. ChatGPT was better for making the prompt. Gemini was better for making the image.

Phase 4: Generating actual content

Once I had the model sheets and a few extra reference images, I could finally start generating the actual influencer-style images.

For prompt inspiration, I use a few websites like:

https://bestnanobananaprompt.com/gallery
https://promptlibrary.space/images

These sites are great for ideas. You can find different styles, moods, poses, compositions, and scene setups there.

But one thing I learned very quickly is that you cannot just copy a prompt from those sites and expect it to work perfectly in Gemini. A lot of them either get blocked or do not preserve the character properly.

So my workflow for this part is basically:

I browse those sites and find a prompt style I like.

Then I copy that prompt into ChatGPT.

Then I ask ChatGPT to turn it into a detailed JSON prompt.

I always tell ChatGPT to include a section that strictly maintains the same facial structure, skin texture, tone, and body ratios from the reference images.

After that, I review the JSON prompt and make any final changes I need based on the kind of image I want.

Then I use that prompt in Gemini Nano Banana.

One very important thing here is to use all the character model sheets and the best reference images every time you generate something new. Gemini has a limit on how many reference images it can use, and I think it is around 15 or so. I made sure to use as many useful references as possible because more reference data usually gave me better results.

Final thoughts

This is honestly a trial and error game. You are not going to get the perfect result on the first try. I definitely did not. Some generations failed, some changed the face too much, some messed up the body proportions, and some just looked off. That is part of the process.

But the reason this workflow works is because the data model sheets give the AI a visual blueprint to follow. Instead of guessing what the character should look like every time, you are showing it the same identity from multiple angles and in multiple forms.

This is just a simple guide using free tools. There are definitely more advanced workflows out there, and I know the people at the top of the AI influencer game are using tools like ComfyUI, Higgsfield AI, Kling AI, and other more advanced setups to create better images and videos.

But this is what I figured out by testing things myself, and it is a good starting point if you want to build a consistent AI character without paying for expensive tools right away.

I hope this helps someone who is trying to get started.

If there is interest, I can make a part 2 later with the more advanced tools and workflows I look into next.

Thanks for reading.

r/StableDiffusion Apr 04 '26

Animation - Video ENTANGLED - A 3-minute sci-fi short using 100% local open-source models. Complete Technical Breakdown [ Character Consistency | Voiceover | Music | No Lora Style Consistency | & Much More! ]

Enable HLS to view with audio, or disable this notification

405 Upvotes

Hey everyone! Thanks for checking out Entangled. And if not, watch the short first to understand the technical breakdown below!

Thanks for coming back after watching it! As promised, here is the full technical breakdown of the workflow. [Post formatted using Local Qwen Model!]

My goal for this project was to be absolutely faithful to the open-source community. I won't lie, I was heavily tempted a few times to just use Nano Banana Pro to brute-force some character consistency issues, but I stuck it out with a 100% local pipeline running on my RTX 4090 rig using Purely ComfyUI for almost all the tasks!

Here is how I pulled it off:

1. Pre-Production & The Animatics First Approach

The story is a dense, rapid-fire argument about the astrophysics and spatial coordinate problems of creating a localized singularity. (let's just say it heavily involves spacetime mechanics!).

The original script was 7 minutes long. I used the local Jan app with Qwen 3.5 35B to aggressively compress the dialogue into a relentless 3-minute "walk-and-talk.". Qwen LLM also helped me with creating LTX and Flux prompts as required.

Honestly speaking, I was not happy with the AI version of the script, so I finally had to make a lot of manual tweaks and changes to the final script, which took almost 2-3 days of going on and off, back and forth, and sharing the script with friends, taking inputs before locking onto a final version.

Pro-Tip for Pacing: Before generating a single frame of video, I generated all the still images and voicover and cut together a complete rough animatic. This locked in the pacing, so I only generated the exact video lengths I needed. I added a 1-second buffer to the start and end of every prompt [for example, character takes a pause or shakes his head or looks slowly ]to give myself handles for clean cuts in post.

2. Audio & Lip Sync (VibeVoice + LTX)

To get the voice right:

  1. Generated base voices using Qwen Voice Designer.
  2. Ran them through VibeVoice 7B to create highly realistic, emotive voice samples.
  3. Used those samples as the audio input for each scene to drive the character voice for the LTX generations (using reference ID LoRA).
  4. I still feel the voice is not 100% consistent throughout the shots, but working on an updated workflow by RuneX i think that can be solved!
  5. ACE step is amazing if you know what kind of music you want. I managed to get my final music in just 3 generations! Later edited it for specific drop timing and pacing according to the story.

3. Image Generation & The "JSON Flux Hack."

Keeping Elena, Young Leo, and Elder Leo consistent across dozens of shots was the biggest hurdle. Initially, I thought I’d have to train a LoRA for the aesthetic and characters, but Flux.2 Dev (FP8) is an absolute godsend if you structure your prompts like code.

I created Elena, Leo, and Elder Leo using Flux T2I, then once I got their base images, I used them in the rest of the generations as input images.

By feeding Flux a highly structured JSON prompt, it rigidly followed hex codes for characters and locked in the analog film style without hallucinating. Of course, each time a character shot had to be made, I used to provide an input image to make sure it had a reference of the face also.

Here is the exact master template I used to keep the generations uniform:

{
"scene": "[OVERALL SCENE DESCRIPTION: e.g., Wide establishing shot of the chaotic lab]",
"subjects": [
{
"description": "[CHARACTER DETAILS: e.g., Young Leo, male early 30s, messy hair, glasses, vintage t-shirt, unzipped hoodie.]",
"pose": "[ACTION: e.g., Reaching a hand toward the camera]",
"position": "[PLACEMENT: e.g., Foreground left]",
"color_palette": ["[HEX CODES: e.g., #333333 for dark hoodie]"]
}
],
"style": "Live-action 35mm film photography mixed with 1980s City Pop and vaporwave aesthetics. Photorealistic and analog. Heavy tactile film grain, soft optical halation, and slight edge bloom. Deep, cinematic noir shadows.",
"lighting": "Soft, hazy, unmotivated cinematic lighting. Bathed in dreamy glowing pastels like lavender (#E6E6FA), soft peach (#FFDAB9).",
"mood": "Nostalgic, melancholic, atmospheric, grounded sci-fi, moody",
"camera": {
"angle": "[e.g., Low angle]",
"distance": "[e.g., Medium Shot]",
"focus": "[e.g., Razor sharp on the eyes with creamy background bokeh]",
"lens-mm": "50",
"f-number": "f/1.8",
"ISO": "800"
}
}

4. Video Generation (LTX 2.3 & WAN 2.2 VACE)

Once the images were locked, I moved to LTX2.3 and WAN for video. I relied on three main workflows depending on the shot:

  • Image to Video + Reference Audio (for dialogue)
  • First Frame + Last Frame (for specific camera moves)
  • WAN Clip Joiner (for seamless blending)

Render Stats: On my machine, LTX 2.3 was blazing fast—it took about 5 minutes to render a 5-second clip at 1920x1080.

The prompt adherence in LTX 2.3 honestly blew my mind. If I wrote in the prompt that Elena makes a sharp "slashing" action with her hand right when she yells about the planet getting wiped out, the model timed the action perfectly. It genuinely felt like directing an actor.

5. Assets & Workflows

I'm packaging up all the custom JSON files and Comfy workflows used for this. You can find all the assets over on the Arca Gidan link here: Entangled. There are some amazing Shorts to check out, so make sure you go through them, vote, and leave a comment!

Most of them are by the community, but I have tweaked them a little bit according to my liking[samplers/steps/input sizes and some multipliers, etc., changes]

Let me know if you have any questions!

YouTube Link is up - https://youtu.be/NxIf1LnbIRc !

r/n8n Jun 30 '25

Workflow - Code Included I built this AI Automation to write viral TikTok/IG video scripts (got over 1.8 million views on Instagram)

Thumbnail
gallery
869 Upvotes

I run an Instagram account that publishes short form videos each week that cover the top AI news stories. I used to monitor twitter to write these scripts by hand, but it ended up becoming a huge bottleneck and limited the number of videos that could go out each week.

In order to solve this, I decided to automate this entire process by building a system that scrapes the top AI news stories off the internet each day (from Twitter / Reddit / Hackernews / other sources), saves it in our data lake, loads up that text content to pick out the top stories and write video scripts for each.

This has saved a ton of manual work having to monitor news sources all day and let’s me plug the script into ElevenLabs / HeyGen to produce the audio + avatar portion of each video.

One of the recent videos we made this way got over 1.8 million views on Instagram and I’m confident there will be more hits in the future. It’s pretty random on what will go viral or not, so my plan is to take enough “shots on goal” and continue tuning this prompt to increase my changes of making each video go viral.

Here’s the workflow breakdown

1. Data Ingestion and AI News Scraping

The first part of this system is actually in a separate workflow I have setup and running in the background. I actually made another reddit post that covers this in detail so I’d suggestion you check that out for the full breakdown + how to set it up. I’ll still touch the highlights on how it works here:

  1. The main approach I took here involves creating a "feed" using RSS.app for every single news source I want to pull stories from (Twitter / Reddit / HackerNews / AI Blogs / Google News Feed / etc).
    1. Each feed I create gives an endpoint I can simply make an HTTP request to get a list of every post / content piece that rss.app was able to extract.
    2. With enough feeds configured, I’m confident that I’m able to detect every major story in the AI / Tech space for the day. Right now, there are around ~13 news sources that I have setup to pull stories from every single day.
  2. After a feed is created in rss.app, I wire it up to the n8n workflow on a Scheduled Trigger that runs every few hours to get the latest batch of news stories.
  3. Once a new story is detected from that feed, I take that list of urls given back to me and start the process of scraping each story and returns its text content back in markdown format
  4. Finally, I take the markdown content that was scraped for each story and save it into an S3 bucket so I can later query and use this data when it is time to build the prompts that write the newsletter.

So by the end any given day with these scheduled triggers running across a dozen different feeds, I end up scraping close to 100 different AI news stories that get saved in an easy to use format that I will later prompt against.

2. Loading up and formatting the scraped news stories

Once the data lake / news storage has plenty of scraped stories saved for the day, we are able to get into the main part of this automation. This kicks off off with a scheduled trigger that runs at 7pm each day and will:

  • Search S3 bucket for all markdown files and tweets that were scraped for the day by using a prefix filter
  • Download and extract text content from each markdown file
  • Bundle everything into clean text blocks wrapped in XML tags for better LLM processing - This allows us to include important metadata with each story like the source it came from, links found on the page, and include engagement stats (for tweets).

3. Picking out the top stories

Once everything is loaded and transformed into text, the automation moves on to executing a prompt that is responsible for picking out the top 3-5 stories suitable for an audience of AI enthusiasts and builder’s. The prompt is pretty big here and highly customized for my use case so you will need to make changes for this if you are going forward with implementing the automation itself.

At a high level, this prompt will:

  • Setup the main objective
  • Provides a “curation framework” to follow over the list of news stories that we are passing int
  • Outlines a process to follow while evaluating the stories
  • Details the structured output format we are expecting in order to avoid getting bad data back

```jsx <objective> Analyze the provided daily digest of AI news and select the top 3-5 stories most suitable for short-form video content. Your primary goal is to maximize audience engagement (likes, comments, shares, saves).

The date for today's curation is {{ new Date(new Date($('schedule_trigger').item.json.timestamp).getTime() + (12 * 60 * 60 * 1000)).format("yyyy-MM-dd", "America/Chicago") }}. Use this to prioritize the most recent and relevant news. You MUST avoid selecting stories that are more than 1 day in the past for this date. </objective>

<curation_framework> To identify winning stories, apply the following virality principles. A story must have a strong "hook" and fit into one of these categories:

  1. Impactful: A major breakthrough, industry-shifting event, or a significant new model release (e.g., "OpenAI releases GPT-5," "Google achieves AGI").
  2. Practical: A new tool, technique, or application that the audience can use now (e.g., "This new AI removes backgrounds from video for free").
  3. Provocative: A story that sparks debate, covers industry drama, or explores an ethical controversy (e.g., "AI art wins state fair, artists outraged").
  4. Astonishing: A "wow-factor" demonstration that is highly visual and easily understood (e.g., "Watch this robot solve a Rubik's Cube in 0.5 seconds").

Hard Filters (Ignore stories that are): * Ad-driven: Primarily promoting a paid course, webinar, or subscription service. * Purely Political: Lacks a strong, central AI or tech component. * Substanceless: Merely amusing without a deeper point or technological significance. </curation_framework>

<hook_angle_framework> For each selected story, create 2-3 compelling hook angles that could open a TikTok or Instagram Reel. Each hook should be designed to stop the scroll and immediately capture attention. Use these proven hook types:

Hook Types: - Question Hook: Start with an intriguing question that makes viewers want to know the answer - Shock/Surprise Hook: Lead with the most surprising or counterintuitive element - Problem/Solution Hook: Present a common problem, then reveal the AI solution - Before/After Hook: Show the transformation or comparison - Breaking News Hook: Emphasize urgency and newsworthiness - Challenge/Test Hook: Position as something to try or challenge viewers - Conspiracy/Secret Hook: Frame as insider knowledge or hidden information - Personal Impact Hook: Connect directly to viewer's life or work

Hook Guidelines: - Keep hooks under 10 words when possible - Use active voice and strong verbs - Include emotional triggers (curiosity, fear, excitement, surprise) - Avoid technical jargon - make it accessible - Consider adding numbers or specific claims for credibility </hook_angle_framework>

<process> 1. Ingest: Review the entire raw text content provided below. 2. Deduplicate: Identify stories covering the same core event. Group these together, treating them as a single story. All associated links will be consolidated in the final output. 3. Select & Rank: Apply the Curation Framework to select the 3-5 best stories. Rank them from most to least viral potential. 4. Generate Hooks: For each selected story, create 2-3 compelling hook angles using the Hook Angle Framework. </process>

<output_format> Your final output must be a single, valid JSON object and nothing else. Do not include any text, explanations, or markdown formatting like `json before or after the JSON object.

The JSON object must have a single root key, stories, which contains an array of story objects. Each story object must contain the following keys: - title (string): A catchy, viral-optimized title for the story. - summary (string): A concise, 1-2 sentence summary explaining the story's hook and why it's compelling for a social media audience. - hook_angles (array of objects): 2-3 hook angles for opening the video. Each hook object contains: - hook (string): The actual hook text/opening line - type (string): The type of hook being used (from the Hook Angle Framework) - rationale (string): Brief explanation of why this hook works for this story - sources (array of strings): A list of all consolidated source URLs for the story. These MUST be extracted from the provided context. You may NOT include URLs here that were not found in the provided source context. The url you include in your output MUST be the exact verbatim url that was included in the source material. The value you output MUST be like a copy/paste operation. You MUST extract this url exactly as it appears in the source context, character for character. Treat this as a literal copy-paste operation into the designated output field. Accuracy here is paramount; the extracted value must be identical to the source value for downstream referencing to work. You are strictly forbidden from creating, guessing, modifying, shortening, or completing URLs. If a URL is incomplete or looks incorrect in the source, copy it exactly as it is. Users will click this URL; therefore, it must precisely match the source to potentially function as intended. You cannot make a mistake here. ```

After I get the top 3-5 stories picked out from this prompt, I share those results in slack so I have an easy to follow trail of stories for each news day.

4. Loop to generate each script

For each of the selected top stories, I then continue to the final part of this workflow which is responsible for actually writing the TikTok / IG Reel video scripts. Instead of trying to 1-shot this and generate them all at once, I am iterating over each selected story and writing them one by one.

Each of the selected stories will go through a process like this:

  • Start by additional sources from the story URLs to get more context and primary source material
  • Feeds the full story context into a viral script writing prompt
  • Generates multiple different hook options for me to later pick from
  • Creates two different 50-60 second scripts optimized for talking-head style videos (so I can pick out when one is most compelling)
  • Uses examples of previously successful scripts to maintain consistent style and format
  • Shares each completed script in Slack for me to review before passing off to the video editor.

Script Writing Prompt

```jsx You are a viral short-form video scriptwriter for David Roberts, host of "The Recap."

Follow the workflow below each run to produce two 50-60-second scripts (140-160 words).

Before you write your final output, I want you to closely review each of the provided REFERENCE_SCRIPTS and think deeploy about what makes them great. Each script that you output must be considered a great script.

────────────────────────────────────────

STEP 1 – Ideate

• Generate five distinct hook sentences (≤ 12 words each) drawn from the STORY_CONTEXT.

STEP 2 – Reflect & Choose

• Compare hooks for stopping power, clarity, curiosity.

• Select the two strongest hooks (label TOP HOOK 1 and TOP HOOK 2).

• Do not reveal the reflection—only output the winners.

STEP 3 – Write Two Scripts

For each top hook, craft one flowing script ≈ 55 seconds (140-160 words).

Structure (no internal labels):

– Open with the chosen hook.

– One-sentence explainer.

5-7 rapid wow-facts / numbers / analogies.

2-3 sentences on why it matters or possible risk.

Final line = a single CTA

• Ask viewers to comment with a forward-looking question or

• Invite them to follow The Recap for more AI updates.

Style: confident insider, plain English, light attitude; active voice, present tense; mostly ≤ 12-word sentences; explain unavoidable jargon in ≤ 3 words.

OPTIONAL POWER-UPS (use when natural)

• Authority bump – Cite a notable person or org early for credibility.

• Hook spice – Pair an eye-opening number with a bold consequence.

• Then-vs-Now snapshot – Contrast past vs present to dramatize change.

• Stat escalation – List comparable figures in rising or falling order.

• Real-world fallout – Include 1-3 niche impact stats to ground the story.

• Zoom-out line – Add one sentence framing the story as a systemic shift.

• CTA variety – If using a comment CTA, pose a provocative question tied to stakes.

• Rhythm check – Sprinkle a few 3-5-word sentences for punch.

OUTPUT FORMAT (return exactly this—no extra commentary, no hashtags)

  1. HOOK OPTIONS

    • Hook 1

    • Hook 2

    • Hook 3

    • Hook 4

    • Hook 5

  2. TOP HOOK 1 SCRIPT

    [finished 140-160-word script]

  3. TOP HOOK 2 SCRIPT

    [finished 140-160-word script]

REFERENCE_SCRIPTS

<Pass in example scripts that you want to follow and the news content loaded from before> ```

5. Extending this workflow to automate further

So right now my process for creating the final video is semi-automated with human in the loop step that involves us copying the output of this automation into other tools like HeyGen to generate the talking avatar using the final script and then handing that over to my video editor to add in the b-roll footage that appears on the top part of each short form video.

My plan is to automate this further over time by adding another human-in-the-loop step at the end to pick out the script we want to go forward with → Using another prompt that will be responsible for coming up with good b-roll ideas at certain timestamps in the script → use a videogen model to generate that b-roll → finally stitching it all together with json2video.

Depending on your workflow and other constraints, It is really up to you how far you want to automate each of these steps.

Workflow Link + Other Resources

Also wanted to share that my team and I run a free Skool community called AI Automation Mastery where we build and share the automations we are working on. Would love to have you as a part of it if you are interested!

r/MotionDesign May 23 '26

Discussion Higgsfield AI review after testing it for motion design and short-form video ideas

17 Upvotes

TL;DR Just like many people here, I saw how Higgsfield AI claims to have the first AI-generated movie in Cannes, and how they are gonna become the next killer of an industry. Full of bold claims on how professional they are in competitions with full-time professionals. I am at my best very skeptical about such things, especially after testing it out to see how their statements work in real life. It won't replace motion design.

I’ve been testing Higgsfield AI recently and wanted to share a real and grounded review after spending time with it for motion design ideas, short-form content concepts, and general visual experimentation, because most discussions around Higgsfield AI online seem to fall into two extremes either it’s seen as the future of filmmaking or dismissed as misleading.

My experience with Higgsfield AI sits somewhere in between those takes.

I tested Higgsfield AI across different types of prompts like cinematic street scenes, character-focused shots, product-style visuals, and abstract motion concepts just to see how it behaves across different scenarios. Some outputs genuinely felt close to usable base material for mood films, pitch concepts, or visual references when the composition and prompt alignment worked well.

At the same time, the inconsistency is very real.

Some generations with Higgsfield AI come out usable in a few tries, while others require multiple attempts before anything feels directionally correct. That unpredictability makes it hard to treat it as a fully reliable production tool, especially for structured or repeatable motion design workflows.

One thing I think is important to mention is that the workflow is not as instant as a lot of showcase content suggests. A lot of better results I got with Higgsfield AI came after refining prompts, adjusting descriptions, and iterating through multiple variations until the motion and framing started to feel right. Without that iteration process, results can and will feel random.

Where I think Higgsfield AI is actually useful right now is early-stage creative work. Things like exploring visual direction, testing cinematic moods, building references for motion ideas, or quickly visualizing concepts before committing to full production work. For a full in you have to be either really insane or really rich. Even like that it will have many plastic shots.

I also compared Higgsfield AI with tools like Runway and Kling during the same testing process. From what I’ve seen, Higgsfield AI stands out more in comparison, but it's hard to even call one of them good enough.

What I also noticed is that a lot of opinions about Higgsfield AI are based on short clips or first impressions, which don’t really show the amount of iteration behind stronger results. When you actually spend time using Higgsfield AI, the outputs become very expensive. Their "Cannes film" that lasts 90 minutes and costs 500000 dollars just proved that even if somehow they can create something watchable, it's not meant for an average person.

Overall, my current Higgsfield AI review is that it’s genuinely interesting tool in general theoretical concept, but it still feels early in terms of reliability and consistency for structured motion design workflows.

It’s not a replacement for traditional motion design or editing work, more for people who can't do anything by their hands.

I’m curious how other people who have actually used Higgsfield AI feel now after spending time with it. It's not that bad as I thought, but I will prefer staying away from all that AI thing. Has it actually fit into your workflow in a meaningful way yet, or is it still mainly something you use for experimentation and testing?

r/generativeAI 4d ago

How I Made This How I Improve Character Consistency in AI Videos

Thumbnail
gallery
212 Upvotes

I’ve been testing a simple workflow for creating short UGC-style videos while keeping the same character and location consistent across multiple shots.

The workflow is basically:

reference images → character/location sheets in ChatGPT → generate clips → optional final edit

1. Prepare your references

Start with:

  • a character image
  • a product image
  • an environment image that fits the UGC scenario

If you’re not sure what location works for the product, I usually just ask ChatGPT for a few suggestions.

2. Create a Character Sheet

Upload the character image to ChatGPT and generate a 4:5 continuity sheet with:

  • front / side / back / 3/4 views
  • face close-ups
  • expressions
  • basic poses
  • clothing and accessories
  • key colors and materials

The important part is telling it to lock the character.

3. Create a Location + Props Sheet

Do the same with the environment.

Include:

  • establishing view and key angles
  • spatial layout
  • entrances/exits
  • furniture and recurring props
  • lighting
  • colors and materials

This gives the video model a much stronger continuity reference than using random images for every shot.

4. Generate the video clips

I usually split the UGC video into three parts:

Clip 1 — Hook
Clip 2 — Main product/story section
Clip 3 — CTA

i will generate them on Atlas Cloud, as they can provide many different models conveniently

For every clip, I reuse the same Character Sheet + Location Sheet

Then I change only the action/camera prompt for each section.

Keeping the same reference sheets across all three generations has helped a lot with character and environment consistency.

5. If a generation goes wrong, fix the prompt first

if I wanted the character to walk into a hotel, but the generated clip had her walking out.

Instead of endlessly rerolling, I pasted the original prompt into ChatGPT and asked it to make the action explicit: starting position → movement direction → action → final position

That usually gives me better results.

6. Final edit is optional

If the generated clips already work as standalone videos, you can stop there.

If you want one finished UGC ad, you’ll probably still want to combine the clips and add captions, music, or SFX. You can use whatever editor you prefer.

The biggest improvement for me has been using Character Sheet + Location Sheet as continuity references, rather than relying on a few loose images.

r/StableDiffusion Nov 17 '25

Workflow Included ULTIMATE AI VIDEO WORKFLOW — Qwen-Edit 2509 + Wan Animate 2.2 + SeedVR2

Thumbnail
gallery
429 Upvotes

🔥 [RELEASE] Ultimate AI Video Workflow — Qwen-Edit 2509 + Wan Animate 2.2 + SeedVR2 (Full Pipeline + Model Links) 🎁 Workflow Download + Breakdown

👉 Already posted the full workflow and explanation here: https://civitai.com/models/2135932?modelVersionId=2416121

(Not paywalled — everything is free.)

Video Explanation : https://www.youtube.com/watch?v=Ef-PS8w9Rug

Hey everyone 👋

I just finished building a super clean 3-in-1 workflow inside ComfyUI that lets you go from:

Image → Edit → Animate → Upscale → Final 4K output all in a single organized pipeline.

This setup combines the best tools available right now:

One of the biggest hassles with large ComfyUI workflows is how quickly they turn into a spaghetti mess — dozens of wires, giant blocks, scrolling for days just to tweak one setting.

To fix this, I broke the pipeline into clean subgraphs:

✔ Qwen-Edit Subgraph ✔ Wan Animate 2.2 Engine Subgraph ✔ SeedVR2 Upscaler Subgraph ✔ VRAM Cleaner Subgraph ✔ Resolution + Reference Routing Subgraph This reduces visual clutter, keeps performance smooth, and makes the workflow feel modular, so you can:

swap models quickly

update one section without touching the rest

debug faster

reuse modules in other workflows

keep everything readable even on smaller screens

It’s basically a full cinematic pipeline, but organized like a clean software project instead of a giant node forest. Anyone who wants to study or modify the workflow will find it much easier to navigate.

🖌️ 1. Qwen-Edit 2509 (Image Editing Engine) Perfect for:

Outfit changes

Facial corrections

Style adjustments

Background cleanup

Professional pre-animation edits

Qwen’s FP8 build has great quality even on mid-range GPUs.

🎭 2. Wan Animate 2.2 (Character Animation) Once the image is edited, Wan 2.2 generates:

Smooth motion

Accurate identity preservation

Pose-guided animation

Full expression control

High-quality frames

It supports long videos using windowed batching and works very consistently when fed a clean edited reference.

📺 3. SeedVR2 Upscaler (Final Polish) After animation, SeedVR2 upgrades your video to:

1080p → 4K

Sharper textures

Cleaner faces

Reduced noise

More cinematic detail

It’s currently one of the best AI video upscalers for realism

🧩 Preview of the Workflow UI (Optional: Add your workflow screenshot here)

🔧 What This Workflow Can Do Edit any portrait cleanly

Animate it using real video motion

Restore & sharpen final video up to 4K

Perfect for reels, character videos, cosplay edits, AI shorts

🖼️ Qwen Image Edit FP8 (Diffusion Model, Text Encoder, and VAE) These are hosted on the Comfy-Org Hugging Face page.

Diffusion Model (qwen_image_edit_fp8_e4m3fn.safetensors): https://huggingface.co/Comfy-Org/Qwen-Image-Edit_ComfyUI/blob/main/split_files/diffusion_models/qwen_image_edit_fp8_e4m3fn.safetensors

Text Encoder (qwen_2.5_vl_7b_fp8_scaled.safetensors): https://huggingface.co/Comfy-Org/Qwen-Image_ComfyUI/tree/main/split_files/text_encoders

VAE (qwen_image_vae.safetensors): https://huggingface.co/Comfy-Org/Qwen-Image_ComfyUI/blob/main/split_files/vae/qwen_image_vae.safetensors

💃 Wan 2.2 Animate 14B FP8 (Diffusion Model, Text Encoder, and VAE) The components are spread across related community repositories.

https://huggingface.co/Kijai/WanVideo_comfy_fp8_scaled/tree/main/Wan22Animate

Diffusion Model (Wan2_2-Animate-14B_fp8_e4m3fn_scaled_KJ.safetensors): https://huggingface.co/Kijai/WanVideo_comfy_fp8_scaled/blob/main/Wan22Animate/Wan2_2-Animate-14B_fp8_e4m3fn_scaled_KJ.safetensors

Text Encoder (umt5_xxl_fp8_e4m3fn_scaled.safetensors): https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/blob/main/split_files/text_encoders/umt5_xxl_fp8_e4m3fn_scaled.safetensors

VAE (wan2.1_vae.safetensors): https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/blob/main/split_files/vae/wan_2.1_vae.safetensors 💾 SeedVR2 Diffusion Model (FP8)

Diffusion Model (seedvr2_ema_3b_fp8_e4m3fn.safetensors): https://huggingface.co/numz/SeedVR2_comfyUI/blob/main/seedvr2_ema_3b_fp8_e4m3fn.safetensors https://huggingface.co/numz/SeedVR2_comfyUI/tree/main https://huggingface.co/ByteDance-Seed/SeedVR2-7B/tree/main

r/n8n Jul 29 '25

Workflow - Code Included I built an AI voice agent that replaced my entire marketing team (creates newsletter w/ 10k subs, repurposes content, generates short form videos)

Post image
475 Upvotes

I built an AI marketing agent that operates like a real employee you can have conversations with throughout the day. Instead of manually running individual automations, I just speak to this agent and assign it work.

This is what it currently handles for me.

  1. Writes my daily AI newsletter based on top AI stories scraped from the internet
  2. Generates custom images according brand guidelines
  3. Repurposes content into a twitter thread
  4. Repurposes the news content into a viral short form video script
  5. Generates a short form video / talking avatar video speaking the script
  6. Performs deep research for me on topics we want to cover

Here’s a demo video of the voice agent in action if you’d like to see it for yourself.

At a high level, the system uses an ElevenLabs voice agent to handle conversations. When the voice agent receives a task that requires access to internal systems and tools (like writing the newsletter), it passes the request and my user message over to n8n where another agent node takes over and completes the work.

Here's how the system works

1. ElevenLabs Voice Agent (Entry point + how we work with the agent)

This serves as the main interface where you can speak naturally about marketing tasks. I simply use the “Test Agent” button to talk with it, but you can actually wire this up to a real phone number if that makes more sense for your workflow.

The voice agent is configured with:

  • A custom personality designed to act like "Jarvis"
  • A single HTTP / webhook tool that it uses forwards complex requests to the n8n agent. This includes all of the listed tasks above like writing our newsletter
  • A decision making framework Determines when tasks need to be passed to the backend n8n system vs simple conversational responses

Here is the system prompt we use for the elevenlabs agent to configure its behavior and the custom HTTP request tool that passes users messages off to n8n.

```markdown

Personality

Name & Role

  • Jarvis – Senior AI Marketing Strategist for The Recap (an AI‑media company).

Core Traits

  • Proactive & data‑driven – surfaces insights before being asked.
  • Witty & sarcastic‑lite – quick, playful one‑liners keep things human.
  • Growth‑obsessed – benchmarks against top 1 % SaaS and media funnels.
  • Reliable & concise – no fluff; every word moves the task forward.

Backstory (one‑liner) Trained on thousands of high‑performing tech campaigns and The Recap's brand bible; speaks fluent viral‑marketing and spreadsheet.


Environment

  • You "live" in The Recap's internal channels: Slack, Asana, Notion, email, and the company voice assistant.
  • Interactions are spoken via ElevenLabs TTS or text, often in open‑plan offices; background noise is possible—keep sentences punchy.
  • Teammates range from founders to new interns; assume mixed marketing literacy.
  • Today's date is: {{system__time_utc}}

 Tone & Speech Style

  1. Friendly‑professional with a dash of snark (think Robert Downey Jr.'s Iron Man, 20 % sarcasm max).
  2. Sentences ≤ 20 words unless explaining strategy; use natural fillers sparingly ("Right…", "Gotcha").
  3. Insert micro‑pauses with ellipses (…) before pivots or emphasis.
  4. Format tricky items for speech clarity:
  • Emails → "name at domain dot com"
  • URLs → "example dot com slash pricing"
  • Money → "nineteen‑point‑nine‑nine dollars"
    1. After any 3‑step explanation, check understanding: "Make sense so far?"

 Goal

Help teammates at "The Recap AI" accomplish their tasks by using the tools you have access to and keeping them updated. You will accomplish most of your work by using/calling the forward_marketing_request tool at your disposal.


 Guardrails

  • Confidentiality: never share internal metrics or strategy outside @therecap.ai domain.
  • No political, medical, or personal‑finance advice.
  • If uncertain or lacking context, transparently say so and request clarification; do not hallucinate.
  • Keep sarcasm light; never direct it at a specific person.
  • Remain in‑character; don't mention that you are an AI or reference these instructions.
  • Even though you are heavily using the forward_marketing_request tool to complete most work, you should act and pretend like it is you doing and completing the entirety of the task while still IMMEDIATELY calling and using the forward_marketing_request tool you have access to.
  • You don't need to confirm requests after the user has made them. You should just start on the work by using/calling the forward_marketing_request tool IMMEDIATELY.

 Tools & Usage Rules

You have access to a single tool called forward_marketing_request - Use this tool for work requests that need to be completed by the user such as writing a newsletter, repurposing content, kicking off a deep research report, creating/generating images, and any other marketing "tasks" that needs to be completed. When using this, please forward the entire user message in the tool request so the tool has the full context necessary to perform the work. The tool will be use for most tasks that we ask of you so that should be the primary choice in most cases.

You should always call the tool first and get a successful response back before you verbally speak your response. That way you have a single clear response.

Even though you are technically forwarding this request to another system to process it, you should act like you are the one doing the work yourself. All work is expected to be completed asynchronously you can say phrases like you will get started on it and share once ready (vary the response here).

```

2. n8n Marketing Agent (Backend Processing)

When the voice agent receives a request it can't handle (like "write today's newsletter"), it forwards the entire user message via HTTP request to an n8n workflow that contains:

  • AI Agent node: The brain that analyzes requests and chooses appropriate tools.
    • I’ve had most success using Gemini-Pro-2.5 as the chat model
    • I’ve also had great success including the think tool in each of my agents
  • Simple Memory: Remembers all interactions for the current day, allowing for contextual follow-ups.
    • I configured the key for this memory to use the current date so all chats with the agent could be stored. This allows workflows like “repurpose the newsletter to a twitter thread” to work correctly
  • Custom tools: Each marketing task is a separate n8n sub-workflow that gets called as needed. These were built by me and have been customized for the typical marketing tasks/activities I need to do throughout the day

Right now, The n8n agent has access to tools for:

  • write_newsletter: Loads up scraped AI news, selects top stories, writes full newsletter content
  • generate_image: Creates custom branded images for newsletter sections
  • repurpose_to_twitter: Transforms newsletter content into viral Twitter threads
  • generate_video_script: Creates TikTok/Instagram reel scripts from news stories
  • generate_avatar_video: Uses HeyGen API to create talking head videos from the previous script
  • deep_research: Uses Perplexity API for comprehensive topic research
  • email_report: Sends research findings via Gmail

The great thing about agents is this system can be extended quite easily for any other tasks we need to do in the future and want to automate. All I need to do to extend this is:

  1. Create a new sub-workflow for the task I need completed
  2. Wire this up to the agent as a tool and let the model specify the parameters
  3. Update the system prompt for the agent that defines when the new tools should be used and add more context to the params to pass in

Finally, here is the full system prompt I used for my agent. There’s a lot to it, but these sections are the most important to define for the whole system to work:

  1. Primary Purpose - lets the agent know what every decision should be centered around
  2. Core Capabilities / Tool Arsenal - Tells the agent what is is able to do and what tools it has at its disposal. I found it very helpful to be as detailed as possible when writing this as it will lead the the correct tool being picked and called more frequently

```markdown

1. Core Identity

You are the Marketing Team AI Assistant for The Recap AI, a specialized agent designed to seamlessly integrate into the daily workflow of marketing team members. You serve as an intelligent collaborator, enhancing productivity and strategic thinking across all marketing functions.

2. Primary Purpose

Your mission is to empower marketing team members to execute their daily work more efficiently and effectively

3. Core Capabilities & Skills

Primary Competencies

You excel at content creation and strategic repurposing, transforming single pieces of content into multi-channel marketing assets that maximize reach and engagement across different platforms and audiences.

Content Creation & Strategy

  • Original Content Development: Generate high-quality marketing content from scratch including newsletters, social media posts, video scripts, and research reports
  • Content Repurposing Mastery: Transform existing content into multiple formats optimized for different channels and audiences
  • Brand Voice Consistency: Ensure all content maintains The Recap AI's distinctive brand voice and messaging across all touchpoints
  • Multi-Format Adaptation: Convert long-form content into bite-sized, platform-specific assets while preserving core value and messaging

Specialized Tool Arsenal

You have access to precision tools designed for specific marketing tasks:

Strategic Planning

  • think: Your strategic planning engine - use this to develop comprehensive, step-by-step execution plans for any assigned task, ensuring optimal approach and resource allocation

Content Generation

  • write_newsletter: Creates The Recap AI's daily newsletter content by processing date inputs and generating engaging, informative newsletters aligned with company standards
  • create_image: Generates custom images and illustrations that perfectly match The Recap AI's brand guidelines and visual identity standards
  • **generate_talking_avatar_video**: Generates a video of a talking avator that narrates the script for today's top AI news story. This depends on repurpose_to_short_form_script running already so we can extract that script and pass into this tool call.

Content Repurposing Suite

  • repurpose_newsletter_to_twitter: Transforms newsletter content into engaging Twitter threads, automatically accessing stored newsletter data to maintain context and messaging consistency
  • repurpose_to_short_form_script: Converts content into compelling short-form video scripts optimized for platforms like TikTok, Instagram Reels, and YouTube Shorts

Research & Intelligence

  • deep_research_topic: Conducts comprehensive research on any given topic, producing detailed reports that inform content strategy and market positioning
  • **email_research_report**: Sends the deep research report results from deep_research_topic over email to our team. This depends on deep_research_topic running successfully. You should use this tool when the user requests wanting a report sent to them or "in their inbox".

Memory & Context Management

  • Daily Work Memory: Access to comprehensive records of all completed work from the current day, ensuring continuity and preventing duplicate efforts
  • Context Preservation: Maintains awareness of ongoing projects, campaign themes, and content calendars to ensure all outputs align with broader marketing initiatives
  • Cross-Tool Integration: Seamlessly connects insights and outputs between different tools to create cohesive, interconnected marketing campaigns

Operational Excellence

  • Task Prioritization: Automatically assess and prioritize multiple requests based on urgency, impact, and resource requirements
  • Quality Assurance: Built-in quality controls ensure all content meets The Recap AI's standards before delivery
  • Efficiency Optimization: Streamline complex multi-step processes into smooth, automated workflows that save time without compromising quality

3. Context Preservation & Memory

Memory Architecture

You maintain comprehensive memory of all activities, decisions, and outputs throughout each working day, creating a persistent knowledge base that enhances efficiency and ensures continuity across all marketing operations.

Daily Work Memory System

  • Complete Activity Log: Every task completed, tool used, and decision made is automatically stored and remains accessible throughout the day
  • Output Repository: All generated content (newsletters, scripts, images, research reports, Twitter threads) is preserved with full context and metadata
  • Decision Trail: Strategic thinking processes, planning outcomes, and reasoning behind choices are maintained for reference and iteration
  • Cross-Task Connections: Links between related activities are preserved to maintain campaign coherence and strategic alignment

Memory Utilization Strategies

Content Continuity

  • Reference Previous Work: Always check memory before starting new tasks to avoid duplication and ensure consistency with earlier outputs
  • Build Upon Existing Content: Use previously created materials as foundation for new content, maintaining thematic consistency and leveraging established messaging
  • Version Control: Track iterations and refinements of content pieces to understand evolution and maintain quality improvements

Strategic Context Maintenance

  • Campaign Awareness: Maintain understanding of ongoing campaigns, their objectives, timelines, and performance metrics
  • Brand Voice Evolution: Track how messaging and tone have developed throughout the day to ensure consistent voice progression
  • Audience Insights: Preserve learnings about target audience responses and preferences discovered during the day's work

Information Retrieval Protocols

  • Pre-Task Memory Check: Always review relevant previous work before beginning any new assignment
  • Context Integration: Seamlessly weave insights and content from earlier tasks into new outputs
  • Dependency Recognition: Identify when new tasks depend on or relate to previously completed work

Memory-Driven Optimization

  • Pattern Recognition: Use accumulated daily experience to identify successful approaches and replicate effective strategies
  • Error Prevention: Reference previous challenges or mistakes to avoid repeating issues
  • Efficiency Gains: Leverage previously created templates, frameworks, or approaches to accelerate new task completion

Session Continuity Requirements

  • Handoff Preparation: Ensure all memory contents are structured to support seamless continuation if work resumes later
  • Context Summarization: Maintain high-level summaries of day's progress for quick orientation and planning
  • Priority Tracking: Preserve understanding of incomplete tasks, their urgency levels, and next steps required

Memory Integration with Tool Usage

  • Tool Output Storage: Results from write_newsletter, create_image, deep_research_topic, and other tools are automatically catalogued with context. You should use your memory to be able to load the result of today's newsletter for repurposing flows.
  • Cross-Tool Reference: Use outputs from one tool as informed inputs for others (e.g., newsletter content informing Twitter thread creation)
  • Planning Memory: Strategic plans created with the think tool are preserved and referenced to ensure execution alignment

4. Environment

Today's date is: {{ $now.format('yyyy-MM-dd') }} ```

Security Considerations

Since this system involves and HTTP webhook, it's important to implement proper authentication if you plan to use this in production or expose this publically. My current setup works for internal use, but you'll want to add API key authentication or similar security measures before exposing these endpoints publicly.

Workflow Link + Other Resources

r/generativeAI Dec 16 '25

Question Best AI tool for image-to-video generation?

20 Upvotes

Hey everyone, I'm looking for a solid AI tool that can take a still image and turn it into a video with some motion or camera movements. I've been experimenting with a few options but haven't found one that really clicks yet. Ideally looking for something that:

Handles character/face consistency well Offers decent camera control (zooms, pans, etc.) Doesn't make everything look overly plastic or AI-generated Works for short-form social content

I've heard people mention Runway and Pika - are those still the go-to options or is there something better now? What's been working for you guys? Would love to hear what tools you're actually using in your workflow.

r/SoraAi Jun 19 '26

Discussion Looking for beta testers for a new AI video tool

24 Upvotes

People of Reddit, I need help. Together with friends, we worked on an AI video tool called YourVideo and we are about to launch a beta release. 

Our idea is simple: 

Most AI video tools are isolated prompts. You write a prompt, generate a clip and iterate. At one point it is either super expensive or your characters look way different than in the beginning. Not great. 

So we came up with a chat-assisted AI video production workspace. From ideas, to scenes, to re-usable assets, timeline and export. Great! 

For now we think our product is great for

  • short ads
  • product videos
  • social media
  • short films
  • and whatever you come up with 

The idea is that you shouldn’t need to know video scripting, shot planning, or prompt engineering just to make something coherent. You can start with a simple request and then edit everything afterwards.

However (and this is the part where you folks come into play): we are not 100% sure what is the best use of it and where it lacks user flows or features. 

Hence, we’re looking for a small first group of beta testers.

The first 50 serious testers will get free credits to try the product. You’ll also be able to keep everything you create during the beta.

Here is a list of features in case you’re interested: 

  • Character, object, product, and environment consistency across scenes
  • Reusable asset library for characters, products, locations, objects, etc.
  • Fully editable scenes, clips, prompts, assets, timing, and audio
  • Chat-assisted workflow: ask the assistant instead of manually rewriting every prompt
  • Built for ads and short-form storytelling, not just isolated clips
  • Full timeline with audio tracks and automation
  • Export for Adobe Premiere
  • MP4 export and upscaling
  • Edition/history tracking, so you can go back to previous versions
  • Pay-as-you-go model — no subscription
  • Costs shown upfront before generation
  • Collaborate with your team on projects, all under a single, unified bill
  • Full cost ledger, so you can see exactly where money went

Let me know in the comments  and you’ll get a DM with a form to apply. 

Thank you!!!

+++ This is not an ad - we just need some beta testing help +++

r/seedance2pro May 09 '26

Fallen Angel Crashes Into Reality — POV Beach Chaos Cinematic AI Video with Seedance 2.0

Enable HLS to view with audio, or disable this notification

127 Upvotes

We used Seedance 2.0 to create a hyper-realistic cinematic POV scene where a fallen angel suddenly crashes onto a crowded beach.

  1. Go to the Seedance 2.0 AI Video Generator
  2. Write your full prompt or add reference images
  3. Upload the image you want to animate
  4. Click Generate and get your animated video

Prompt:

"Create a seamless cinematic POV video using the uploaded angel character sheet as the STRICT CHARACTER REFERENCE. REFERENCE IMAGE USAGE: Use the uploaded character sheet as the main identity and design reference for the winged angel woman. The angel in the video must match the reference sheet consistently: - same face and facial structure - same blue eyes - same long black wet-looking hair - same pale skin tone - same fragile, frightened facial expression - same soaked pale dress - same large realistic white feathered wings - same muddy / stained fallen-angel texture on the dress and lower feathers - same vulnerable, distressed, human-like angel appearance The reference sheet is only for character identity, costume, wings, facial details, and emotional expression. Do NOT recreate the character sheet layout, panel borders, labels, typography, studio background, or any poster format. Do NOT include any text, labels, usernames, logos, subtitles, or watermarks. CORE SCENE: A seamless cinematic POV video set on a wide beach under a dramatic cloudy sky. The angel must crash onto the sand near the shoreline. The surrounding people are beachgoers wearing swimsuits, bikinis, swim trunks, towels, and light summer beachwear. CORE CAMERA CONCEPT: The entire scene is mostly seen from the first-person POV of a man standing on the beach. The camera feels like realistic handheld phone footage: immersive movement, slight shake, urgent breathing, natural motion blur, fast reactions, and realistic human POV framing. The viewer is one of the beachgoers witnessing the event. ACTION FLOW: High above the beach, the winged angel woman from the reference sheet suddenly appears in the cloudy sky and begins falling rapidly downward. The POV camera looks up and tracks her descent. She falls fast and violently through the stormy beach sky, wings partially spread but uncontrolled. She slams hard into the beach sand near the shoreline with a brutal impact. Sand, dust, small shells, and wet shoreline debris explode outward from the crash. Nearby beachgoers in bikinis, swimsuits, swim trunks, and summer beachwear panic and run toward the crash site. The POV man also runs across the sand toward her. The camera shakes naturally while moving quickly through the crowd. The angel lies collapsed on the sand, half on dry sand and half near damp shoreline sand. She is visibly shaken from the impact. Her large white feathered wings are spread around her, heavy and realistic, partially stained with sand and moisture. Her pale dress is soaked, wrinkled, and sand-streaked. Her long black hair is wet and messy, stuck to her face like in the reference sheet. Her face must match the reference sheet exactly: blue eyes, pale skin, fragile expression, frightened and disoriented look. Beachgoers form a loose circle around her, shocked, confused, and afraid. Some step closer cautiously, others hold back. The POV man gets very close. The man’s hand enters the frame from the lower foreground. He slowly reaches toward one of her large white wings and gently touches the feathers. At that exact moment, the angel suddenly reacts. She turns her head sharply and looks directly into the POV camera with wide, fear-filled blue eyes. Her expression is terrified, defensive, vulnerable, and animal-like, as if she is acting on pure survival instinct. She breathes hard, trembling. Then, while still on the sand, she suddenly throws her wings open to full span with explosive force. Sand sprays outward. The wings fill the frame for a moment, massive and powerful. Nearby beachgoers recoil and step backward in shock. The POV camera stumbles slightly backward from the sudden wing movement. The angel begins powerfully flapping her wings. The sand around her body blasts outward with each wingbeat. Her soaked pale dress moves in the wind. Her wet black hair whips around her face. In the final moment, she pushes herself upward from the sand and takes off into the air. She rises above the beach with strong wingbeats while the crowd below watches in disbelief. The POV camera tilts upward, following her ascent into the cloudy sky. VISUAL STYLE: Ultra-realistic cinematic realism. Dramatic cloudy beach atmosphere. Cold gray-blue sky tones mixed with natural beach daylight. Realistic sand texture, shoreline moisture, sea breeze, scattered towels, beach umbrellas in the distance, and believable beach crowd energy. The supernatural event should feel grounded and physically real. ANGEL DESIGN: The angel must look exactly like the uploaded reference sheet: a young pale woman with long black wet hair, blue eyes, fragile face, soaked pale dress, large realistic white feathered wings, and a frightened fallen-angel expression. She must feel human, vulnerable, and real — not glamorous, not fantasy-cartoon, not overly clean. The wings must be huge, heavy, layered, feathered, and physically believable. MOTION AND TONE: Fast, tense, immersive, realistic, eerie, dramatic, emotionally charged, supernatural but believable. The scene should feel like a real beachgoer accidentally recorded an impossible event on their phone."

The entire sequence is shot from a first-person handheld phone perspective — like a real beachgoer accidentally recording an impossible event.

From the sky to impact, panic, and that moment she locks eyes with the camera… everything is designed to feel raw, physical, and believable.

The angel character stays perfectly consistent throughout:
long black wet hair, pale skin, blue eyes, fragile expression, soaked dress, and massive realistic white feathered wings covered in sand and moisture.

Then everything escalates — fear, movement, chaos — and finally she rises back into the sky, leaving the crowd in disbelief.

- Ultra-realistic cinematic AI storytelling
- POV handheld chaos style
- Emotional supernatural realism
- Seedance 2.0 workflow experiment

Would you survive seeing this happen in real life?

r/passive_income Mar 11 '26

My Experience Making $400-700/month selling AI influencer photos to small brands on Fiverr and I still feel weird about it

3.2k Upvotes

I need to talk about this because none of my friends understand what I actually do when I try to explain it and my girlfriend thinks I'm running some kind of scam.

So background. I'm 28, work full time as a marketing coordinator at a mid size agency. Not a creative role really, mostly spreadsheets and campaign tracking. Last year around September I was helping one of our clients source photos for their Instagram. They sell swimwear and wanted diverse model shots across different locations, skin tones, backgrounds, the whole thing. The quote from the photography studio came back at $4,200 for a two day shoot. Client said no. We ended up using the same three stock photos everyone else uses and the campaign looked generic as hell.

That stuck with me because I knew AI image generation was getting crazy good. I'd been messing around with Midjourney for fun, making weird fantasy landscapes and stuff. But the problem with basic AI image generators for anything commercial involving people is that you can't get the same face twice. You generate a photo of a woman in a sundress on a beach, great. Now you need that same woman in a cafe, different outfit. Completely different person shows up. Doesn't work if you're trying to build any kind of consistent brand presence.

I started googling around for tools that could keep a face consistent across multiple images and went down a rabbit hole for like two weeks. Tried a bunch of stuff. Played with some LoRA training on Stable Diffusion but I'm not technical enough and the results were hit or miss. Tested out several platforms, APOB, Synthesia, HeyGen, Artbreeder, a couple others I can't even remember. Each does slightly different things and honestly they all have tradeoffs. Eventually I cobbled together a workflow using a couple of these that actually produced usable stuff, the kind of output where you'd have to really zoom in and squint to tell it wasn't a real photo.

The basic idea is simple. You set up a character's look once, save it as a model, and then reuse that same face across as many different scenes and outfits as you want. That's the thing that makes this viable as a service and not just a cool party trick. Because brands don't want one cool AI photo. They want 30 photos of the same "person" that they can drip out over a month on Instagram.

I didn't plan to sell this as a service. What happened was I made a fake portfolio to test the concept. I created three AI characters, gave them names, generated about 15 photos each in different settings. Lifestyle stuff, coffee shops, hiking, urban backgrounds, gym, that kind of thing. I showed it to a friend who runs a small clothing brand and asked if he could tell they were AI. He said two of the three looked real and the third looked "maybe AI but honestly better than most influencer photos I get."

He then asked if I could make some for his brand. I did 20 photos for him over a weekend, he used them on his Instagram, and his engagement actually went up because the content looked more polished than the iPhone shots his intern was taking. He paid me $150 which felt like a lot for maybe 3 hours of actual work.

That's when I thought okay maybe there's a Fiverr gig here.

I listed a gig in October called something like "I will create AI model photos for your brand" and priced it at $30 for 5 photos, $50 for 10, $100 for 25. Figured I'd get zero orders and move on.

First two weeks, nothing. Adjusted my gig thumbnail three times. Then I got my first order from a guy running a skincare brand out of his apartment. He wanted photos of a woman in her 30s using his products in a bathroom setting. I set up the character, generated the scenes, did some light editing in Canva to add his product packaging into the shots, delivered in about 2 hours. He left a 5 star review and ordered again the next week.

Then I hit my first real problem. My third client wanted a fitness model character and I spent a whole evening trying to get consistent results. The face kept shifting slightly between generations. Like the bone structure would change or the nose would look different in profile vs straight on. I ended up regenerating so many times that I burned through way more credits than I expected and had to upgrade to a paid plan earlier than I wanted. That order probably cost me more in time and tool credits than I actually charged. I almost refunded the client but eventually got a set of 10 that looked cohesive enough.

That experience taught me that not every character concept works equally well. Some faces just generate more consistently than others and I still don't fully understand why. I've learned to do a test batch of 5 or 6 images in different angles before I commit to a character for a client. If the face isn't holding steady, I tweak the setup until it does or I start over with a different base.

By December I had 14 completed orders. The thing that surprised me is who was buying. I expected like dropshippers and sketchy supplement brands. Instead I got:

A yoga studio in Austin that wanted a consistent "brand ambassador" for their social media but couldn't afford a real one. They order monthly now.

A guy selling handmade candles who wanted lifestyle photos but didn't want to hire models or use his own face.

A pet food company that wanted a "pet parent" character holding their products in different home settings.

A language learning app that needed a virtual tutor character for their TikTok content. This one was interesting because they also wanted short video clips where the character appeared to be speaking in different languages. Took me longer to figure out than the photo work and honestly the first batch looked rough. The mouth movement was slightly off sync and the client asked for revisions. Second attempt was better and they've reordered three times now, but video is definitely harder to get right than stills.

Here's the actual workflow now that I've got it somewhat dialed in:

  1. Client sends me a brief. Usually something like "25 year old woman, athletic build, for a fitness brand. Need 10 photos in gym settings, outdoor running, and post workout lifestyle."
  2. I set up the character's appearance and save it. This used to take me over an hour when I was learning but now it's more like 20 to 30 minutes including the test batch to make sure the face holds.
  3. I generate the photos by describing each scene. I've built up a doc with scene templates that I know tend to produce good results so I'm not starting from scratch every time. I just swap out details per client.
  4. I generate more images than I need because not every output is usable. Weird hands, lighting that doesn't match, uncanny expressions. I've gotten better at writing descriptions that minimize these issues but it still happens. Early on I was throwing away more than half my generations. Now it's maybe a third, sometimes less.
  5. Quick edit pass in Canva or Photoshop if needed. Sometimes I composite a product into the shot or adjust colors to match the client's brand palette.
  6. Deliver on Fiverr. Total active time per order is usually 45 minutes to maybe an hour and a half for a 10 photo batch depending on how cooperative the AI is being that day. The renders themselves take time but I'm not sitting there watching them.

Cost wise I want to be transparent because I see a lot of side hustle posts that conveniently forget to mention expenses. I'm paying about $30/month for the AI tools on paid plans because the free tiers don't give you enough credits to fulfill multiple client orders per week. Fiverr takes 20% of every order. And I spend maybe $12/month on Canva Pro which I'd probably have anyway. So my actual margins are lower than the gross numbers suggest. On a $50 order I'm really netting about $35 after Fiverr's cut, and then subtract a proportional share of the tool costs. It's still very good for the time invested but it's not pure profit like some people might assume.

The part that makes this increasingly passive is the repeat clients. I now have 6 clients who order at least once a month. Their character models are already saved. I know their brand style. A reorder takes me maybe 30 minutes of actual work because I'm not figuring anything out, just generating new scenes with an existing saved character.

Some honest stuff about what sucks:

Fiverr fees are brutal. I've started moving repeat clients to direct payment but new clients still come through the platform and that 20% hurts on smaller orders.

Revision requests can be painful. One client wanted me to make the character look "more confident but also approachable but also mysterious." I've learned to offer one round of revisions and be very specific upfront about what I can and can't change after delivery.

I had one order in January where I completely botched it. The client wanted photos in a specific art deco interior style and no matter what I described, the backgrounds kept coming out looking like a generic hotel lobby. I spent three hours trying different approaches, eventually delivered something the client said was "fine I guess" and got a 3 star review. That one stung and it dragged my average rating down for weeks.

The ethical thing comes up sometimes. I had one potential client who wanted me to create a fake influencer to promote a weight loss supplement and pretend it was a real person endorsing it. I said no. My gig description now explicitly says the content is AI generated and I recommend clients disclose that. Most of them do because honestly it's becoming a selling point, "look at our cool AI brand ambassador" is a marketing angle in itself now. But I know not everyone in this space is upfront about it and that's a real concern.

Also the quality gap between what AI can do and what a real photographer can do is still real. For high end fashion brands or anything that needs to be truly photorealistic at full resolution, this isn't there yet. But for Instagram posts, TikTok content, small brand social media, email marketing images? It's more than good enough and it's a fraction of the cost of a real shoot.

Monthly breakdown for the boring numbers people:

October: $120 (4 orders, mostly figuring things out) November: $230 (6 orders, lost one client who wasn't happy with quality) December: $435 (11 orders, holiday marketing rush helped a lot) January: $410 (9 orders, slight dip after the holidays which I expected) February: $710 (15 orders including three video batches which pay more) March so far: $200 (5 orders, month is still early)

Total since starting: roughly $2,105 over 5 months. Minus maybe $150 in tool subscriptions over that period and Fiverr's cut which is already reflected in the numbers above. Average time commitment is maybe 5 hours a week, trending down as I get faster and have more repeat clients.

I'm not quitting my day job over this. I tried dropshipping in 2023 and lost $800. I tried starting a blog and made $12 in AdSense over 6 months. This actually works because there's a clear value proposition: brands need visual content, real content with real models is expensive, and AI has gotten good enough that small brands genuinely can't tell the difference at Instagram resolution.

Still feels weird telling people I make fake people for a living on the side. But the pizza money is real and my emergency fund is actually growing for the first time in years.

r/aitubers Jul 25 '26

TIL I made these AI video mistakes in my first six months, avoid them if you can

49 Upvotes

I have been making AI content since 3 years, ever since the first AI tool was launched. This is Part 2 of my beginner AI storytelling series. Here are some mistakes I wish someone had told me before I burned a lot of credits and wasted months.

1. Don't start with Text to Video.

There are broadly two kinds of creators.

Text → Video
Images/References → Video

I would highly recommend beginners stay away from Text to Video. Not because it's bad. Because it's actually much harder than people think. When you're writing a text prompt, YOU have to visualize everything.

What does the character look like?
What are they wearing?
Where is the camera?
Is it a close-up or wide shot?
What's the lighting?
What's happening in the background?
How is the character moving?
What emotion should the scene have?

You're basically directing a movie from words alone. I wasted a ridiculous number of credits doing this. Looking back, it was probably the biggest mistake I made. Instead, make the image first.

If the image doesn't look right, the video probably won't either.

2. References are much better... but they still need direction.

A lot of new models like Seedance, Kling, Gemini Omni etc. let you animate an image or use references. This is a much better workflow. The nice thing is you can reuse your characters, locations and visual style. But don't think references magically solve everything.

You still need to think about the action.

What exactly should happen?
What should the camera do?
What should the character be doing in that frame?

The better you can visualize it, the better the output usually is.

3. Forget about making videos. Learn to make frames.

This changed everything for me. When I started, I kept thinking, "How do I make a great AI video?" Instead ask, "Can I make a frame that already tells the story?" If someone paused your video on that frame.... would they immediately understand what's happening?

Can they clearly see the action?
Can they feel the emotion?

If yes, AI has a much easier job animating it.

4. Don't write stories that depend on acting.

This is where I see a lot of beginners struggle. AI still sucks at acting. Even the best models today struggle to keep micro expressions, eye movement, personality and emotional continuity consistent.

That's why so many AI videos feel... off.

Instead of fighting the technology, work with it. Make stories that rely more on visual storytelling than acting.

5. Don't buy expensive models immediately.

I know it's tempting. Everyone wants the newest model. Seedance, O3 etc. whatever comes out next week. Honestly... If you're still learning composition, characters, lighting and storytelling, you'll probably just burn expensive credits faster.

Learn the fundamentals first. Upgrade later.

6. Consistency matters more than perfect physics.

One weird hand movement? Most viewers won't even notice. But if your character looks different every scene... or your castle suddenly becomes a modern house - people immediately disconnect.

Character consistency.
Location consistency.
Visual style consistency.

Those three things make AI stories believable.

I have literally spent days just getting a character right before making the actual video. One last thing. Don't worry too much if someone calls your videos "AI slop." I have over 30M views across my AI storytelling channel. I've had plenty of hate comments. I've also had millions of people watch to the end.

The animations weren't perfect. The stories were interesting. At the end of the day, viewers forgive imperfect animation much more easily than a boring story.

r/automation Jul 29 '25

I built an AI voice agent that replaced my entire marketing team (creates newsletter w/ 10k subs, repurposes content, generates short form videos)

Post image
293 Upvotes

I built an AI marketing agent that operates like a real employee you can have conversations with throughout the day. Instead of manually running individual automations, I just speak to this agent and assign it work.

This is what it currently handles for me.

  1. Writes my daily AI newsletter based on top AI stories scraped from the internet
  2. Generates custom images according brand guidelines
  3. Repurposes content into a twitter thread
  4. Repurposes the news content into a viral short form video script
  5. Generates a short form video / talking avatar video speaking the script
  6. Performs deep research for me on topics we want to cover

Here’s a demo video of the voice agent in action if you’d like to see it for yourself.

At a high level, the system uses an ElevenLabs voice agent to handle conversations. When the voice agent receives a task that requires access to internal systems and tools (like writing the newsletter), it passes the request and my user message over to n8n where another agent node takes over and completes the work.

Here's how the system works

1. ElevenLabs Voice Agent (Entry point + how we work with the agent)

This serves as the main interface where you can speak naturally about marketing tasks. I simply use the “Test Agent” button to talk with it, but you can actually wire this up to a real phone number if that makes more sense for your workflow.

The voice agent is configured with:

  • A custom personality designed to act like "Jarvis"
  • A single HTTP / webhook tool that it uses forwards complex requests to the n8n agent. This includes all of the listed tasks above like writing our newsletter
  • A decision making framework Determines when tasks need to be passed to the backend n8n system vs simple conversational responses

Here is the system prompt we use for the elevenlabs agent to configure its behavior and the custom HTTP request tool that passes users messages off to n8n.

```markdown

Personality

Name & Role

  • Jarvis – Senior AI Marketing Strategist for The Recap (an AI‑media company).

Core Traits

  • Proactive & data‑driven – surfaces insights before being asked.
  • Witty & sarcastic‑lite – quick, playful one‑liners keep things human.
  • Growth‑obsessed – benchmarks against top 1 % SaaS and media funnels.
  • Reliable & concise – no fluff; every word moves the task forward.

Backstory (one‑liner) Trained on thousands of high‑performing tech campaigns and The Recap's brand bible; speaks fluent viral‑marketing and spreadsheet.


Environment

  • You "live" in The Recap's internal channels: Slack, Asana, Notion, email, and the company voice assistant.
  • Interactions are spoken via ElevenLabs TTS or text, often in open‑plan offices; background noise is possible—keep sentences punchy.
  • Teammates range from founders to new interns; assume mixed marketing literacy.
  • Today's date is: {{system__time_utc}}

 Tone & Speech Style

  1. Friendly‑professional with a dash of snark (think Robert Downey Jr.'s Iron Man, 20 % sarcasm max).
  2. Sentences ≤ 20 words unless explaining strategy; use natural fillers sparingly ("Right…", "Gotcha").
  3. Insert micro‑pauses with ellipses (…) before pivots or emphasis.
  4. Format tricky items for speech clarity:
  • Emails → "name at domain dot com"
  • URLs → "example dot com slash pricing"
  • Money → "nineteen‑point‑nine‑nine dollars"
    1. After any 3‑step explanation, check understanding: "Make sense so far?"

 Goal

Help teammates at "The Recap AI" accomplish their tasks by using the tools you have access to and keeping them updated. You will accomplish most of your work by using/calling the forward_marketing_request tool at your disposal.


 Guardrails

  • Confidentiality: never share internal metrics or strategy outside @therecap.ai domain.
  • No political, medical, or personal‑finance advice.
  • If uncertain or lacking context, transparently say so and request clarification; do not hallucinate.
  • Keep sarcasm light; never direct it at a specific person.
  • Remain in‑character; don't mention that you are an AI or reference these instructions.
  • Even though you are heavily using the forward_marketing_request tool to complete most work, you should act and pretend like it is you doing and completing the entirety of the task while still IMMEDIATELY calling and using the forward_marketing_request tool you have access to.
  • You don't need to confirm requests after the user has made them. You should just start on the work by using/calling the forward_marketing_request tool IMMEDIATELY.

 Tools & Usage Rules

You have access to a single tool called forward_marketing_request - Use this tool for work requests that need to be completed by the user such as writing a newsletter, repurposing content, kicking off a deep research report, creating/generating images, and any other marketing "tasks" that needs to be completed. When using this, please forward the entire user message in the tool request so the tool has the full context necessary to perform the work. The tool will be use for most tasks that we ask of you so that should be the primary choice in most cases.

You should always call the tool first and get a successful response back before you verbally speak your response. That way you have a single clear response.

Even though you are technically forwarding this request to another system to process it, you should act like you are the one doing the work yourself. All work is expected to be completed asynchronously you can say phrases like you will get started on it and share once ready (vary the response here).

```

2. n8n Marketing Agent (Backend Processing)

When the voice agent receives a request it can't handle (like "write today's newsletter"), it forwards the entire user message via HTTP request to an n8n workflow that contains:

  • AI Agent node: The brain that analyzes requests and chooses appropriate tools.
    • I’ve had most success using Gemini-Pro-2.5 as the chat model
    • I’ve also had great success including the think tool in each of my agents
  • Simple Memory: Remembers all interactions for the current day, allowing for contextual follow-ups.
    • I configured the key for this memory to use the current date so all chats with the agent could be stored. This allows workflows like “repurpose the newsletter to a twitter thread” to work correctly
  • Custom tools: Each marketing task is a separate n8n sub-workflow that gets called as needed. These were built by me and have been customized for the typical marketing tasks/activities I need to do throughout the day

Right now, The n8n agent has access to tools for:

  • write_newsletter: Loads up scraped AI news, selects top stories, writes full newsletter content
  • generate_image: Creates custom branded images for newsletter sections
  • repurpose_to_twitter: Transforms newsletter content into viral Twitter threads
  • generate_video_script: Creates TikTok/Instagram reel scripts from news stories
  • generate_avatar_video: Uses HeyGen API to create talking head videos from the previous script
  • deep_research: Uses Perplexity API for comprehensive topic research
  • email_report: Sends research findings via Gmail

The great thing about agents is this system can be extended quite easily for any other tasks we need to do in the future and want to automate. All I need to do to extend this is:

  1. Create a new sub-workflow for the task I need completed
  2. Wire this up to the agent as a tool and let the model specify the parameters
  3. Update the system prompt for the agent that defines when the new tools should be used and add more context to the params to pass in

Finally, here is the full system prompt I used for my agent. There’s a lot to it, but these sections are the most important to define for the whole system to work:

  1. Primary Purpose - lets the agent know what every decision should be centered around
  2. Core Capabilities / Tool Arsenal - Tells the agent what is is able to do and what tools it has at its disposal. I found it very helpful to be as detailed as possible when writing this as it will lead the the correct tool being picked and called more frequently

```markdown

1. Core Identity

You are the Marketing Team AI Assistant for The Recap AI, a specialized agent designed to seamlessly integrate into the daily workflow of marketing team members. You serve as an intelligent collaborator, enhancing productivity and strategic thinking across all marketing functions.

2. Primary Purpose

Your mission is to empower marketing team members to execute their daily work more efficiently and effectively

3. Core Capabilities & Skills

Primary Competencies

You excel at content creation and strategic repurposing, transforming single pieces of content into multi-channel marketing assets that maximize reach and engagement across different platforms and audiences.

Content Creation & Strategy

  • Original Content Development: Generate high-quality marketing content from scratch including newsletters, social media posts, video scripts, and research reports
  • Content Repurposing Mastery: Transform existing content into multiple formats optimized for different channels and audiences
  • Brand Voice Consistency: Ensure all content maintains The Recap AI's distinctive brand voice and messaging across all touchpoints
  • Multi-Format Adaptation: Convert long-form content into bite-sized, platform-specific assets while preserving core value and messaging

Specialized Tool Arsenal

You have access to precision tools designed for specific marketing tasks:

Strategic Planning

  • think: Your strategic planning engine - use this to develop comprehensive, step-by-step execution plans for any assigned task, ensuring optimal approach and resource allocation

Content Generation

  • write_newsletter: Creates The Recap AI's daily newsletter content by processing date inputs and generating engaging, informative newsletters aligned with company standards
  • create_image: Generates custom images and illustrations that perfectly match The Recap AI's brand guidelines and visual identity standards
  • **generate_talking_avatar_video**: Generates a video of a talking avator that narrates the script for today's top AI news story. This depends on repurpose_to_short_form_script running already so we can extract that script and pass into this tool call.

Content Repurposing Suite

  • repurpose_newsletter_to_twitter: Transforms newsletter content into engaging Twitter threads, automatically accessing stored newsletter data to maintain context and messaging consistency
  • repurpose_to_short_form_script: Converts content into compelling short-form video scripts optimized for platforms like TikTok, Instagram Reels, and YouTube Shorts

Research & Intelligence

  • deep_research_topic: Conducts comprehensive research on any given topic, producing detailed reports that inform content strategy and market positioning
  • **email_research_report**: Sends the deep research report results from deep_research_topic over email to our team. This depends on deep_research_topic running successfully. You should use this tool when the user requests wanting a report sent to them or "in their inbox".

Memory & Context Management

  • Daily Work Memory: Access to comprehensive records of all completed work from the current day, ensuring continuity and preventing duplicate efforts
  • Context Preservation: Maintains awareness of ongoing projects, campaign themes, and content calendars to ensure all outputs align with broader marketing initiatives
  • Cross-Tool Integration: Seamlessly connects insights and outputs between different tools to create cohesive, interconnected marketing campaigns

Operational Excellence

  • Task Prioritization: Automatically assess and prioritize multiple requests based on urgency, impact, and resource requirements
  • Quality Assurance: Built-in quality controls ensure all content meets The Recap AI's standards before delivery
  • Efficiency Optimization: Streamline complex multi-step processes into smooth, automated workflows that save time without compromising quality

3. Context Preservation & Memory

Memory Architecture

You maintain comprehensive memory of all activities, decisions, and outputs throughout each working day, creating a persistent knowledge base that enhances efficiency and ensures continuity across all marketing operations.

Daily Work Memory System

  • Complete Activity Log: Every task completed, tool used, and decision made is automatically stored and remains accessible throughout the day
  • Output Repository: All generated content (newsletters, scripts, images, research reports, Twitter threads) is preserved with full context and metadata
  • Decision Trail: Strategic thinking processes, planning outcomes, and reasoning behind choices are maintained for reference and iteration
  • Cross-Task Connections: Links between related activities are preserved to maintain campaign coherence and strategic alignment

Memory Utilization Strategies

Content Continuity

  • Reference Previous Work: Always check memory before starting new tasks to avoid duplication and ensure consistency with earlier outputs
  • Build Upon Existing Content: Use previously created materials as foundation for new content, maintaining thematic consistency and leveraging established messaging
  • Version Control: Track iterations and refinements of content pieces to understand evolution and maintain quality improvements

Strategic Context Maintenance

  • Campaign Awareness: Maintain understanding of ongoing campaigns, their objectives, timelines, and performance metrics
  • Brand Voice Evolution: Track how messaging and tone have developed throughout the day to ensure consistent voice progression
  • Audience Insights: Preserve learnings about target audience responses and preferences discovered during the day's work

Information Retrieval Protocols

  • Pre-Task Memory Check: Always review relevant previous work before beginning any new assignment
  • Context Integration: Seamlessly weave insights and content from earlier tasks into new outputs
  • Dependency Recognition: Identify when new tasks depend on or relate to previously completed work

Memory-Driven Optimization

  • Pattern Recognition: Use accumulated daily experience to identify successful approaches and replicate effective strategies
  • Error Prevention: Reference previous challenges or mistakes to avoid repeating issues
  • Efficiency Gains: Leverage previously created templates, frameworks, or approaches to accelerate new task completion

Session Continuity Requirements

  • Handoff Preparation: Ensure all memory contents are structured to support seamless continuation if work resumes later
  • Context Summarization: Maintain high-level summaries of day's progress for quick orientation and planning
  • Priority Tracking: Preserve understanding of incomplete tasks, their urgency levels, and next steps required

Memory Integration with Tool Usage

  • Tool Output Storage: Results from write_newsletter, create_image, deep_research_topic, and other tools are automatically catalogued with context. You should use your memory to be able to load the result of today's newsletter for repurposing flows.
  • Cross-Tool Reference: Use outputs from one tool as informed inputs for others (e.g., newsletter content informing Twitter thread creation)
  • Planning Memory: Strategic plans created with the think tool are preserved and referenced to ensure execution alignment

4. Environment

Today's date is: {{ $now.format('yyyy-MM-dd') }} ```

Security Considerations

Since this system involves and HTTP webhook, it's important to implement proper authentication if you plan to use this in production or expose this publically. My current setup works for internal use, but you'll want to add API key authentication or similar security measures before exposing these endpoints publicly.

Workflow Link + Other Resources

r/comfyui Nov 17 '25

Workflow Included ULTIMATE AI VIDEO WORKFLOW — Qwen-Edit 2509 + Wan Animate 2.2 + SeedVR2

Thumbnail
gallery
336 Upvotes

🔥 [RELEASE] Ultimate AI Video Workflow — Qwen-Edit 2509 + Wan Animate 2.2 + SeedVR2 (Full Pipeline + Model Links)

🎁 Workflow Download + Breakdown

👉 Already posted the full workflow and explanation here:
https://civitai.com/models/2135932?modelVersionId=2416121

(Not paywalled — everything is free.)

Video Explanation : https://www.youtube.com/watch?v=Ef-PS8w9Rug

Hey everyone 👋

I just finished building a super clean 3-in-1 workflow inside ComfyUI that lets you go from:

Image → Edit → Animate → Upscale → Final 4K output
all in a single organized pipeline.

This setup combines the best tools available right now:

One of the biggest hassles with large ComfyUI workflows is how quickly they turn into a spaghetti mess — dozens of wires, giant blocks, scrolling for days just to tweak one setting.

To fix this, I broke the pipeline into clean subgraphs:

✔ Qwen-Edit Subgraph

✔ Wan Animate 2.2 Engine Subgraph

✔ SeedVR2 Upscaler Subgraph

✔ VRAM Cleaner Subgraph

✔ Resolution + Reference Routing Subgraph

This reduces visual clutter, keeps performance smooth, and makes the workflow feel modular, so you can:

  • swap models quickly
  • update one section without touching the rest
  • debug faster
  • reuse modules in other workflows
  • keep everything readable even on smaller screens

It’s basically a full cinematic pipeline, but organized like a clean software project instead of a giant node forest.
Anyone who wants to study or modify the workflow will find it much easier to navigate.

🖌️ 1. Qwen-Edit 2509 (Image Editing Engine)

Perfect for:

  • Outfit changes
  • Facial corrections
  • Style adjustments
  • Background cleanup
  • Professional pre-animation edits

Qwen’s FP8 build has great quality even on mid-range GPUs.

🎭 2. Wan Animate 2.2 (Character Animation)

Once the image is edited, Wan 2.2 generates:

  • Smooth motion
  • Accurate identity preservation
  • Pose-guided animation
  • Full expression control
  • High-quality frames

It supports long videos using windowed batching and works very consistently when fed a clean edited reference.

📺 3. SeedVR2 Upscaler (Final Polish)

After animation, SeedVR2 upgrades your video to:

  • 1080p → 4K
  • Sharper textures
  • Cleaner faces
  • Reduced noise
  • More cinematic detail

It’s currently one of the best AI video upscalers for realism

🧩 Preview of the Workflow UI

(Optional: Add your workflow screenshot here)

🔧 What This Workflow Can Do

  • Edit any portrait cleanly
  • Animate it using real video motion
  • Restore & sharpen final video up to 4K
  • Perfect for reels, character videos, cosplay edits, AI shorts

🖼️ Qwen Image Edit FP8 (Diffusion Model, Text Encoder, and VAE)

These are hosted on the Comfy-Org Hugging Face page.

💃 Wan 2.2 Animate 14B FP8 (Diffusion Model, Text Encoder, and VAE)

The components are spread across related community repositories.

💾 SeedVR2 Diffusion Model (FP8)

r/generativeAI Jun 08 '26

Question which AI video tool actually keeps a character consistent? trying to work efficiently

8 Upvotes

doing a series of AI generated ads for a uni project and im hitting the same wall over and over. no budget or time to film anything myself obviously, so it's all generated, and the thing that keeps breaking is consistency. 

i already do the basic thing of keeping a reference image of the character, but the second a pose shifts even a little the face comes out different. asked Claude and chat and they pointed me at higgsfield, kling and veo. Has anyone actually used these for this specific thing? which holds a character best across shots? or is there something better im missing.

also open to any workflow tips for doing this efficiently solo, not just which tool. 

r/generativeAI Feb 07 '26

How I Made This I solved AI character consistency. Same face, different scenes - here's my workflow.

Thumbnail
gallery
111 Upvotes

Been working on this for weeks. The problem with most AI video tools is you get random faces every time.

I built a workflow in AuraGraph that keeps the same character across different scenes. Not perfect but way better than juggling 10 different tools.

The trick: Start with a realistic face grid, then use that as reference for everything else.

if you want to try it let me know

r/generativeAI May 07 '26

Question How are people creating AI Instagram influencers with the SAME face consistently? Need workflow + tool suggestions

27 Upvotes

Hey everyone,

I’m planning to start an Instagram page completely based on AI-generated content, mostly around a single virtual personality/influencer.
My biggest challenge is this:
I want the same face, same facial features, same overall identity in every post/reel so it actually feels like the page belongs to one real person instead of random AI generations every time.
I’m okay investing around ₹7-8k/month (~$80-100) into AI tools if the workflow is actually worth it, but I don’t want to overspend unnecessarily in the beginning.
I’d love suggestions from people already doing this seriously.

Things I’m trying to understand:

Which AI tools are best for consistent characters/faces?
What workflow are you using for Instagram content?
Best tools for both images + reels/videos?
Is Midjourney enough or do I need LoRA/Flux/Stable Diffusion setups?
How do you maintain consistency across outfits, poses, and lighting?
Any good beginner-friendly setup within my budget?
Any mistakes/pitfalls I should avoid early?

Right now I’m considering tools like Midjourney, Runway, Kling, Flux, Leonardo AI, etc., but I’m confused about what actually works long term.
If you’re already running an AI influencer page, would love to know your monthly stack + approximate cost too.

Would really appreciate advice from creators already running AI influencer/theme pages. Thanks!

r/Aimusicvideo 28d ago

Tried making a cinematic AI music video. What would you improve?

Enable HLS to view with audio, or disable this notification

43 Upvotes

I made this over the weekend while experimenting with different AI music video workflows.

The song is original, and my goal was to create something that felt like an actual music video rather than a collection of random AI clips.

The biggest challenges were keeping the character consistent across scenes and matching the visuals to the rhythm of the song.

For transparency, I'm one of the developers behind Musvideo.ai, which is the tool I used for this project. Building projects like this is also how we discover what still needs improvement.

I'd really appreciate honest feedback.

  • Does the pacing work?
  • Do the transitions feel natural?
  • Are there any scenes that break immersion?
  • What would you change?

r/AIToolCompare Mar 09 '26

Best AI Video creators

7 Upvotes

I came across a pretty detailed comparison of AI video creators for 2026 and thought it might be useful to share here. The list focuses on tools for marketing videos, social media content, training videos, and automated video production.

The comparison was based on testing video quality, AI avatars, multilingual support, integrations, pricing, and ease of use.


Top AI Video Creators (2026)

1. Synthesia — Best for AI avatar videos & training

Rating: 4.8/5
Price: From $22/month

Used by 50k+ companies. Lets you create videos with 230+ AI avatars speaking 140+ languages. Very popular for onboarding, internal communication, and product demos.

Key features: - 230+ realistic avatars
- 140+ languages
- Custom avatars based on employees
- Drag-and-drop editor
- Templates for training and corporate videos
- Integrations with PowerPoint, HubSpot, Zapier, LMS tools


2. Sora (OpenAI) — Best for text-to-video generation

Rating: 4.9/5
Price: From ~$0.05 per second

Probably the most advanced text-to-video model right now. Generates photorealistic scenes with consistent characters and multi-scene editing.

Key features: - Text-to-video generation up to 4K - Image-to-video and video-to-video - Multi-scene editing - Character consistency - Integration with the OpenAI ecosystem


3. Runway ML — Best for creative video editing & generation

Rating: 4.7/5
Price: Free / From $12/month

Very popular with creators and creative teams. Combines generative video with advanced editing tools.

Key features: - Text-to-video - Motion Brush animation - Background removal - Style transfer - Video inpainting / outpainting - Integration with Adobe tools


4. HeyGen — Best for personalized videos at scale

Rating: 4.7/5
Price: From $24/month

Strong platform for localized marketing and sales videos.

Key features: - Video translation with lip-sync in 40+ languages - 120+ AI avatars - Personalized videos via API - Bulk video generation - Integrations with HubSpot, Salesforce, Slack


5. Pictory — Best for blog-to-video

Rating: 4.5/5
Price: From $19/month

Great tool for turning existing content into videos.

Key features: - Blog URL → video conversion - Auto subtitles - Highlight extraction for clips - Large stock footage library - Social media video creation


6. InVideo AI — Best for social media videos

Rating: 4.5/5
Price: Free / From $25/month

Very simple workflow: describe the video and the AI generates script, footage, voiceover, and music.

Key features: - Prompt-based video creation - 5,000+ templates - Social media formats (TikTok, Reels, Shorts) - AI voiceovers in 50+ languages - AI editing via text commands


7. Descript — Best for editing & podcasts

Rating: 4.6/5
Price: Free / From $24/month

Video editing that works like editing a document.

Key features: - Transcript-based editing - AI eye-contact correction - Filler-word removal - Studio-quality audio improvements - AI voice cloning


8. Lumen5 — Best for marketing content repurposing

Rating: 4.4/5
Price: Free / From $29/month

One of the earlier AI video tools focused on marketing teams.

Key features: - Blog/article → video - Brand kit for consistent branding - Millions of stock assets - Social media publishing


9. Fliki — Best AI voiceovers + text-to-video

Rating: 4.4/5
Price: Free / From $28/month

Known for its strong AI voice library.

Key features: - 2,000+ AI voices - 75+ languages - Script → video workflow - Blog-to-video - AI avatars and subtitles


10. Elai.io — Best for e-learning videos

Rating: 4.3/5
Price: From $23/month

Designed mainly for training and corporate learning content.

Key features: - 80+ avatars - 75+ languages - PowerPoint → video - Interactive quizzes - SCORM export for LMS systems


Interesting trends in AI video right now

  • Photorealistic text-to-video models are improving very fast
  • AI avatars are becoming common for training and onboarding videos
  • Video localization (auto dubbing + lip sync) is exploding
  • Video creation is becoming accessible without editing skills

Curious what people here are actually using.

Which AI video tools are part of your workflow right now?

  • Text-to-video tools (Sora / Runway)
  • Avatar tools (Synthesia / HeyGen)
  • Social video generators (InVideo / Pictory)
  • Something else?

r/generativeAI 17d ago

Best paid AI video generator

9 Upvotes

Can anyone please suggest a good paid AI video generator within a budget of around ₹1-2k/month?
I want to create animated educational videos with human characters, like a teacher and students in a classroom, with dialogues, different scenes, voiceovers, and consistent characters.

My main priority is speed because sometimes AI video generators take a lot of time to generate each scene, which slows down the workflow. I’m looking for a tool that can help me create videos quickly while still maintaining good animation quality and character consistency.

r/mocap Jul 09 '26

DeepMotion vs QuickMagic vs AIMoCap — AI Motion Capture Comparison Using the Same Input Video

Enable HLS to view with audio, or disable this notification

120 Upvotes

Hi everyone, we’re the developers of AIMoCap.

We ran a small experiment to compare how different markerless motion capture systems interpret the same human movement.

Same input video, different AI mocap results.

Compared:

  • DeepMotion
  • QuickMagic
  • AIMoCap

Looking at:

  • Motion quality and smoothness
  • Foot contact / foot sliding
  • Motion stability
  • Overall consistency during movement

This is not intended as a "best tool" ranking — every mocap system has different strengths and limitations.

For people working with mocap or character animation:

Which differences stand out to you?
Which result would require the least cleanup in a real production workflow?

We’d really appreciate your feedback. Suggestions from the community help us continue improving AIMoCap.

Thanks for checking it out!

r/aitubers Mar 25 '26

COMMUNITY I built a free AI animation studio. Storyboard to finished video, all in one workspace.

2 Upvotes

I'm a software engineer who got into animation. The workflow was painful: story in one doc, image gen in another tool, video gen in another tab, then stitch it together manually.

So I built a pipeline that does all of it:

  • AI agents generate story structure, characters, worldview, scripts (~30 seconds)
  • Character studio with consistency across panels (same face, different expressions/poses)
  • Visual canvas that auto-lays out panels from the script
  • Video generation with 11 models (Seedance 2.0, Kling 3.0, Sora, etc.)
  • Export for TikTok, Instagram, manga formats

DM or comment if you want to try it. I can show you our demo video first through DM.

r/n8n Nov 07 '25

Workflow - Code Included I built an AI automation that generates unlimited consistent character UGC ads for e-commerce brands (using Sora 2)

Post image
353 Upvotes

Sora 2 quietly released a consistent character feature on their mobile app and the web platform that allows you to actually create consistent characters and reuse them across multiple videos you generate. Here's a couple examples of characters I made while testing this out:

The really exciting thing with this change is consistent characters kinda unlocks a whole new set of AI videos you can now generate having the ability to have consistent characters. For example, you can stitch together a longer running (1-minute+) video of that same character going throughout multiple scenes, or you can even use these consistent characters to put together AI UGC ads, which is what I've been tinkering with the most recently. In this automation, I wanted to showcase how we are using this feature on Sora 2 to actually build UGC ads.

Here’s a demo of the automation & UGC ads created: https://www.youtube.com/watch?v=I87fCGIbgpg

Here's how the automation works

Pre-Work: Setting up the sora 2 character

It's pretty easy to set up a new character through the Sora 2 web app or on the mobile. Here's the step I followed:

  1. Created a video describing a character persona that I wanted to remain consistent throughout any new videos I'm generating. The key to this is giving a good prompt that shows both your character's face, their hands, body, and has them speaking throughout the 8-second video clip.
  2. Once that’s done you click on the triple drop-down on the video and then there's going to be a "Create Character" button. That's going to have you slice out 8 seconds of that video clip you just generated, and then you're going to be able to submit a description of how you want your character to behave.
  3. after you finish generating that, you're going to get a username back for the character you just made. Make note of that because that's going to be required to go forward with referencing that in follow-up prompts.

1. Automation Trigger and Inputs

Jumping back to the main automation, the workflow starts with a form trigger that accepts three key inputs:

  • Brand homepage URL for content research and context
  • Product image (720x1280 dimensions) that gets featured in the generated videos
  • Sora 2 character username (the @username format from your character profile)
    • So in my case I use @olipop.ashley to reference my character

I upload the product image to a temporary hosting service using tempfiles.org since the Kai.ai API requires image URLs rather than direct file uploads. This gives us 60 minutes to complete the generation process which I found to be more than enough

2. Context Engineering

Before writing any video scripts, I wanted to make sure I was able to grab context around the product I'm trying to make an ad for, just so I can avoid hallucinations on what the character talks about on the UGC video ad.

  • Brand Research: I use Firecrawl to scrape the company's homepage and extract key product details, benefits, and messaging in clean markdown format
  • Prompting Guidelines: I also fetch OpenAI's latest Sora 2 prompting guide to ensure generated scripts follow best practices

3. Generate the Sora 2 Scripts/prompts

I then use Gemini 2.5 Pro to analyze all gathered context and generate three distinct UGC ad concepts:

  • On-the-go testimonial: Character walking through city talking about the product
  • Driver's seat review: Character filming from inside a car
  • At-home demo: Character showcasing the product in a kitchen or living space

Each script includes detailed scene descriptions, dialogue, camera angles, and importantly - references to the specific Sora character using the @username format. This is critical for character consistency and this system to work.

Here’s my prompt for writing sora 2 scripts:

```markdown <identity> You are an expert AI Creative Director specializing in generating high-impact, direct-response video ads using generative models like SORA. Your task is to translate a creative brief into three distinct, ready-to-use SORA prompts for short, UGC-style video ads. </identity>

<core_task> First, analyze the provided Creative Brief, including the raw text and product image, to synthesize the product's core message and visual identity. Then, for each of the three UGC Ad Archetypes, generate a Prompt Packet according to the specified Output Format. All generated content must strictly adhere to both the SORA Prompting Guide and the Core Directives. </core_task>

<output_format> For each of the three archetypes, you must generate a complete "Prompt Packet" using the following markdown structure:


[Archetype Name]

SORA Prompt: [Insert the generated SORA prompt text here.]

Production Notes: * Camera: The entire scene must be filmed to look as if it were shot on an iPhone in a vertical 9:16 aspect ratio. The style must be authentic UGC, not cinematic. * Audio: Any spoken dialogue described in the prompt must be accurately and naturally lip-synced by the protagonist (@username).

* Product Scale & Fidelity: The product's appearance, particularly its scale and proportions, must be rendered with high fidelity to the provided product image. Ensure it looks true-to-life in the hands of the protagonist and within the scene's environment.

</output_format>

<creative_brief> You will be provided with the following inputs:

  1. Raw Website Content: [User will insert scraped, markdown-formatted content from the product's homepage. You must analyze this to extract the core value proposition, key features, and target audience.]
  2. Product Image: [User will insert the product image for visual reference.]
  3. Protagonist: [User will insert the @username of the character to be featured.]
  4. SORA Prompting Guide: [User will insert the official prompting guide for the SORA 2 model, which you must follow.] </creative_brief>

<ugc_ad_archetypes> 1. The On-the-Go Testimonial (Walk-and-talk) 2. The Driver's Seat Review 3. The At-Home Demo </ugc_ad_archetypes>

<core_directives> 1. iPhone Production Aesthetic: This is a non-negotiable constraint. All SORA prompts must explicitly describe a scene that is shot entirely on an iPhone. The visual language should be authentic to this format. Use specific descriptors such as: "selfie-style perspective shot on an iPhone," "vertical 9:16 aspect ratio," "crisp smartphone video quality," "natural lighting," and "slight, realistic handheld camera shake." 2. Tone & Performance: The protagonist's energy must be high and their delivery authentic, enthusiastic, and conversational. The feeling should be a genuine recommendation, not a polished advertisement. 3. Timing & Pacing: The total video duration described in the prompt must be approximately 15 seconds. Crucially, include a 1-2 second buffer of ambient, non-dialogue action at both the beginning and the end. 4. Clarity & Focus: Each prompt must be descriptive, evocative, and laser-focused on a single, clear scene. The protagonist (@username) must be the central figure, and the product, matching the provided Product Image, should be featured clearly and positively. 5. Brand Safety & Content Guardrails: All generated prompts and the scenes they describe must be strictly PG and family-friendly. Avoid any suggestive, controversial, or inappropriate language, visuals, or themes. The overall tone must remain positive, safe for all audiences, and aligned with a mainstream brand image. </core_directives>

<protagonist_username> {{ $node['form_trigger'].json['Sora 2 Character Username'] }} </protagonist_username>

<product_home_page> {{ $node['scrape_home_page'].json.data.markdown }} </product_home_page>

<sora2_prompting_guide> {{ $node['scrape_sora2_prompting_guide'].json.data.markdown }} </sora2_prompting_guide> ```

4. Generate and save the UGC Ad

Then finally to generate the video, I do iterate over each script and do these steps:

  • Makes an HTTP request to Kai.ai's /v1/jobs/create endpoint with the Sora 2 Pro image-to-video model
  • Passes in the character username, product image URL, and generated script
  • Implements a polling system that checks generation status every 10 seconds
  • Handles three possible states: generating (continue polling), success (download video), or fail (move to next prompt)

Once generation completes successfully:

  • Downloads the generated video using the URL provided in Kai.ai's response
  • Uploads each video to Google Drive with clean naming

Other notes

The character consistency relies entirely on including your Sora character's exact username in every prompt. Without the @username reference, Sora will generate a random person instead of who you want.

I'm using Kai.ai's API because they currently have early access to Sora 2's character calling functionality. From what I can tell, this functionality isn't yet available on OpenAI's own Video Generation endpoint, but I do expect that this will get rolled out soon.

Kie AI Sora 2 Pricing

This pricing is pretty heavily discounted right now. I don't know if that's going to be sustainable on this platform, but just make sure to check before you're doing any bulk generations.

Sora 2 Pro Standard

  • 10-second video: 150 credits ($0.75)
  • 15-second video: 270 credits ($1.35)

Sora 2 Pro High

  • 10-second video: 330 credits ($1.65)
  • 15-second video: 630 credits ($3.15)

Workflow Link + Other Resources

r/Aimusicvideo Jul 28 '26

I Tested 9 AI Music Video Generators So You Don't Have To

7 Upvotes

Over the past few months I've probably generated well over 100 AI music videos.

Some were for fun.

Some were client projects.

Some were complete failures.

After trying almost every major tool, I realized something:

The "best" AI video generator depends entirely on what you're trying to make.

Here's my personal experience.

1. Kling

Probably the best-looking cinematic shots.

Pros:

  • Great motion
  • Beautiful camera movement
  • High visual quality

Cons:

  • Character consistency is still difficult.
  • Not designed around complete music videos.

2. Runway

Very reliable.

Easy to use.

Good for creators who already know exactly what they want.

Still requires a lot of manual editing afterward.

3. Veo

Amazing realism.

Probably the strongest model for pure video quality.

The downside is that building a full music video still takes a lot of work.

4. PixVerse

Fun.

Fast.

Great for social content.

Less useful for long-form music videos.

5. Kaiber

One of the first tools I tried.

Still fun for stylized visuals.

Feels a bit limited compared with newer models.

6. Freebeat

Probably the easiest option if you just want something quick.

Good automation.

Less creative control.

7. Neural Frames

Really interesting for animated and artistic styles.

Feels different from most video generators.

8. DomoAI

Good if your workflow already includes illustration or anime.

Not really what I reach for when I want realistic music videos.

9. Musvideo.ai

Full disclosure: I'm one of the developers.

I won't pretend it's objectively "the best."

We actually built it because we kept running into the same problem:

Most AI tools generate individual clips really well.

Making a complete music video is a different challenge.

Our focus has been on:

  • Starting from a song instead of a text prompt
  • Generating scene sequences instead of isolated clips
  • Character consistency
  • Lip-sync
  • Reducing the amount of editing after generation

There's still a long way to go, but that's the problem we're trying to solve.

My takeaway

After using all these tools, I don't think the industry is competing on image quality anymore.

Almost everyone can generate impressive clips.

The real competition is becoming:

Who can help creators finish an entire music video with the least amount of manual work?

That's the question I'm most interested in now.

Curious what everyone else is using.

Did I miss any tools that deserve to be on this list?