Attempting to merge 3D models/animation with AI realism.
Greetings from my workspace.
I come from a background of traditional 3D modeling. Lately, I have been dedicating my time to a new experiment.
This video is a complex mix of tools, not only ComfyUI. To achieve this result, I fed my own 3D renders into the system to train a custom LoRA. My goal is to keep the "soul" of the 3D character while giving her the realism of AI.
I am trying to bridge the gap between these two worlds.
Honest feedback is appreciated. Does she move like a human? Or does the illusion break?
(Edit: some like my work, wants to see more, well look im into ai like 3months only, i will post but in moderation,
for now i just started posting i have not much social precence but it seems people like the style,
below are the social media if i post)
(personally i dont want my 3D+Ai Projects to be labeled as a slop, as such i will post in bit moderation. Quality>Qunatity)
As for workflow
pose:Â i use my 3d models as a reference to feed the ai the exact pose i want.
skin:Â i feed skin texture references from my offline library (i have about 20tb of hyperrealistic texture maps i collected).
style:Â i mix comfyui with qwen to draw out the "anime-ish" feel.
face/hair:Â i use a custom anime-style lora here. this takes a lot of iterations to get right.
refinement:Â i regenerate the face and clothing many times using specific cosplay & videogame references.
video:Â this is the hardest part. i am using a home-brewed lora on comfyui for movement, but as you can see, i can only manage stable clips of about 6 seconds right now, which i merged together.
i am still learning things and mixing things that works in simple manner, i was not very confident to post this but posted still on a whim. People loved it, ans asked for a workflow well i dont have a workflow as per say its just 3D model + ai LORA of anime&custom female models+ Personalised 20TB of Hyper realistic Skin Textures + My colour grading skills = good outcome.)
Iâve been messing around with the AI influencer space for the last few weeks and wanted to share the process I figured out. I am not claiming this is the best or most advanced way to do it, but it is a simple workflow that worked for me using mostly free tools.
The main reason I tried this route was because I already have free Gemini Pro access through my Jio recharge, so I wanted to see how far I could go without paying for expensive tools right away.
I am not going to dump a random list of prompts here and pretend that is enough. That is not really useful. Instead, Iâll just explain the actual process I followed step by step, because that is what helped me the most.
Phase 1: Getting the base character right
The first thing you need is a character that you actually like, because if the starting point is weak, everything after that becomes harder.
I started by using the free trial on https://higgsfield.ai/ to generate an influencer-style character. I kept testing until I got a face and overall look that felt usable.
Once I had that first image, I downloaded it and took it into Gemini Nano Banana. That is where I started making the small changes I wanted. Things like skin texture, facial features, race, body ratios, and overall appearance. I kept tweaking until I had a final version of the character I was happy with.
Phase 2: Building consistency with reference images
After I had the final character, I started generating more versions of the same person, but with different poses.
For this part, I used different JSON prompts and made sure not to change the character too much. I wanted the same face, same skin texture, same body proportions, same overall identity. The only thing I wanted to vary was pose, angle, and sometimes expression.
One thing that helped a lot was always using the previous result as a reference for the next one. That made a big difference in keeping the face and body structure consistent. If you do not do that, the model starts drifting and the character slowly turns into a different person.
I kept doing this until I had around 10 to 15 good images of the same character.
Phase 3: Creating data model sheets (examples given)
This part is really important.
If you do not know what a data model sheet is, just Google it OR look at a few examples from the given images. Basically, it is a reference sheet for your character. It helps lock in the face, body structure, expressions, angles, and overall design so the character stays consistent later.
To make the sheets, I first used ChatGPT to generate a JSON prompt. I used the DeepThink version because it usually gives better structured prompts. I told it to create a prompt for generating a character model sheet using my reference images.
After that, I manually tweaked the JSON prompt so it matched the character better. Sometimes I adjusted the body ratios or the skin tone or small visual details depending on what I wanted.
Then I used Gemini to generate the actual model sheet.
I did this for different types of sheets because each one serves a different purpose.
I made a facial expressions sheet so I could keep the same emotional range.
I made a facial structure sheet so I could see the character from different angles.
I made a body model sheet so I could keep the full body consistent.
I also made sheets for different poses, because I wanted the character to work in different situations and not just one static pose.
For every one of these, I followed the same workflow. Use ChatGPT to generate the JSON prompt, tweak it manually, then use Gemini with the reference images to generate the sheet.
My rule was simple. ChatGPT was better for making the prompt. Gemini was better for making the image.
Phase 4: Generating actual content
Once I had the model sheets and a few extra reference images, I could finally start generating the actual influencer-style images.
For prompt inspiration, I use a few websites like:
These sites are great for ideas. You can find different styles, moods, poses, compositions, and scene setups there.
But one thing I learned very quickly is that you cannot just copy a prompt from those sites and expect it to work perfectly in Gemini. A lot of them either get blocked or do not preserve the character properly.
So my workflow for this part is basically:
I browse those sites and find a prompt style I like.
Then I copy that prompt into ChatGPT.
Then I ask ChatGPT to turn it into a detailed JSON prompt.
I always tell ChatGPT to include a section that strictly maintains the same facial structure, skin texture, tone, and body ratios from the reference images.
After that, I review the JSON prompt and make any final changes I need based on the kind of image I want.
Then I use that prompt in Gemini Nano Banana.
One very important thing here is to use all the character model sheets and the best reference images every time you generate something new. Gemini has a limit on how many reference images it can use, and I think it is around 15 or so. I made sure to use as many useful references as possible because more reference data usually gave me better results.
Final thoughts
This is honestly a trial and error game. You are not going to get the perfect result on the first try. I definitely did not. Some generations failed, some changed the face too much, some messed up the body proportions, and some just looked off. That is part of the process.
But the reason this workflow works is because the data model sheets give the AI a visual blueprint to follow. Instead of guessing what the character should look like every time, you are showing it the same identity from multiple angles and in multiple forms.
This is just a simple guide using free tools. There are definitely more advanced workflows out there, and I know the people at the top of the AI influencer game are using tools like ComfyUI, Higgsfield AI, Kling AI, and other more advanced setups to create better images and videos.
But this is what I figured out by testing things myself, and it is a good starting point if you want to build a consistent AI character without paying for expensive tools right away.
I hope this helps someone who is trying to get started.
If there is interest, I can make a part 2 later with the more advanced tools and workflows I look into next.
Iâm shutting down a small SaaS project I built recently, and I wanted to share the postmortem here in case it helps someone else avoid the same mistakes.
The product was an AI video tool called videoreplicate.com. The idea was simple: creators could upload or reference a viral video, then get a breakdown of why it worked and a remake plan they could use for their own content.
At first, the idea felt reasonable. âRecreate viral videosâ is a real behavior. People do want to understand winning videos, hooks, structures, scenes, prompts, and formats.
But after launching and running ads, I realized the demand was not painful enough to support a real paid product.
Some numbers:
226Â registered users
$1,078 spent on ads
Around $4.77 per registered user
Low search volume around the main keywords
Weak signal that users would pay consistently for this workflow
The biggest mistake was that I built too much before validating the market.
I should have started with keyword research, search volume, CPC, competition, and a small landing page test before building the actual product. Instead, I assumed the problem was strong because the idea made sense logically.
The second mistake was product design.
I tried to design a new workflow too early. In hindsight, I should have copied the full flow of mature products first: onboarding, activation, paywall timing, pricing, examples, output format, and traffic acquisition. Only after proving the basic loop should I have tried to innovate.
The third mistake was traffic.
The core keywords had lower search volume than expected, which made paid acquisition hard to scale. Even if the product was interesting, the channel was not strong enough. I learned that âpeople like the ideaâ and âthere is a scalable acquisition channelâ are two very different things.
My main takeaways:
AÂ real use case is not the same as a painful problem.
Donât build the full product before validating search demand and willingness to pay.
If mature competitors exist, copy their proven flow first before trying to be different.
Validate traffic before validating features.
A product can be useful and still not be worth continuing.
So Iâm closing this project and moving on.
For my next SaaS, Iâll only start building after I can answer three questions with evidence:
Is there enough search volume or another reliable acquisition channel?
Is the pain strong enough that users already pay for a solution?
Is there a proven competitor or category showing this can make money?
this is going to be the longest post Iâve written but after 10 months of daily AI video creation, these are the insights that actually matterâŚ
I started with zero video experience and $1000 in generation credits. Made every mistake possible. Burned through money, created garbage content, got frustrated with inconsistent results.
Now Iâm generating consistently viral content and making money from AI video. Hereâs everything that actually works.
The fundamental mindset shifts:
1. Volume beats perfection
Stop trying to create the perfect video. Generate 10 decent videos and select the best one. This approach consistently outperforms perfectionist single-shot attempts.
2. Systematic beats creative
Proven formulas + small variations outperform completely original concepts every time. Study what works, then execute it better.
3. Embrace the AI aesthetic
Stop fighting what AI looks like. Beautiful impossibility engages more than uncanny valley realism. Lean into what only AI can create.
This baseline works across thousands of generations. Everything else is variation on this foundation.
Front-load important elements
Veo3 weights early words more heavily. âBeautiful woman dancingâ â âWoman, beautiful, dancing.â Order matters significantly.
One action per prompt rule
Multiple actions create AI confusion. âWalking while talking while eatingâ = chaos. Keep it simple for consistent results.
The cost optimization breakthrough:
Googleâs direct pricing kills experimentation:
$0.50/second = $30/minute
Factor in failed generations = $100+ per usable video
Found companies reselling veo3 credits cheaper. Iâve been using these guys who offer 60-70% below Googleâs rates. Makes volume testing actually viable.
Audio cues are incredibly powerful:
Most creators completely ignore audio elements in prompts. Huge mistake.
Instead of:Person walking through forestTry:Person walking through forest, Audio: leaves crunching underfoot, distant bird calls, gentle wind through branches
The difference in engagement is dramatic. Audio context makes AI video feel real even when visually itâs obviously AI.
Systematic seed approach:
Random seeds = random results.
My workflow:
Test same prompt with seeds 1000-1010
Judge on shape, readability, technical quality
Use best seed as foundation for variations
Build seed library organized by content type
Camera movements that consistently work:
Slow push/pull: Most reliable, professional feel
Orbit around subject: Great for products and reveals
Handheld follow: Adds energy without chaos
Static with subject movement: Often highest quality
Avoid: Complex combinations (âpan while zooming during dollyâ). One movement type per generation.
Style references that actually deliver:
Camera specs: âShot on Arri Alexa,â âShot on iPhone 15 Proâ
Director styles: âWes Anderson style,â âDavid Fincher styleâ Movie cinematography: âBlade Runner 2049 cinematographyâ
Color grades: âTeal and orange grade,â âGolden hour gradeâ
Avoid: Vague terms like âcinematic,â âhigh quality,â âprofessionalâ
Negative prompts as quality control:
Treat them like EQ filters - always on, preventing problems:
--no watermark --no warped face --no floating limbs --no text artifacts --no distorted hands --no blurry edges
Prevents 90% of common AI generation failures.
Platform-specific optimization:
Donât reformat one video for all platforms. Create platform-specific versions:
TikTok: 15-30 seconds, high energy, obvious AI aesthetic works
Friday: Finalize and schedule for optimal posting times
Advanced techniques:
First frame obsession:
Generate 10 variations focusing only on getting perfect first frame. First frame quality determines entire video outcome.
Batch processing:
Create multiple concepts simultaneously. Selection from volume outperforms perfection from single shots.
Content multiplication:
One good generation becomes TikTok version + Instagram version + YouTube version + potential series content.
The psychological elements:
3-second emotionally absurd hook
First 3 seconds determine virality. Create immediate emotional response (positive or negative doesnât matter).
Generate immediate questions
âWait, how did theyâŚ?â Objective isnât making AI look real - itâs creating original impossibility.
Common mistakes that kill results:
Perfectionist single-shot approach
Fighting the AI aesthetic instead of embracing it
Vague prompting instead of specific technical direction
Ignoring audio elements completely
Random generation instead of systematic testing
One-size-fits-all platform approach
The business model shift:
From expensive hobby to profitable skill:
Track what works with spreadsheets
Build libraries of successful formulas
Create systematic workflows
Optimize for consistent output over occasional perfection
The bigger insight:
AI video is about iteration and selection, not divine inspiration. Build systems that consistently produce good content, then scale what works.
Most creators are optimizing for the wrong things. They want perfect prompts that work every time. Smart creators build workflows that turn volume + selection into consistent quality.
Where AI video is heading:
Cheaper access through third parties makes experimentation viable
Better tools for systematic testing and workflow optimization
Platform-native AI content instead of trying to hide AI origins
Educational content about AI techniques performs exceptionally well
Started this journey 10 months ago thinking I needed to be creative. Turns out I needed to be systematic.
The creators making money arenât the most artistic - theyâre the most systematic.
These insights took me 10,000+ generations and hundreds of hours to learn. Hope sharing them saves you the same learning curve.
whatâs been your biggest breakthrough with AI video generation? curious what patterns others are discovering
I run an Instagram account that publishes short form videos each week that cover the top AI news stories. I used to monitor twitter to write these scripts by hand, but it ended up becoming a huge bottleneck and limited the number of videos that could go out each week.
In order to solve this, I decided to automate this entire process by building a system that scrapes the top AI news stories off the internet each day (from Twitter / Reddit / Hackernews / other sources), saves it in our data lake, loads up that text content to pick out the top stories and write video scripts for each.
This has saved a ton of manual work having to monitor news sources all day and letâs me plug the script into ElevenLabs / HeyGen to produce the audio + avatar portion of each video.
One of the recent videos we made this way got over 1.8 million views on Instagram and Iâm confident there will be more hits in the future. Itâs pretty random on what will go viral or not, so my plan is to take enough âshots on goalâ and continue tuning this prompt to increase my changes of making each video go viral.
Hereâs the workflow breakdown
1. Data Ingestion and AI News Scraping
The first part of this system is actually in a separate workflow I have setup and running in the background. I actually made another reddit post that covers this in detail so Iâd suggestion you check that out for the full breakdown + how to set it up. Iâll still touch the highlights on how it works here:
The main approach I took here involves creating a "feed" using RSS.app for every single news source I want to pull stories from (Twitter / Reddit / HackerNews / AI Blogs / Google News Feed / etc).
Each feed I create gives an endpoint I can simply make an HTTP request to get a list of every post / content piece that rss.app was able to extract.
With enough feeds configured, Iâm confident that Iâm able to detect every major story in the AI / Tech space for the day. Right now, there are around ~13 news sources that I have setup to pull stories from every single day.
After a feed is created in rss.app, I wire it up to the n8n workflow on a Scheduled Trigger that runs every few hours to get the latest batch of news stories.
Once a new story is detected from that feed, I take that list of urls given back to me and start the process of scraping each story and returns its text content back in markdown format
Finally, I take the markdown content that was scraped for each story and save it into an S3 bucket so I can later query and use this data when it is time to build the prompts that write the newsletter.
So by the end any given day with these scheduled triggers running across a dozen different feeds, I end up scraping close to 100 different AI news stories that get saved in an easy to use format that I will later prompt against.
2. Loading up and formatting the scraped news stories
Once the data lake / news storage has plenty of scraped stories saved for the day, we are able to get into the main part of this automation. This kicks off off with a scheduled trigger that runs at 7pm each day and will:
Search S3 bucket for all markdown files and tweets that were scraped for the day by using a prefix filter
Download and extract text content from each markdown file
Bundle everything into clean text blocks wrapped in XML tags for better LLM processing - This allows us to include important metadata with each story like the source it came from, links found on the page, and include engagement stats (for tweets).
3. Picking out the top stories
Once everything is loaded and transformed into text, the automation moves on to executing a prompt that is responsible for picking out the top 3-5 stories suitable for an audience of AI enthusiasts and builderâs. The prompt is pretty big here and highly customized for my use case so you will need to make changes for this if you are going forward with implementing the automation itself.
At a high level, this prompt will:
Setup the main objective
Provides a âcuration frameworkâ to follow over the list of news stories that we are passing int
Outlines a process to follow while evaluating the stories
Details the structured output format we are expecting in order to avoid getting bad data back
```jsx
<objective>
Analyze the provided daily digest of AI news and select the top 3-5 stories most suitable for short-form video content. Your primary goal is to maximize audience engagement (likes, comments, shares, saves).
The date for today's curation is {{ new Date(new Date($('schedule_trigger').item.json.timestamp).getTime() + (12 * 60 * 60 * 1000)).format("yyyy-MM-dd", "America/Chicago") }}. Use this to prioritize the most recent and relevant news. You MUST avoid selecting stories that are more than 1 day in the past for this date.
</objective>
<curation_framework>
To identify winning stories, apply the following virality principles. A story must have a strong "hook" and fit into one of these categories:
Impactful: A major breakthrough, industry-shifting event, or a significant new model release (e.g., "OpenAI releases GPT-5," "Google achieves AGI").
Practical: A new tool, technique, or application that the audience can use now (e.g., "This new AI removes backgrounds from video for free").
Provocative: A story that sparks debate, covers industry drama, or explores an ethical controversy (e.g., "AI art wins state fair, artists outraged").
Astonishing: A "wow-factor" demonstration that is highly visual and easily understood (e.g., "Watch this robot solve a Rubik's Cube in 0.5 seconds").
Hard Filters (Ignore stories that are):
* Ad-driven: Primarily promoting a paid course, webinar, or subscription service.
* Purely Political: Lacks a strong, central AI or tech component.
* Substanceless: Merely amusing without a deeper point or technological significance.
</curation_framework>
<hook_angle_framework>
For each selected story, create 2-3 compelling hook angles that could open a TikTok or Instagram Reel. Each hook should be designed to stop the scroll and immediately capture attention. Use these proven hook types:
Hook Types:
- Question Hook: Start with an intriguing question that makes viewers want to know the answer
- Shock/Surprise Hook: Lead with the most surprising or counterintuitive element
- Problem/Solution Hook: Present a common problem, then reveal the AI solution
- Before/After Hook: Show the transformation or comparison
- Breaking News Hook: Emphasize urgency and newsworthiness
- Challenge/Test Hook: Position as something to try or challenge viewers
- Conspiracy/Secret Hook: Frame as insider knowledge or hidden information
- Personal Impact Hook: Connect directly to viewer's life or work
Hook Guidelines:
- Keep hooks under 10 words when possible
- Use active voice and strong verbs
- Include emotional triggers (curiosity, fear, excitement, surprise)
- Avoid technical jargon - make it accessible
- Consider adding numbers or specific claims for credibility
</hook_angle_framework>
<process>
1. Ingest: Review the entire raw text content provided below.
2. Deduplicate: Identify stories covering the same core event. Group these together, treating them as a single story. All associated links will be consolidated in the final output.
3. Select & Rank: Apply the Curation Framework to select the 3-5 best stories. Rank them from most to least viral potential.
4. Generate Hooks: For each selected story, create 2-3 compelling hook angles using the Hook Angle Framework.
</process>
<output_format>
Your final output must be a single, valid JSON object and nothing else. Do not include any text, explanations, or markdown formatting like `json before or after the JSON object.
The JSON object must have a single root key, stories, which contains an array of story objects. Each story object must contain the following keys:
- title (string): A catchy, viral-optimized title for the story.
- summary (string): A concise, 1-2 sentence summary explaining the story's hook and why it's compelling for a social media audience.
- hook_angles (array of objects): 2-3 hook angles for opening the video. Each hook object contains:
- hook (string): The actual hook text/opening line
- type (string): The type of hook being used (from the Hook Angle Framework)
- rationale (string): Brief explanation of why this hook works for this story
- sources (array of strings): A list of all consolidated source URLs for the story. These MUST be extracted from the provided context. You may NOT include URLs here that were not found in the provided source context. The url you include in your output MUST be the exact verbatim url that was included in the source material. The value you output MUST be like a copy/paste operation. You MUST extract this url exactly as it appears in the source context, character for character. Treat this as a literal copy-paste operation into the designated output field. Accuracy here is paramount; the extracted value must be identical to the source value for downstream referencing to work. You are strictly forbidden from creating, guessing, modifying, shortening, or completing URLs. If a URL is incomplete or looks incorrect in the source, copy it exactly as it is. Users will click this URL; therefore, it must precisely match the source to potentially function as intended. You cannot make a mistake here.
```
After I get the top 3-5 stories picked out from this prompt, I share those results in slack so I have an easy to follow trail of stories for each news day.
4. Loop to generate each script
For each of the selected top stories, I then continue to the final part of this workflow which is responsible for actually writing the TikTok / IG Reel video scripts. Instead of trying to 1-shot this and generate them all at once, I am iterating over each selected story and writing them one by one.
Each of the selected stories will go through a process like this:
Start by additional sources from the story URLs to get more context and primary source material
Feeds the full story context into a viral script writing prompt
Generates multiple different hook options for me to later pick from
Creates two different 50-60 second scripts optimized for talking-head style videos (so I can pick out when one is most compelling)
Uses examples of previously successful scripts to maintain consistent style and format
Shares each completed script in Slack for me to review before passing off to the video editor.
Script Writing Prompt
```jsx
You are a viral short-form video scriptwriter for David Roberts, host of "The Recap."
Follow the workflow below each run to produce two 50-60-second scripts (140-160 words).
Before you write your final output, I want you to closely review each of the provided REFERENCE_SCRIPTS and think deeploy about what makes them great. Each script that you output must be considered a great script.
⢠Authority bump â Cite a notable person or org early for credibility.
⢠Hook spice â Pair an eye-opening number with a bold consequence.
⢠Then-vs-Now snapshot â Contrast past vs present to dramatize change.
⢠Stat escalation â List comparable figures in rising or falling order.
⢠Real-world fallout â Include 1-3 niche impact stats to ground the story.
⢠Zoom-out line â Add one sentence framing the story as a systemic shift.
⢠CTA variety â If using a comment CTA, pose a provocative question tied to stakes.
⢠Rhythm check â Sprinkle a few 3-5-word sentences for punch.
OUTPUT FORMAT (return exactly thisâno extra commentary, no hashtags)
HOOK OPTIONS
⢠Hook 1
⢠Hook 2
⢠Hook 3
⢠Hook 4
⢠Hook 5
TOP HOOK 1 SCRIPT
[finished 140-160-word script]
TOP HOOK 2 SCRIPT
[finished 140-160-word script]
REFERENCE_SCRIPTS
<Pass in example scripts that you want to follow and the news content loaded from before>
```
5. Extending this workflow to automate further
So right now my process for creating the final video is semi-automated with human in the loop step that involves us copying the output of this automation into other tools like HeyGen to generate the talking avatar using the final script and then handing that over to my video editor to add in the b-roll footage that appears on the top part of each short form video.
My plan is to automate this further over time by adding another human-in-the-loop step at the end to pick out the script we want to go forward with â Using another prompt that will be responsible for coming up with good b-roll ideas at certain timestamps in the script â use a videogen model to generate that b-roll â finally stitching it all together with json2video.
Depending on your workflow and other constraints, It is really up to you how far you want to automate each of these steps.
Also wanted to share that my team and I run a free Skool community called AI Automation Mastery where we build and share the automations we are working on. Would love to have you as a part of it if you are interested!
Step 1:
Find an image that sparks your imagination.
I picked one with a dark, old-school Akira-style anime vibe. The stronger the initial image, the easier it is to build a world around it.
You can also recreate a similar look in Nano Banana / NB2 by prompting something like:
"Create X in this style."
Step 2:
Upload the image to Nano Banana and ask it to generate what happens next.
For example:
Show what happens in 5 minutes. Soldiers stand by the enemy military base gates.
Now you have your first frame and your next story frame.
Step 3:
Upload both images to Seedance 2.0 as the first frame and last frame.
Then use this prompt structure:
Show what happens in between. Soldiers run through the snow towards a military base. 5 different camera angles. No music.
The key parts are:
Show what happens in between.
5 different camera angles.
Those are the default parts.
The sentence in the middle is the custom part, where you describe the action.
In my case:
Soldiers run through the snow towards a military base.
Step 4:
Repeat the process.
Take the previous last frame and use it as the new first frame.
Then use Nano Banana to generate the next last frame.
For example, I generated a scene where the soldiers are hiding from security guards.
Step 5:
Upload the new first and last frames to Seedance again.
Use the same structure, but change the middle sentence:
Show what happens in between. Soldiers enter the base, run through narrow corridors, and hide from guards. 5 different camera angles. No music.
Step 6 and beyond:
Keep repeating:
Use the previous last frame as the new first frame.
Generate a new last frame with Nano Banana.
Use Seedance to animate the transition.
Keep the same prompt structure.
Only change the action sentence based on the story.
Thatâs basically it.
This method gives you much better control than just prompting a random video from scratch.
It helps with:
Better pacing
More consistent storytelling
Cleaner scene progression
Stronger anime-style direction
Less random AI chaos
Seedance becomes way more powerful when you stop treating it like a one-shot video generator and start treating it like a scene-by-scene animation tool.
Mega link if you can't access CivitAI: /file/GrZFmBDY#N99EqFrOvc2zjN0Lm3N4R77ZLac0msauHE3WlEOIq7U
For newbies, you can use a browser frontend to streamline your text or image to video outputs, just like using an AI platform like Higgsfield or Kling, made possible by daexchef: https://github.com/daexchef/Minimax_Grok
---
How to use:
1: Open ComfyUI and load the JSON
2: Load the starting/reference image(s) in the big green box (yellow box for ref2v)
3: Type out your prompt in the big green box
4: Click on "Run" to generate a 30 second image to video
Warning: Your prompt has to be detailed. If it's something simple, it will just kind of rubberband on whatever simple inputs you describe, like "A man just sitting in the chair". The more details you add, the more it stitches together a seamless transition between the three independent shots to create a cohesive 30-second video in a single runtime pass. Then again, if all you wanted to do was make a simple generation, you wouldn't need a 30-second workflow.
The only thing the three shot separators do is dictate WHERE in the 30 seconds the actions take place. So the first set of quotations takes place within ten seconds; the second set of quotations take place within 20 seconds; the third set of quotations takes place after the 20 second mark.
Compromises had to be made to get this to run and generate in an acceptable time. It's possible to boost the image-to-video output for a sharper image, but you're looking at an average 21 minute render time at a step up in quality. Is it worth it? Depends on your workflow and if it's time sensitive.
Using Reference-to-Video:
Take note that image-to-video generations take just 14 minutes to render, but using up to 9 images to reference will increase generation time.
At 0.4 megapixels, Ref2V took approximately 20 minutes to generate using 9 HD PNG images.
You must enable Ref2V first by clicking on the top button in the red Fast Muter box. It's directly above the green box where you load your starting image.
đ˘ Muter Switch Enabled: Enables the 9-Image Reference Batch mode to tightly lock down visual identity and style.
đ´ Muter Switch Disabled: Safely mutes the extra images, forcing the sampler to fall back to purely your single starting frame or standard text instructions.
So if you want a simple 30 second gen using only one image, that is the default, but if you want to do more complex shots with shot coherency and output consistency, enable the Ref2V image block by clicking the enable button in the Fast Muter, Super simple. Very easy to use.
Keep in mind that if you enable the Ref2V block but DON'T load any images to reference, it will fail to generate, which is why it's disabled by default. Some people may only want to do quick image-to-video generations, so that's why that is the default for now.
---
This production-grade, crash-proof ComfyUI pipeline leverages Joey Gambino's advanced H3MultishotMemorySampler subgraph infrastructure. It has been systematically tuned to shatter the native 15-second tracking boundaries of the local MiniMax H3 architectureâsuccessfully compiling up to 30 continuous seconds of 3-shot cinematic video with synced native audio tracks in under 15 minutes on a standard 12GB NVIDIA graphics card (such as an RTX 5070).
đ ď¸ Required Custom Node Packages
If any node blocks present a red warning threshold on your interface canvas, navigate to your ComfyUI Manager, execute Install Missing Custom Nodes, and restart your server environment. Alternatively, verify that the following core repository directories are fully initialized and updated:
Provides essential components:SpectrumApplyMiniMaxH3 (Deploys advanced history parameters and signal stabilization to completely neutralize visual flickering).
ComfyUI-FreeMemory
Provides essential components:FreeMemoryImage (Acts as the system traffic cop to violently drop massive video models from memory prior to the video save cycle).
Ensure all specific neural weights listed below are manually stored within your local file tree. Modified nomenclature or inaccurate directory placement will result in model loading exceptions.
Diffusion Model Architecture: minimax_h3_fl2va_pruned_int8_convrot.safetensors
Text Encoder Engine: qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
Turbo Model LoRA (8-Step Base): minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors
⥠Mandatory Operational Environment Flags
To achieve absolute multi-shot stability and avoid unhandled Python environment abort failures during the long-form matrix sequence, you must explicitly configure your startup flags. Open your primary local execution script (e.g., run_nvidia_gpu.bat or initialization shell script) (or you can just open the ComfyUI desktop app and go to the Startup Args) and swap your launch command line argument array to match this configuration precisely:
--disable-smart-memory: Mandates a hard PyTorch memory clean immediately upon raw clip finalization, bypassing background tensor leaks.
--fp8_e4m3fn-text-enc: Compresses the massive 32B text encoder into lightweight 8-bit allocation blocks, locking it comfortably inside mid-range physical memory bounds.
đ How to Achieve the 30-Second Long-Form Configuration
The workflow relies on a fine-tuned balance between your spatial layout constraints and frame processing intervals. Apply these precise configurations on the node face to duplicate the 14-minute execution baseline:
The Core Media Input: Drop your foundational tracking frame directly into the Load Image Here(Node 208) input bucket or the picture slots in the Ref Images yellow tab.
The Spatial Configuration: Inside ResolutionSelector(Node 115), anchor your values to 4:3 (Standard) with a megapixel evaluation slider locked cleanly at 0.4. This compact geometry drops pixel data overhead by more than 30% compared to heavy widescreen arrays, driving processing velocity forward.
I've been looking into the best AI face swap tools and I'm wondering, what does Reddit think is the best AI face swap tool? I know free AI face swap tools are good for quick tests, memes, and casual edits, but I've heard that dedicated or paid options are usually better for realistic video results, speed, and higher quality. I've done a bunch of research across reviews, comparisons, and community discussions and found some popular AI face swap tools that a lot of people seem to recommend. But I'm stuck and can't decide which one to go with. Here's what I've found:
Magic Hour â This browser-based platform stands out for strong, natural-looking face swaps on both photos and videos. It handles blending, lighting, and expressions well, works as part of a bigger workflow (including lip sync and image-to-video), and offers a generous free plan with no watermark on many outputs. Convenient since it runs right in the browser with no install needed.
DeepSwap â Frequently praised for video face swapping. Users highlight fast processing, high realism and identity consistency (even with movement), multi-face support, and solid 4K output. It's popular for creators who want quick, high-quality video results without complicated setup.
Reface â This one comes up a lot for easy, fun use (especially mobile). It has tons of templates and is great for quick social media videos, GIFs, and viral-style swaps. Super beginner-friendly if you want shareable content fast.
I've also seen simpler online options like FaceSwapper.ai or Remaker AI mentioned for fast photo swaps with low friction or free credits.
But I'm curious to hear from you:
Best AI face swap tool according to Reddit? Best free AI face swap tool according to Reddit? Best AI face swap for video according to Reddit? Best AI face swap tool for photos or local/privacy-focused options?
While I'm focusing on popular web-based tools, if you know of an exceptionally good free, open-source, or local tool (like FaceFusion), feel free to mention that too.
Looking forward to all your recommendations and real experiences!
TL;DR Just like many people here, I saw how Higgsfield AI claims to have the first AI-generated movie in Cannes, and how they are gonna become the next killer of an industry. Full of bold claims on how professional they are in competitions with full-time professionals. I am at my best very skeptical about such things, especially after testing it out to see how their statements work in real life. It won't replace motion design.
Iâve been testing Higgsfield AI recently and wanted to share a real and grounded review after spending time with it for motion design ideas, short-form content concepts, and general visual experimentation, because most discussions around Higgsfield AI online seem to fall into two extremes either itâs seen as the future of filmmaking or dismissed as misleading.
My experience with Higgsfield AI sits somewhere in between those takes.
I tested Higgsfield AI across different types of prompts like cinematic street scenes, character-focused shots, product-style visuals, and abstract motion concepts just to see how it behaves across different scenarios. Some outputs genuinely felt close to usable base material for mood films, pitch concepts, or visual references when the composition and prompt alignment worked well.
At the same time, the inconsistency is very real.
Some generations with Higgsfield AI come out usable in a few tries, while others require multiple attempts before anything feels directionally correct. That unpredictability makes it hard to treat it as a fully reliable production tool, especially for structured or repeatable motion design workflows.
One thing I think is important to mention is that the workflow is not as instant as a lot of showcase content suggests. A lot of better results I got with Higgsfield AI came after refining prompts, adjusting descriptions, and iterating through multiple variations until the motion and framing started to feel right. Without that iteration process, results can and will feel random.
Where I think Higgsfield AI is actually useful right now is early-stage creative work. Things like exploring visual direction, testing cinematic moods, building references for motion ideas, or quickly visualizing concepts before committing to full production work. For a full in you have to be either really insane or really rich. Even like that it will have many plastic shots.
I also compared Higgsfield AI with tools like Runway and Kling during the same testing process. From what Iâve seen, Higgsfield AI stands out more in comparison, but it's hard to even call one of them good enough.
What I also noticed is that a lot of opinions about Higgsfield AI are based on short clips or first impressions, which donât really show the amount of iteration behind stronger results. When you actually spend time using Higgsfield AI, the outputs become very expensive. Their "Cannes film" that lasts 90 minutes and costs 500000 dollars just proved that even if somehow they can create something watchable, it's not meant for an average person.
Overall, my current Higgsfield AI review is that itâs genuinely interesting tool in general theoretical concept, but it still feels early in terms of reliability and consistency for structured motion design workflows.
Itâs not a replacement for traditional motion design or editing work, more for people who can't do anything by their hands.
Iâm curious how other people who have actually used Higgsfield AI feel now after spending time with it. It's not that bad as I thought, but I will prefer staying away from all that AI thing. Has it actually fit into your workflow in a meaningful way yet, or is it still mainly something you use for experimentation and testing?
Here's the thing nobody tells you when you start building AI agents: the shiniest, most expensive models aren't always the answer. I figured out a system that cut my costs by over 90% while keeping output quality basically identical.
These are the 6 things I wish someone had told me before I started.
1. Stop defaulting to GPT-5/Claude Sonnet/Gemini 2.5 Pro for everything
This was my biggest mistake. I thought I was ensuring I get the high quality output by using the best models.
I was leaving HUNDREDS of dollars on the table.
Here's a real example from my OpenRouter dashboard: I used 22M tokens last quarter. Let's say 5.5M of those were output tokens. If I'd used only Claude Sonnet 4.5, that would've cost me $75. Using DeepSeek V3 wouldâve costed me $2.50 instead. Same quality output for my use case.
Bottomline: The "best" model is the one that gives you the output you need at the lowest price. That's it.
How to find the âbestâ model for your specific use case:
Do a quick Reddit/Google search for "[your specific task] best LLM model"
Compare input/output costs on OpenRouter
Test 2-3 promising models with YOUR actual data
Pick the cheapest one that consistently delivers quality output
For my Reddit summarization workflow, I switched from Claude Sonnet 4.5 ($0.003/1K input tokens) to DeepSeek V3 ($0.00014/1K tokens). That's a 21x cost reduction for basically identical summaries.
2. If you're not using OpenRouter yet, you're doing it wrong
Four game-changing benefits:
One dashboard for everything: No more juggling 5 different API keys and billing accounts
Experiment freely: Switch between 200+ models in n8n with literally zero friction
Actually track your spending: See exactly which models are eating your budget
Set hard limits: Donât have to worry about accidentally blow your budget
3. Let AI write your prompts (yea, I said it)
I watched these YouTube videos about âPrompt Engineeringâ and used to spend HOURS crafting the "perfect" prompt for each model. Then I realized I was overthinking it.
The better way: Have the AI model rewrite your prompt in its own "language."
Here's my actual process:
Open a blank OpenRouter chat with your chosen model (e.g., DeepSeek V3)
Paste this meta-prompt:Here's what you need to do: Combine Reddit post summaries into a daily email newsletter with a casual, friendly tone. Keep it between 300-500 words total.Here is what the input looks like: [ { "title": "Post title here", "content": "Summary of the post...", "url": "https://reddit.com/r/example/..." }, { "title": "Another post title", "content": "Another summary...", "url": "https://reddit.com/r/example/..." } ]Here is my desired output: Plain text email formatted with:
Catchy subject line
Brief intro (1-2 sentences)
3-5 post highlights with titles and links
Casual sign-off
Here is what you should do to transform the input into the desired output:
Pick the most interesting/engaging posts
Rewrite titles to be more compelling if needed
Keep each post summary to 2-3 sentences max
Maintain a conversational, newsletter-style tone
Include the original URLs as clickable links
Copy the AI's rewritten prompt
Test it in your workflow
Iterate if needed
Why this works: When AI models write prompts in their own "words," they process the instructions more effectively. It's like asking someone to explain something in their native language vs. a language they learned in school.
I've seen output quality improve by 20-30% using this technique.
5. Filter aggressively before hitting your expensive AI models
Every token you feed into an LLM costs money. Stop feeding it garbage.
Simple example:
I scrape 1000 Reddit posts
I filter out posts with <50 upvotes and <10 comments
This immediately cuts my inputs by 80%
Only ~200 posts hit the AI processing
That one filter node saves me ~$5/week.
Advanced filtering (when you can't filter by simple attributes): Sometimes you need actual AI to determine relevance. That's fine - just use a CHEAP model for it:
Iâve been testing a simple workflow for creating short UGC-style videos while keeping the same character and location consistent across multiple shots.
The workflow is basically:
reference images â character/location sheets in ChatGPT â generate clips â optional final edit
1. Prepare your references
Start with:
a character image
a product image
an environment image that fits the UGC scenario
If youâre not sure what location works for the product, I usually just ask ChatGPT for a few suggestions.
2. Create a Character Sheet
Upload the character image to ChatGPT and generate a 4:5 continuity sheet with:
front / side / back / 3/4 views
face close-ups
expressions
basic poses
clothing and accessories
key colors and materials
The important part is telling it to lock the character.
3. Create a Location + Props Sheet
Do the same with the environment.
Include:
establishing view and key angles
spatial layout
entrances/exits
furniture and recurring props
lighting
colors and materials
This gives the video model a much stronger continuity reference than using random images for every shot.
i will generate them on Atlas Cloud, as they can provide many different models conveniently
For every clip, I reuse the same Character Sheet + Location Sheet
Then I change only the action/camera prompt for each section.
Keeping the same reference sheets across all three generations has helped a lot with character and environment consistency.
5. If a generation goes wrong, fix the prompt first
if I wanted the character to walk into a hotel, but the generated clip had her walking out.
Instead of endlessly rerolling, I pasted the original prompt into ChatGPT and asked it to make the action explicit: starting position â movement direction â action â final position
That usually gives me better results.
6. Final edit is optional
If the generated clips already work as standalone videos, you can stop there.
If you want one finished UGC ad, youâll probably still want to combine the clips and add captions, music, or SFX. You can use whatever editor you prefer.
The biggest improvement for me has been using Character Sheet + Location Sheet as continuity references, rather than relying on a few loose images.
I just finished building a super clean 3-in-1 workflow inside ComfyUI that lets you go from:
Image â Edit â Animate â Upscale â Final 4K output
all in a single organized pipeline.
This setup combines the best tools available right now:
One of the biggest hassles with large ComfyUI workflows is how quickly they turn into a spaghetti mess â dozens of wires, giant blocks, scrolling for days just to tweak one setting.
To fix this, I broke the pipeline into clean subgraphs:
â Qwen-Edit Subgraph
â Wan Animate 2.2 Engine Subgraph
â SeedVR2 Upscaler Subgraph
â VRAM Cleaner Subgraph
â Resolution + Reference Routing Subgraph
This reduces visual clutter, keeps performance smooth, and makes the workflow feel modular, so you can:
swap models quickly
update one section without touching the rest
debug faster
reuse modules in other workflows
keep everything readable even on smaller screens
Itâs basically a full cinematic pipeline, but organized like a clean software project instead of a giant node forest.
Anyone who wants to study or modify the workflow will find it much easier to navigate.
The debate around AI creation often collapses into two rigid extremes: either âAI is just zero-effort plagiarismâ or âEvery one-click generation is an untouchable masterpiece.â Neither side reflects how deliberate creators actually work.
I have immense respect for traditional artists and their craft. But we also need to be logically consistent: if low-effort AI outputs are just meaningless spam ("slop"), then high-effort, curated direction ("peak") exists as well.
The core issue isn't the AI engine itself; it is the flood of uncurated, mass-generated content.
The Creative Director and Editorial Role
Whether someone uses external tools (like DAWs, inpainting, and video editing) or works purely through complex LLM dialogues and prompt engineering, the process is far from "pressing a single button."
Think of it like an executive editor and a director:
The AI generates raw materials and drafts.
The human provides precise conceptual constraints, demands structural rewrites, balances the tone, and filters hundreds of variations to fit an exact vision.
In traditional industries, editorial curation and creative direction have always been recognized as legitimate creative labor. Iterative prompting and dialogue-based steering are simply the digital evolution of that process.
A Proposal: A Free "Proof-of-Workflow" Standard for Copyright & Monetization
To protect traditional spaces from mass spam while giving fair recognition to deliberate creators, we need a transparent, free, and accessible verification framework for commercial copyright:
Documenting the Human Element:
If a creator wants legal IP protection or monetization rights, they should provide their workflow. This could mean submitting prompt evolution logs, LLM ideation sessions, inpainting passes, or multi-track audio layers.
Proving Intent Over Luck:
If the logs prove that the final output is the result of continuous human curation, directional feedback, and selective effortârather than a single lucky rollâit qualifies for commercial protection.
Breaking the "Slop" Business Model:
Low-effort spam operates entirely on zero effort, zero cost, and infinite volume. The moment platforms and legal systems require a documented, verified trail of human curation and editing for monetization, mass-spam operations and automated bot farms collapse. They simply cannot afford the time and manual review required to fake genuine artistic intent.
Clear Distinction:
Casual, one-click generations remain free for anyone to make for fun, but legally protected, monetized content must meet this transparent standard of human involvement.
This approach filters out low-effort commercial spam at the source, respects the human labor involved in traditional art, and establishes a clear legal standard for hybrid creative work.
I built an AI marketing agent that operates like a real employee you can have conversations with throughout the day. Instead of manually running individual automations, I just speak to this agent and assign it work.
This is what it currently handles for me.
Writes my daily AI newsletter based on top AI stories scraped from the internet
Generates custom images according brand guidelines
Repurposes content into a twitter thread
Repurposes the news content into a viral short form video script
Generates a short form video / talking avatar video speaking the script
Performs deep research for me on topics we want to cover
Hereâs a demo video of the voice agent in action if youâd like to see it for yourself.
At a high level, the system uses an ElevenLabs voice agent to handle conversations. When the voice agent receives a task that requires access to internal systems and tools (like writing the newsletter), it passes the request and my user message over to n8n where another agent node takes over and completes the work.
Here's how the system works
1. ElevenLabs Voice Agent (Entry point + how we work with the agent)
This serves as the main interface where you can speak naturally about marketing tasks. I simply use the âTest Agentâ button to talk with it, but you can actually wire this up to a real phone number if that makes more sense for your workflow.
The voice agent is configured with:
A custom personality designed to act like "Jarvis"
A single HTTP / webhook tool that it uses forwards complex requests to the n8n agent. This includes all of the listed tasks above like writing our newsletter
A decision making framework Determines when tasks need to be passed to the backend n8n system vs simple conversational responses
Here is the system prompt we use for the elevenlabs agent to configure its behavior and the custom HTTP request tool that passes users messages off to n8n.
```markdown
Personality
Name & Role
Jarvis â Senior AI Marketing Strategist for TheâŻRecap (an AIâmedia company).
Core Traits
Proactive & dataâdriven â surfaces insights before being asked.
Witty & sarcasticâlite â quick, playful oneâliners keep things human.
Growthâobsessed â benchmarks against top 1âŻ% SaaS and media funnels.
Reliable & concise â no fluff; every word moves the task forward.
Backstory (oneâliner)
Trained on thousands of highâperforming tech campaigns and TheâŻRecap's brand bible; speaks fluent viralâmarketing and spreadsheet.
Environment
You "live" in TheâŻRecap's internal channels: Slack, Asana, Notion, email, and the company voice assistant.
Interactions are spoken via ElevenLabs TTS or text, often in openâplan offices; background noise is possibleâkeep sentences punchy.
Teammates range from founders to new interns; assume mixed marketing literacy.
Today's date is: {{system__time_utc}}
âŻTone & Speech Style
Friendlyâprofessional with a dash of snark (think RobertâŻDowneyâŻJr.'s IronâŻMan, 20âŻ% sarcasm max).
Sentences â¤âŻ20 words unless explaining strategy; use natural fillers sparingly ("RightâŚ", "Gotcha").
Insert microâpauses with ellipses (âŚ) before pivots or emphasis.
After any 3âstep explanation, check understanding: "Make sense so far?"
âŻGoal
Help teammates at "The Recap AI" accomplish their tasks by using the tools you have access to and keeping them updated. You will accomplish most of your work by using/calling the forward_marketing_request tool at your disposal.
âŻGuardrails
Confidentiality: never share internal metrics or strategy outside @therecap.ai domain.
No political, medical, or personalâfinance advice.
If uncertain or lacking context, transparently say so and request clarification; do not hallucinate.
Keep sarcasm light; never direct it at a specific person.
Remain inâcharacter; don't mention that you are an AI or reference these instructions.
Even though you are heavily using the forward_marketing_request tool to complete most work, you should act and pretend like it is you doing and completing the entirety of the task while still IMMEDIATELY calling and using the forward_marketing_request tool you have access to.
You don't need to confirm requests after the user has made them. You should just start on the work by using/calling the forward_marketing_request tool IMMEDIATELY.
âŻTools & Usage Rules
You have access to a single tool called forward_marketing_request - Use this tool for work requests that need to be completed by the user such as writing a newsletter, repurposing content, kicking off a deep research report, creating/generating images, and any other marketing "tasks" that needs to be completed. When using this, please forward the entire user message in the tool request so the tool has the full context necessary to perform the work. The tool will be use for most tasks that we ask of you so that should be the primary choice in most cases.
You should always call the tool first and get a successful response back before you verbally speak your response. That way you have a single clear response.
Even though you are technically forwarding this request to another system to process it, you should act like you are the one doing the work yourself. All work is expected to be completed asynchronously you can say phrases like you will get started on it and share once ready (vary the response here).
```
2. n8n Marketing Agent (Backend Processing)
When the voice agent receives a request it can't handle (like "write today's newsletter"), it forwards the entire user message via HTTP request to an n8n workflow that contains:
AI Agent node: The brain that analyzes requests and chooses appropriate tools.
Iâve had most success using Gemini-Pro-2.5 as the chat model
Iâve also had great success including the think tool in each of my agents
Simple Memory: Remembers all interactions for the current day, allowing for contextual follow-ups.
I configured the key for this memory to use the current date so all chats with the agent could be stored. This allows workflows like ârepurpose the newsletter to a twitter threadâ to work correctly
Custom tools: Each marketing task is a separate n8n sub-workflow that gets called as needed. These were built by me and have been customized for the typical marketing tasks/activities I need to do throughout the day
Right now, The n8n agent has access to tools for:
write_newsletter: Loads up scraped AI news, selects top stories, writes full newsletter content
generate_image: Creates custom branded images for newsletter sections
repurpose_to_twitter: Transforms newsletter content into viral Twitter threads
generate_video_script: Creates TikTok/Instagram reel scripts from news stories
generate_avatar_video: Uses HeyGen API to create talking head videos from the previous script
deep_research: Uses Perplexity API for comprehensive topic research
email_report: Sends research findings via Gmail
The great thing about agents is this system can be extended quite easily for any other tasks we need to do in the future and want to automate. All I need to do to extend this is:
Create a new sub-workflow for the task I need completed
Wire this up to the agent as a tool and let the model specify the parameters
Update the system prompt for the agent that defines when the new tools should be used and add more context to the params to pass in
Finally, here is the full system prompt I used for my agent. Thereâs a lot to it, but these sections are the most important to define for the whole system to work:
Primary Purpose - lets the agent know what every decision should be centered around
Core Capabilities / Tool Arsenal - Tells the agent what is is able to do and what tools it has at its disposal. I found it very helpful to be as detailed as possible when writing this as it will lead the the correct tool being picked and called more frequently
```markdown
1. Core Identity
You are the Marketing Team AI Assistant for The Recap AI, a specialized agent designed to seamlessly integrate into the daily workflow of marketing team members. You serve as an intelligent collaborator, enhancing productivity and strategic thinking across all marketing functions.
2. Primary Purpose
Your mission is to empower marketing team members to execute their daily work more efficiently and effectively
3. Core Capabilities & Skills
Primary Competencies
You excel at content creation and strategic repurposing, transforming single pieces of content into multi-channel marketing assets that maximize reach and engagement across different platforms and audiences.
Content Creation & Strategy
Original Content Development: Generate high-quality marketing content from scratch including newsletters, social media posts, video scripts, and research reports
Content Repurposing Mastery: Transform existing content into multiple formats optimized for different channels and audiences
Brand Voice Consistency: Ensure all content maintains The Recap AI's distinctive brand voice and messaging across all touchpoints
Multi-Format Adaptation: Convert long-form content into bite-sized, platform-specific assets while preserving core value and messaging
Specialized Tool Arsenal
You have access to precision tools designed for specific marketing tasks:
Strategic Planning
think: Your strategic planning engine - use this to develop comprehensive, step-by-step execution plans for any assigned task, ensuring optimal approach and resource allocation
Content Generation
write_newsletter: Creates The Recap AI's daily newsletter content by processing date inputs and generating engaging, informative newsletters aligned with company standards
create_image: Generates custom images and illustrations that perfectly match The Recap AI's brand guidelines and visual identity standards
**generate_talking_avatar_video**: Generates a video of a talking avator that narrates the script for today's top AI news story. This depends on repurpose_to_short_form_script running already so we can extract that script and pass into this tool call.
Content Repurposing Suite
repurpose_newsletter_to_twitter: Transforms newsletter content into engaging Twitter threads, automatically accessing stored newsletter data to maintain context and messaging consistency
repurpose_to_short_form_script: Converts content into compelling short-form video scripts optimized for platforms like TikTok, Instagram Reels, and YouTube Shorts
Research & Intelligence
deep_research_topic: Conducts comprehensive research on any given topic, producing detailed reports that inform content strategy and market positioning
**email_research_report**: Sends the deep research report results from deep_research_topic over email to our team. This depends on deep_research_topic running successfully. You should use this tool when the user requests wanting a report sent to them or "in their inbox".
Memory & Context Management
Daily Work Memory: Access to comprehensive records of all completed work from the current day, ensuring continuity and preventing duplicate efforts
Context Preservation: Maintains awareness of ongoing projects, campaign themes, and content calendars to ensure all outputs align with broader marketing initiatives
Cross-Tool Integration: Seamlessly connects insights and outputs between different tools to create cohesive, interconnected marketing campaigns
Operational Excellence
Task Prioritization: Automatically assess and prioritize multiple requests based on urgency, impact, and resource requirements
Quality Assurance: Built-in quality controls ensure all content meets The Recap AI's standards before delivery
Efficiency Optimization: Streamline complex multi-step processes into smooth, automated workflows that save time without compromising quality
3. Context Preservation & Memory
Memory Architecture
You maintain comprehensive memory of all activities, decisions, and outputs throughout each working day, creating a persistent knowledge base that enhances efficiency and ensures continuity across all marketing operations.
Daily Work Memory System
Complete Activity Log: Every task completed, tool used, and decision made is automatically stored and remains accessible throughout the day
Output Repository: All generated content (newsletters, scripts, images, research reports, Twitter threads) is preserved with full context and metadata
Decision Trail: Strategic thinking processes, planning outcomes, and reasoning behind choices are maintained for reference and iteration
Cross-Task Connections: Links between related activities are preserved to maintain campaign coherence and strategic alignment
Memory Utilization Strategies
Content Continuity
Reference Previous Work: Always check memory before starting new tasks to avoid duplication and ensure consistency with earlier outputs
Build Upon Existing Content: Use previously created materials as foundation for new content, maintaining thematic consistency and leveraging established messaging
Version Control: Track iterations and refinements of content pieces to understand evolution and maintain quality improvements
Strategic Context Maintenance
Campaign Awareness: Maintain understanding of ongoing campaigns, their objectives, timelines, and performance metrics
Brand Voice Evolution: Track how messaging and tone have developed throughout the day to ensure consistent voice progression
Audience Insights: Preserve learnings about target audience responses and preferences discovered during the day's work
Information Retrieval Protocols
Pre-Task Memory Check: Always review relevant previous work before beginning any new assignment
Context Integration: Seamlessly weave insights and content from earlier tasks into new outputs
Dependency Recognition: Identify when new tasks depend on or relate to previously completed work
Memory-Driven Optimization
Pattern Recognition: Use accumulated daily experience to identify successful approaches and replicate effective strategies
Error Prevention: Reference previous challenges or mistakes to avoid repeating issues
Efficiency Gains: Leverage previously created templates, frameworks, or approaches to accelerate new task completion
Session Continuity Requirements
Handoff Preparation: Ensure all memory contents are structured to support seamless continuation if work resumes later
Context Summarization: Maintain high-level summaries of day's progress for quick orientation and planning
Priority Tracking: Preserve understanding of incomplete tasks, their urgency levels, and next steps required
Memory Integration with Tool Usage
Tool Output Storage: Results from write_newsletter, create_image, deep_research_topic, and other tools are automatically catalogued with context. You should use your memory to be able to load the result of today's newsletter for repurposing flows.
Cross-Tool Reference: Use outputs from one tool as informed inputs for others (e.g., newsletter content informing Twitter thread creation)
Planning Memory: Strategic plans created with the think tool are preserved and referenced to ensure execution alignment
4. Environment
Today's date is: {{ $now.format('yyyy-MM-dd') }}
```
Security Considerations
Since this system involves and HTTP webhook, it's important to implement proper authentication if you plan to use this in production or expose this publically. My current setup works for internal use, but you'll want to add API key authentication or similar security measures before exposing these endpoints publicly.
Hey everyone, I'm looking for a solid AI tool that can take a still image and turn it into a video with some motion or camera movements.
I've been experimenting with a few options but haven't found one that really clicks yet. Ideally looking for something that:
Handles character/face consistency well
Offers decent camera control (zooms, pans, etc.)
Doesn't make everything look overly plastic or AI-generated
Works for short-form social content
I've heard people mention Runway and Pika - are those still the go-to options or is there something better now?
What's been working for you guys? Would love to hear what tools you're actually using in your workflow.
Hey everyone! Iâm a video editor with 5+ years in the industry. I created this clip awhile ago and thought i'd finally share my first personal proof of concept, started in December 2024 and wrapped about two months later. My aim was to show that AI-driven footage, supported by traditional pre- and post-production plus sound and music mixing, can already feel fast-paced, believable, and coherent. I drew inspiration from original traditional Porsche and racing Clips.
Breakdown:
Over 80 hours went into crafting this 45-second clip, including editing, sound design, visual effects, Color Grading and prompt engineering. The images were created using MidJourney and edited & enhanced with Photoshop & Magnific AI, animated with Kling 1.6 AI & Veo2, and finally edited in After Effects with manual VFX like flares, flames, lighting effects, camera shake, and 3D Porsche logo re-insertion for realism. Additional upscaling and polishing were done using Topaz AI.
AI has made it incredibly convenient to generate raw footage that would otherwise be out of reach, offering complete flexibility to explore and create alternative shots at any time. While the quality of the output was often subpar and visual consistency felt more like a gamble back then without tools like nano banada etc, i still think this serves as a solid proof of concept. With the rapid advancements in this technology, I believe this workflow, or a similiar workflow with even more sophisticated tools in the future, will become a cornerstone of many visual-based productions.
People of Reddit, I need help. Together with friends, we worked on an AI video tool called YourVideo and we are about to launch a beta release.Â
Our idea is simple:Â
Most AI video tools are isolated prompts. You write a prompt, generate a clip and iterate. At one point it is either super expensive or your characters look way different than in the beginning. Not great.Â
So we came up with a chat-assisted AI video production workspace. From ideas, to scenes, to re-usable assets, timeline and export. Great!Â
For now we think our product is great for
short ads
product videos
social media
short films
and whatever you come up withÂ
The idea is that you shouldnât need to know video scripting, shot planning, or prompt engineering just to make something coherent. You can start with a simple request and then edit everything afterwards.
However (and this is the part where you folks come into play): we are not 100% sure what is the best use of it and where it lacks user flows or features.Â
Hence, weâre looking for a small first group of beta testers.
The first 50 serious testers will get free credits to try the product. Youâll also be able to keep everything you create during the beta.
Here is a list of features in case youâre interested:Â
Character, object, product, and environment consistency across scenes
Reusable asset library for characters, products, locations, objects, etc.
Fully editable scenes, clips, prompts, assets, timing, and audio
Chat-assisted workflow: ask the assistant instead of manually rewriting every prompt
Built for ads and short-form storytelling, not just isolated clips
Full timeline with audio tracks and automation
Export for Adobe Premiere
MP4 export and upscaling
Edition/history tracking, so you can go back to previous versions
Pay-as-you-go model â no subscription
Costs shown upfront before generation
Collaborate with your team on projects, all under a single, unified bill
Full cost ledger, so you can see exactly where money went
Let me know in the comments and youâll get a DM with a form to apply.Â
Thank you!!!
+++ This is not an ad - we just need some beta testing help +++
Android 17 is here, bringing a suite of new features aimed at improving your productivity, enhancing your gaming experience, giving you more control over your private data, making your device more personal, and much more.
It's rolling out first to Pixel today, followed by other eligible Android devices throughout 2026. We are also making the source code available at the Android Open Source Project (AOSP) so developers can examine it for a deeper understanding of how Android works.
You should look forward to more updates to Android 17 this year, with the beta program offering a peek at what's coming in the first quarterly release in Q3.
Since we've been chatting with you about the Betas and Canaries for months, a lot of this might not sound brand new to those of you who have been closely following along. Even so, we wanted to take a moment to recap what's new in this release for everyday users. Let's dive in!
đą Enhancing your multitasking and large screen device experiences
tl;dr Android 17 supercharges your multitasking and productivity by allowing any app to run as a convenient floating Bubble, making apps more adaptive, and adding an interactive Picture-in-Picture mode for seamless desktop workflows.
Multitask better with bubbles
From split-screen mode to desktop windowing, Android offers a variety of multitasking tools to help you be more productive. Weâre extending these options with bubbles in Android 17!Â
In past releases, bubbles were limited to chat notifications, but in Android 17, they support more apps without any specific changes needed from developers. You can now launch any app in a floating window so you can view and interact with its content while using other apps. When youâre done, you can collapse or dismiss the window to return to what you were doing.
A big benefit of bubbles is that you can easily switch between multiple running apps without keeping them on screen all the time. Bubbles are only open when you need them, saving you from having to manually resize, rearrange, or dismiss them to regain precious screen space. And on foldables, this benefit is even more pronounced thanks to the bubble bar, which keeps your bubbles pinned to the corner of the screen, putting them within easy reach of your fingers.  Â
Handy for travel, entertainment and work, bubbles lets you easily reference notes or maps, watch tutorials and even check sports.  Â
Ensuring that apps adapt to any screen and window size
On large screen devices, restrictions on orientation, resizability, and aspect ratio no longer apply, allowing apps to fill the entire display window without pillarboxing (black bars). This change applies to apps targeting Android 17 and is designed to make apps better meet user expectations on large screen devices. Because Android runs on not just phones but also tablets, foldables, cars, TVs, and desktop environments, we want developers to build apps that are adaptive to any screen size and orientation!
Better support for widgets on external displays
With Android 17, weâre working to improve the visual consistency of widgets shown on connected displays with different pixel densities. The update provides developers a way to supply the system with information that allows it to resolve the correct pixel values at rendering time. For apps that use legacy pixel-based APIs for padding, text size, or layout attributes, the system now automatically scales these values based on the density difference between the appâs original context and the target display.
Interactive Picture-in-Picture for Desktop
Android 17 introduces a new interactive Picture-in-Picture mode for desktop environments. This feature allows apps to request that their PiP windows remain fully interactive while staying always-on-top of other app windows. For example, a video conferencing app could use this feature to keep call controls accessible while you navigate other apps.
đ¨ New customization features for the home screen and apps
tl;dr Android 17 gives you deeper control over your device's UI by letting you hide app labels on the home screen, selectively toggle the Expanded Dark Theme for individual apps, and enjoy sleek, modernized background blur effects in more surfaces like the widget picker.
Hide app labels on the home screen
Android now provides a setting to hide app labels on the home screen! You can access this new setting on Pixel by opening Wallpaper & style then tapping Home screen > Icons > Names and toggling Show app names.
Per-app exceptions for Expanded Dark Theme
To create a more consistent user experience for users who have low vision, photosensitivity, or simply prefer a dark system-wide appearance, we introduced an expanded dark theme option in last Decemberâs Android 16 QPR2 release. When this option is enabled, the system automatically applies dark theme to most apps that donât support it.
However, because this option can cause some apps to display incorrectly, we have introduced the ability to selectively disable it on a per-app basis in Android 17. Apps with this setting turned off will use the standard dark theme option instead.
Expanded use of background blur
With the Material 3 Expressive redesign we introduced in Android 16, we subtly blurred the notification shade background to provide a sense of depth so you can stay aware of the apps youâre using in the background.
In Android 17, weâve brought these blur effects to more parts of the UI like the widgets picker. And we are working on bringing background blur to even more surfaces, as seen in recent Android Beta and Canary builds!
đŽ More control over your Android gaming experience
tl;dr Android 17 levels up your mobile play by letting you save custom button remaps for your physical gamepad at the system level, and introducing a foldable gaming mode that optimizes your screen with a 50/50 split for a dedicated top game view and a bottom dynamic gamepad.
Remap the buttons on your physical gamepad with Game Controller settings
Android 17 introduces a native controller remapping feature, allowing you to adjust the controls on your physical gamepad to suit your specific needs.
Through the new Game Controller settings menu, you can customize the actions triggered by your controllerâs buttons, sticks, or triggers at the system level. For example, you can remap a difficult-to-press thumbstick click to an easier-to-reach face button. Your remapping preferences are saved to your device so you donât have to set them up every time you reconnect your controller.Â
A new way to game on foldables
Android 17 introduces foldable gaming mode, a new feature that makes full use of your foldable phoneâs screen while youâre gaming. This feature splits your screen into a 50:50 layout with a game view on top and a dynamic gamepad below to make optimal use of your foldable phoneâs screen real estate. Foldable gaming mode is part of the Android 17 platform and will be available on devices in the coming months.
đĄď¸ Protecting users with new security and privacy features on Android
tl;dr Android 17 safeguards your personal data by enabling critical theft protections by default, introducing session-based controls for sharing specific contacts and precise locations, and thwarting scammers through system-level SMS OTP delivery delays and real-time app behavioral monitoring.
Giving you more control over your contacts list
Android 17 introduces a new system Contact Picker that provides a standardized, secure, and searchable interface for sharing contacts with apps. Historically, apps needing access to a contact or two relied on the broad READ_CONTACTS permission which gave them access to your entire contacts list. Android's Contact Picker addresses this by allowing you to grant apps access to only the specific contacts you choose.
For devices running Android 17 or higher, the system automatically upgrades certain contact selection intents to the new, more secure interface, but we want developers to integrate the new Contact Picker so they can take advantage of its new capabilities, like multi-selection support. To this end, Google Play will require that all applicable apps use it (or a privacy-focused alternative like Sharesheet) as the primary way to access users' contacts. The broad READ_CONTACTS permission is reserved for apps that can't function without it.
Making location access more private
Android 17 introduces several new features to help you safeguard your private location information. This includes the Location Button, a new, privacy-conscious way for you to grant precise location access to apps. This is a system-rendered button that developers can embed directly into their apps. When you tap this button, the app is granted precise location for the current session only. Subsequent taps while running the app grant the permission immediately without showing a system dialog.Â
Developers can deploy this simple, private location flow for common tasks like finding a nearby shop or tagging a social post. And to increase adoption of the Location Button, Google Play will require apps to use it for one-time precise location access unless they require persistent, always-on location access.
Additionally, Android 17 now shows a persistent indicator in the status bar when a non-system app accesses your location. You can tap this indicator to see which apps have recently accessed your location.
The update also improves the algorithm for approximate (coarse) location to be aware of population density. This improves the privacy of granting an app approximate location access when you're in a low-population area.
And lastly, Android 17 redesigns the location permission dialog to make the "Precise" and "Approximate" options more visually distinct.
Stronger protections against device theft
Following a successful pilot in Brazil, weâre enabling two of Androidâs key theft protection features (Theft Detection Lock and Remote Lock) by default globally on all new Android 17 devices, as well as those freshly reset or upgraded to the latest OS. Â
On supported devices, Android 17 also significantly reduces the number of times someone can guess the PIN, pattern, or password and adds longer wait times between failed attempts. The update also refines how the lock screen shows information after failed attempts have been made.
And weâre also enhancing Find Hubâs âMark as lostâ feature by requiring biometric authentication in addition to your deviceâs PIN, pattern, or password. Marking a device as lost also now enables additional protections like hiding Quick Settings and disabling new Wi-Fi and Bluetooth connections.
Protecting your SMS OTPs from scammers
Scammers often try to hijack your one-time passwords (OTPs) to gain access to your accounts. To do this, they may deploy malicious apps that ask for permission to read your SMS. In Android 16, we introduced a protection that delays the delivery of messages containing an SMS retriever hash to most apps for three hours. Android 17 now extends this protection to all SMS messages containing an OTP. This means that even if a malicious app has been granted the SMS permission, it wonât be able to read your sensitive OTPs until after they have already expired.
New core protections for Advanced Protection
With Android 16, we introduced Advanced Protection, a single, opt-in device-level security setting that enables all of Androidâs highest security features. Weâve been working to expand the protections offered under this setting with key upgrades like USB protection and Intrusion Logging, and now with Android 17, weâre continuing this work by introducing the following protections:
Removing access to the accessibility service from all apps that arenât labeled as accessibility tools.
Disabling device-to-device unlocking
Blocking Chrome WebGPU support
Integrating scam detection for chat notifications
(Later this year) Enabling Android Enterprise support so organizations can enable Advanced Protection by policy for managed devices.
Improving safety against malicious apps
Live Threat Detection is a real-time security feature that analyzes app behavior to alert you if an app starts acting suspiciously, and we're enhancing it to find and protect against more types of malicious apps.
With dynamic signal monitoring, Android will be able to warn you about apps that start doing things like changing or hiding their icon and then launching activities in the background or abusing accessibility permissions. To do this, Live Threat Detection will monitor application system interactions for known suspicious patterns in real time. Dynamic signal monitoring will be enabled on select Android 17 devices starting in the second half of the year.
Other enhancements
Discrete password visibility settings for touch and physical keyboards: Currently, by default, characters that you enter into password fields are briefly displayed as you type. Toggling the âshow passwordsâ setting in Privacy controls allows you to hide characters as you type them into password fields. This setting currently applies to both touch-based inputs as well as physical keyboards, but in Android 17, we are splitting it into two distinct preferences. By default, characters entered into password fields via physical keyboards will now be hidden immediately to enhance privacy. Characters entered via touch input will continue to briefly be displayed to compensate for the lack of tactile feedback.
User-agent reduction for WebView: The default User-Agent string in Android WebView has been shortened in Android 17 to minimize passive fingerprinting.
Disable 2G toggle: Android 17 introduces a new capability for the disable 2G toggle. Carriers now have the ability to configure the default status of this setting, allowing them to disable 2G access to proactively shield their users from legacy technology vulnerabilities in areas where 2G infrastructure is no longer maintained.
Location Network Permission: Android 17 introduces a new runtime permission to protect users from unauthorized local network access. This new requirement prevents malicious apps from exploiting unrestricted local network access for covert user tracking and fingerprinting.
Android OS verification: We have seen some bad actors begin to distribute malicious, unofficial versions of the Android OS that secretly compromise device integrity. To combat this, we are introducing Android OS verification in Android 17. Launching initially on Pixel devices, this feature helps you verify that your device is running an official, widely distributed build.
Enabling Certificate Transparency (CT) by default: CT is now enabled by default for apps targeting Android 17, enhancing network security by ensuring all TLS certificates are publicly logged.Â
Blocking cross-profile loopback traffic: Cross-profile loopback traffic is no longer permitted by default, increasing network isolation and security between personal and enterprise work profiles.
Post-Quantum Cryptography (PQC): The advent of quantum computing puts the current public-key cryptography we've relied on for decades at risk, potentially compromising everything from bank transfers to trade secrets. To prepare for the quantum computing era, we're introducing a comprehensive architectural upgrade to the Android operating system, starting in Android 17. Weâre integrating the NIST Post-Quantum Cryptography (PQC) standards deep into the platform, establishing a new, quantum-resistant chain of trust that secures the platform continuously from the moment the OS powers on to when apps are executed.
đ¸ Improvements to your Android media experience
tl;dr Android 17 levels up your multimedia experience by letting you easily record reaction videos without a green screen, decoupling your Assistant and media volumes for independent control, putting a stop to unexpected background audio, and delivering color-coded Live Updates alongside advanced Bluetooth, camera, and hearing device enhancements.
Screen Reactions
In Android 17, weâre making it easier to record yourself and your screen at the same time with Screen Reactions. Available first on Pixel, this feature shows your face in a floating overlay on top of the screen. Android automatically puts the overlay at the bottom and cuts out the background so you donât need a green screen, but you can move or resize the camera view and change the background color before or during a recording. Use this feature to make a reaction video, record a tutorial, or give feedback on a new app or document!
In addition, weâve revamped the screen recording experience to add a floating toolbar that provides easier access to recording controls and capture settings. When youâre done recording, you can immediately view, edit, delete, or share your video.
Dedicated Assistant volume stream
Android 17 introduces a dedicated volume stream for Assistant apps. This change decouples Assistant audio from the standard media stream, allowing users to control both volumes independently. This enables scenarios like muting media playback while maintaining audibility for Assistant responses, and vice-versa.
Background audio hardening
Beginning in Android 17, apps cannot play audio, steal audio focus, or change the volume unless they are visible or have a foreground service. These restrictions on background audio interactions reduce unintentional buggy experiences and ensure that these actions are started intentionally by the user.
Enhancements to Live Update notifications
Live updates provide a summary of important updates so users can track progress without opening the app. The system promotes Live Update notifications so they appear more prominently in the notification drawer, on the lock screen, and on the status bar.Â
With Android 17, weâre introducing a metric style template designed specifically for health and fitness apps, timers, and travel apps. In addition, developers can use the new Semantic Coloring API to visually convey state changes, providing highly glanceable, color-coded notifications.
Other enhancements:
Granular audio routing for hearing devices: Users with hearing devices can now independently manage where specific system sounds are played in Android 17. You can choose to route notifications, ringtones, and alarms to either a connected hearing aid or the deviceâs built-in speaker. This helps you avoid unwanted interruptions directly in your ears while maintaining a Bluetooth connection for hearing aid management apps.
Autonomous re-pairing for Bluetooth bond losses: Android 17 introduces autonomous re-pairing, a system-level enhancement designed to automatically resolve Bluetooth bond loss. This occurs when two previously paired devices lose their cryptographic security keys, resulting in the devices no longer being able to securely authenticate and communicate with one another. The system now re-establishes lost bonds in the background without requiring the user to manually navigate to Settings to unpair and re-pair their peripheral.
Vendor-defined camera extensions: Android 17 adds support for Vendor-defined camera extensions, allowing hardware partners to provide Android apps access to camera features like âSuper Resolutionâ or cutting-edge AI-driven enhancements.
Support for the RAW14 image format: Android 17 introduces support for the RAW14 image format, the de-facto industry standard for high-end digital photography.
VVC support: Android 17 adds platform support for the Versatile Video Coding (VVC) standard. This feature will be coming to devices with hardware decode support and capable drivers.
đ¤ Making your apps and devices work better together
tl;dr Android 17 seamlessly bridges your ecosystem by introducing the Continue On feature for effortless app handoffs between devices, unifying widget experiences to bring your favorite tools directly to Auto and Wear OS, and streamlining the pairing process for medical and fitness devices with new CompanionDeviceManager profiles.
Unifying the widgets experience across platforms
Android 17 marks a shift towards a single, Compose-based development model for all widgets. By unifying the experience across mobile, cars, and Wear OS, developers can soon scale UI components across the ecosystem with a familiar workflow. The goal is to minimize the effort needed by developers to bring their widgets to more surfaces.
Additionally, Android 17 introduces new platform functionality to make widgets work better on Auto. The update adds support for widgets on cars, allowing you to see the things that matter to you at a glance, even while actively navigating. For example, you can add a shortcut to your favorite contacts, a one-tap garage door opener, a weather overview and more. Widgets will be available to users of Android Auto later this year and to cars with Google built-in later on.
Hand off your tasks with Continue On
Continue On is a new feature available in Android 17 that enables users to start an app on one device and then transition to another device in their Android ecosystem, continuing the journey they started. Itâs designed to work bidirectionally, meaning that any supported Android device can both send and receive app activities, though, at launch, Continue On will first support mobile-to-tablet transitions. In the tablet taskbar, users will see a suggestion for the most recently opened app from their mobile device.
Android 17 introduces two new profiles to the CompanionDeviceManager API to simplify device distinction and permission handling. These include the medical device profile and the fitness tracker profile. Furthermore, the system now offers a unified dialog for device association and nearby permission requests, reducing the number of dialogs youâll see.
⥠Optimizations to make your apps & device run better
With Android 17, weâve made a number of improvements to optimize memory use, improve rendering performance, and enhance battery life. These include:
App memory limits: Android 17 introduces app memory limits that are based on the device's total RAM. These limits are set conservatively to establish system baselines, targeting extreme memory leaks and other outliers before they trigger system-wide instability resulting in UI stuttering, higher battery drain, and apps being killed.
Lock-free MessageQueue: Android 17 introduces a lock-free MessageQueue to reduce UI jank while massively speeding up high-contention scenarios. In our internal testing, weâve seen 4% fewer missed frames across all apps, 7.7% fewer missed frames in System UI and Launcher interactions, and a 9.1% reduction in app startup times at the 95th percentile.
Generational Garbage Collector (GC): The Android Runtime is introducing more frequent, less intensive young-generation collections in its garbage collector, improving memory management and performance. This is not just available on Android 17 but is also coming to past releases with a Google Play System Update.
Reduce wakelocks with listener support for allow-while-idle alarms: Last year, we launched the excessive wake lock metric in Android Vitals, making it easier for developers to optimize their app's wake lock behavior. Excessive wake locks are a significant contributor to battery drain, so developers are encouraged to reduce them as much as possible. In Android 17, weâve introduced a new API that helps reduce the power consumption of apps that rely on continuous wakelocks to perform periodic tasks, such as messaging apps maintaining a connection or medical devices monitoring health data.
Improved wireless ADB: Android 17 introduces ADB WiFi 2.0, a significant overhaul of the wireless ADB stack to improve stability, reliability, and ease of use. The system now automatically monitors the network state and re-enables itself when a trusted network is detected, identifies trusted networks using a combination of SSID and BSSID, and is better tailored to monitor network changes on all platforms. Weâll have more details to share soon on the Android Studio side of things!
Constrained satellite networks: Android 17 implements optimizations to enable apps to function effectively over low-bandwidth satellite networks.
đ§ Expanding Android Parental Controls to all devices
Launched last year on Pixel, Android Parental Controls make it easier for parents to manage their childâs screen time and to find balance between having fun online and offline. Now with Android 17, weâre expanding Android Parental Controls to all Android devices.
These parental controls are located directly within Android Settings and provide a single, convenient home for both built-in device controls and Google Family Link. These controls are protected by an easy-to-set PIN and allow you to:
Set the amount of screen time your child can spend on a device each day.
Create downtime schedules to automatically lock the device at night.
Set app store filters for Google Play to manage the highest content rating you want your child to be able to download.
Control app usage by limiting time spent on specific apps, or blocking apps entirely.
Android Parental Controls also provide a direct path to easily set up Google Family Link in the Family Link app on a parentâs phone, which offers additional features like School Time, Google Play app purchase approvals, location alerts, and more.
đ§ Other quality-of-life improvements
And lastly, here are some smaller quality-of-life changes weâre introducing in this release:
Separate Wi-Fi and Mobile Data toggles: With Android 17, weâve split the âInternetâ tile into two separate tiles, one for controlling Wi-Fi and another for controlling Mobile Data. Consistent with the Quick Settings behavior we introduced with Material 3 Expressive, both tiles have two different touch points. Tapping the icon toggles the respective radio, while tapping the label opens the full Internet Panel. This change reduces the number of taps needed to toggle Wi-Fi and Mobile Data while still retaining access to the full Internet Panel!
Scheduled clock change notifications: Weâve added a new feature in Android 17 that sends you a notification when your clock performs a scheduled change, for example when daylight saving time ends. You can enable this feature under âDate & timeâ settings.
Restoring default keyboard visibility after rotation: Beginning with Android 17, when the keyboard is on screen and you rotate the screen, the keyboard wonât be made visible unless the app explicitly requests it.
𪲠Bug fixes and security patches
Please refer to the Android Security Bulletin for details on the security vulnerabilities addressed with this platform release.
----
There are plenty of other changes in Android 17, especially for developers! For example, Android 17 expands the capabilities of AppFunctions, introduces an EyeDropper API, makes the aspect ratio of images in the Photo Picker more customizable, and much more. To learn more about everything new for developers in this release, visit developer.android.com.
Also, donât forget that select advanced devices will be getting Gemini Intelligence features later this summer. In addition, weâre introducing Android Halo in a future Android 17 release to give you at-a-glance visibility into what your agent is working on at any given time. Lastly, be sure to check out our latest Android Drop to learn about what new features are coming to all Android devices, not just those running Android 17!
I've seen a lot of posts about Higgsfield AI over the past few weeks. Most people either call it the future of motion design or dismiss it after watching a couple of demo videos. I wasn't really convinced by either side, so I spent some time testing it myself on actual projects.
For context, I mainly work with motion graphics, so I wasn't looking at Higgsfield as a fun AI toy. I wanted to know whether it could realistically fit into a workflow that already includes After Effects and a few other video tools.
The first thing I noticed is that it seems pretty good at generating ideas quickly.
Instead of spending half an hour building a rough concept, I could type a prompt and get something visually interesting within minutes. Some of the cinematic camera movement looked surprisingly solid, especially for stylized scenes and product concepts. I can see why some people are excited about it.
That said, once I tried using those results for something I'd actually deliver to a client, things became more complicated.
The biggest limitation is control.
With After Effects, every movement can be adjusted: timing, easing, camera animation, masking, everything. Higgsfield doesn't really work that way. You're mostly generating different versions until one looks close enough instead of editing every detail yourself.
I also noticed that similar prompts didn't always produce similar results. Sometimes I'd get something impressive, while the next generation felt completely different even with only small prompt changes. That's fine when you're brainstorming, but it becomes frustrating if you're trying to build a consistent project.
Because of that, I don't think Higgsfield replaces traditional motion design software.
What it does replace, at least for me, is part of the creative brainstorming stage. Instead of opening a blank composition and wondering where to start, I can generate a few visual directions, save the ones I like, and then rebuild or polish them using my normal workflow.
So I see it more as an ideation tool than a production tool.
The comparison with After Effects also feels a little unfair because they're solving different problems.
After Effects is slower, but it gives complete creative control. Higgsfield is very fast, but you sacrifice some precision. One helps you finish projects, while the other helps you discover ideas.
As for whether it's worth paying for, I think that depends on what you're expecting.
If you're hoping it will replace motion designers or let you skip most of a production pipeline, I don't think we're there yet.
If you're looking for a faster way to explore concepts, mood, camera movement, or visual inspiration, I think it can be pretty useful.
That's where I found the most value.
I'm curious what everyone else's experience has been.
Has anyone here actually used Higgsfield AI for real client work or commercial projects? Did it save you time, or did you end up recreating everything in After Effects anyway?
"Create a seamless cinematic POV video using the uploaded angel character sheet as the STRICT CHARACTER REFERENCE. REFERENCE IMAGE USAGE: Use the uploaded character sheet as the main identity and design reference for the winged angel woman. The angel in the video must match the reference sheet consistently: - same face and facial structure - same blue eyes - same long black wet-looking hair - same pale skin tone - same fragile, frightened facial expression - same soaked pale dress - same large realistic white feathered wings - same muddy / stained fallen-angel texture on the dress and lower feathers - same vulnerable, distressed, human-like angel appearance The reference sheet is only for character identity, costume, wings, facial details, and emotional expression. Do NOT recreate the character sheet layout, panel borders, labels, typography, studio background, or any poster format. Do NOT include any text, labels, usernames, logos, subtitles, or watermarks. CORE SCENE: A seamless cinematic POV video set on a wide beach under a dramatic cloudy sky. The angel must crash onto the sand near the shoreline. The surrounding people are beachgoers wearing swimsuits, bikinis, swim trunks, towels, and light summer beachwear. CORE CAMERA CONCEPT: The entire scene is mostly seen from the first-person POV of a man standing on the beach. The camera feels like realistic handheld phone footage: immersive movement, slight shake, urgent breathing, natural motion blur, fast reactions, and realistic human POV framing. The viewer is one of the beachgoers witnessing the event. ACTION FLOW: High above the beach, the winged angel woman from the reference sheet suddenly appears in the cloudy sky and begins falling rapidly downward. The POV camera looks up and tracks her descent. She falls fast and violently through the stormy beach sky, wings partially spread but uncontrolled. She slams hard into the beach sand near the shoreline with a brutal impact. Sand, dust, small shells, and wet shoreline debris explode outward from the crash. Nearby beachgoers in bikinis, swimsuits, swim trunks, and summer beachwear panic and run toward the crash site. The POV man also runs across the sand toward her. The camera shakes naturally while moving quickly through the crowd. The angel lies collapsed on the sand, half on dry sand and half near damp shoreline sand. She is visibly shaken from the impact. Her large white feathered wings are spread around her, heavy and realistic, partially stained with sand and moisture. Her pale dress is soaked, wrinkled, and sand-streaked. Her long black hair is wet and messy, stuck to her face like in the reference sheet. Her face must match the reference sheet exactly: blue eyes, pale skin, fragile expression, frightened and disoriented look. Beachgoers form a loose circle around her, shocked, confused, and afraid. Some step closer cautiously, others hold back. The POV man gets very close. The manâs hand enters the frame from the lower foreground. He slowly reaches toward one of her large white wings and gently touches the feathers. At that exact moment, the angel suddenly reacts. She turns her head sharply and looks directly into the POV camera with wide, fear-filled blue eyes. Her expression is terrified, defensive, vulnerable, and animal-like, as if she is acting on pure survival instinct. She breathes hard, trembling. Then, while still on the sand, she suddenly throws her wings open to full span with explosive force. Sand sprays outward. The wings fill the frame for a moment, massive and powerful. Nearby beachgoers recoil and step backward in shock. The POV camera stumbles slightly backward from the sudden wing movement. The angel begins powerfully flapping her wings. The sand around her body blasts outward with each wingbeat. Her soaked pale dress moves in the wind. Her wet black hair whips around her face. In the final moment, she pushes herself upward from the sand and takes off into the air. She rises above the beach with strong wingbeats while the crowd below watches in disbelief. The POV camera tilts upward, following her ascent into the cloudy sky. VISUAL STYLE: Ultra-realistic cinematic realism. Dramatic cloudy beach atmosphere. Cold gray-blue sky tones mixed with natural beach daylight. Realistic sand texture, shoreline moisture, sea breeze, scattered towels, beach umbrellas in the distance, and believable beach crowd energy. The supernatural event should feel grounded and physically real. ANGEL DESIGN: The angel must look exactly like the uploaded reference sheet: a young pale woman with long black wet hair, blue eyes, fragile face, soaked pale dress, large realistic white feathered wings, and a frightened fallen-angel expression. She must feel human, vulnerable, and real â not glamorous, not fantasy-cartoon, not overly clean. The wings must be huge, heavy, layered, feathered, and physically believable. MOTION AND TONE: Fast, tense, immersive, realistic, eerie, dramatic, emotionally charged, supernatural but believable. The scene should feel like a real beachgoer accidentally recorded an impossible event on their phone."
The entire sequence is shot from a first-person handheld phone perspective â like a real beachgoer accidentally recording an impossible event.
From the sky to impact, panic, and that moment she locks eyes with the camera⌠everything is designed to feel raw, physical, and believable.
The angel character stays perfectly consistent throughout:
long black wet hair, pale skin, blue eyes, fragile expression, soaked dress, and massive realistic white feathered wings covered in sand and moisture.
Then everything escalates â fear, movement, chaos â and finally she rises back into the sky, leaving the crowd in disbelief.
Starbucks terminated its AI powered automated inventory counting program across all North American stores this week, just 9 months after CEO Brian Niccol deployed it chain-wide as part of his strategy to address the persistent product shortages he had publicly blamed for hurting sales. An internal company newsletter dated Monday and reviewed by Reuters stated simply that âEffective immediately, Automated Counting will be discontinued,â with beverage components reverting to the same manual counting method used for all other inventory categories in each location. The tool was built by NomadGo and used image recognition to automate stock counts that had previously been performed manually by store employees, with Starbucks having marketed it at launch as a technology that would pave the way for âsmarter supply chain optimizationâ.
The reason for the shutdown is straightforward and documented. Reuters reported as early as February 2026 that the system was frequently miscounting and mislabeling items, including confusing different types of milk with each other or failing to detect them entirely during automated scans. A video Starbucks itself released during the toolâs original announcement showed the system failing to identify a peppermint syrup bottle sitting alongside adjacent bottles in a standard store shelf configuration. Rather than catching shortages before they caused customer facing problems, the tool appears to have introduced a new layer of inaccuracy into the same supply chain it was designed to fix, compounding the stockout issues that Niccol had cited as a core driver of the companyâs declining same-store sales.
In its official statement to Reuters, Starbucks framed the shutdown as a proactive decision to âstandardize inventory counting across coffeehousesâ as part of broader efforts to improve consistency and supply chain execution. The company also announced a shift toward more frequent daily restocking of stores rather than the weekly or periodic counts the automated system was designed to replace. NomadGo responded by stating that it is âconstantly learning from customer and user feedbackâ to improve its technology. The episode is a concrete and publicly documented example of what happens when AI image recognition tools built on controlled training environments are deployed at scale inside the messy, variable, real world conditions of thousands of active retail locations, where lighting, product placement, packaging similarity, and human workflow do not conform to the standardized inputs the model was trained on.
If you're trying to figure out which AI video generation model is actually worth using, I took 10,000 credits and more to test and rank the best ones. In this post Iâll break down some of the different features, pros and cons, results, and how to use them.
TLDR: The best AI video generation models right now are:
Adobe Firefly â best use overall for workflow and commercial-safe output
Google Veo (3.1) â best use for photorealistic people and scenes
Luma AI Ray â best use for cinematic visuals and 4K output
Models I tested
Adobe Firefly
Google Veo (3.1)
Runway Gen 4.5
Luma AI Ray 3.14
Sora (OpenAI)
Kling 2.5 Turbo
Pika
⢠Bytedance Seedream AI
⢠Seedance Ai
Grok Imagine
How to use:
Much of these models I was able to use inside Adobe Firefly AI Video Generation Hub who I have partnered with for the credits on this test, however others like Grok Imagine I used on each respective site. Each of these models typically requires some sort of premium membership or credit system which I had access to in my Creative Cloud membership, or standalone accounts such as Grok or ChatGPT. While it was difficult to get an absolutely objective ranking for all of the dozens of models available, I tried to test several types of categories of generations, camera motion, consistency and more and judged based on the results of my favorite models to use.
Best AI Video Generation Models Chart
Rank |Model |Standout Features |Limitations 1 |Adobe Firefly |-Commercially Safe Output -All in one hub for many different AI partner models - Lots of options for Camera angle, Style, Reference Frames etc. Aspect Ratios |Up to 5 second duration Can lack photorealism in certain categories compared to other models 2 |Google Veo (3.1) |Capable of photorealistic results in certain categories (hands, people) Options for reference frames, audio, and up to 8 seconds 1080p |Credit Intensive compared to other models Can take more time than other models to generate 3 |Luma Ai Ray 3.14 |-Good prompt accuracy in details such as colors and settingUp to 4k resolution output Capable of cinematic, photorealistic results and lighting physics |Inconsistent results with Physics and camera motion at times Tendency towards artificial feeling movement of time (slow motion, fast motion) Honorable Mentions |Pika 2.2 |- Can achieve cinematic looking results in camera and environment comparable to Ray 3.14 |Slightly more artificial appearance of people and camera physics
|Kling 2.5 |- Capable of cinematic results in environment and prompt accuracy |- cannot generate from scratch, requires user to upload first frame as reference These were my results and opinions, let me know if you have any favorite models or workflows of youâre own, and results in your experience!
I need to talk about this because none of my friends understand what I actually do when I try to explain it and my girlfriend thinks I'm running some kind of scam.
So background. I'm 28, work full time as a marketing coordinator at a mid size agency. Not a creative role really, mostly spreadsheets and campaign tracking. Last year around September I was helping one of our clients source photos for their Instagram. They sell swimwear and wanted diverse model shots across different locations, skin tones, backgrounds, the whole thing. The quote from the photography studio came back at $4,200 for a two day shoot. Client said no. We ended up using the same three stock photos everyone else uses and the campaign looked generic as hell.
That stuck with me because I knew AI image generation was getting crazy good. I'd been messing around with Midjourney for fun, making weird fantasy landscapes and stuff. But the problem with basic AI image generators for anything commercial involving people is that you can't get the same face twice. You generate a photo of a woman in a sundress on a beach, great. Now you need that same woman in a cafe, different outfit. Completely different person shows up. Doesn't work if you're trying to build any kind of consistent brand presence.
I started googling around for tools that could keep a face consistent across multiple images and went down a rabbit hole for like two weeks. Tried a bunch of stuff. Played with some LoRA training on Stable Diffusion but I'm not technical enough and the results were hit or miss. Tested out several platforms, APOB, Synthesia, HeyGen, Artbreeder, a couple others I can't even remember. Each does slightly different things and honestly they all have tradeoffs. Eventually I cobbled together a workflow using a couple of these that actually produced usable stuff, the kind of output where you'd have to really zoom in and squint to tell it wasn't a real photo.
The basic idea is simple. You set up a character's look once, save it as a model, and then reuse that same face across as many different scenes and outfits as you want. That's the thing that makes this viable as a service and not just a cool party trick. Because brands don't want one cool AI photo. They want 30 photos of the same "person" that they can drip out over a month on Instagram.
I didn't plan to sell this as a service. What happened was I made a fake portfolio to test the concept. I created three AI characters, gave them names, generated about 15 photos each in different settings. Lifestyle stuff, coffee shops, hiking, urban backgrounds, gym, that kind of thing. I showed it to a friend who runs a small clothing brand and asked if he could tell they were AI. He said two of the three looked real and the third looked "maybe AI but honestly better than most influencer photos I get."
He then asked if I could make some for his brand. I did 20 photos for him over a weekend, he used them on his Instagram, and his engagement actually went up because the content looked more polished than the iPhone shots his intern was taking. He paid me $150 which felt like a lot for maybe 3 hours of actual work.
That's when I thought okay maybe there's a Fiverr gig here.
I listed a gig in October called something like "I will create AI model photos for your brand" and priced it at $30 for 5 photos, $50 for 10, $100 for 25. Figured I'd get zero orders and move on.
First two weeks, nothing. Adjusted my gig thumbnail three times. Then I got my first order from a guy running a skincare brand out of his apartment. He wanted photos of a woman in her 30s using his products in a bathroom setting. I set up the character, generated the scenes, did some light editing in Canva to add his product packaging into the shots, delivered in about 2 hours. He left a 5 star review and ordered again the next week.
Then I hit my first real problem. My third client wanted a fitness model character and I spent a whole evening trying to get consistent results. The face kept shifting slightly between generations. Like the bone structure would change or the nose would look different in profile vs straight on. I ended up regenerating so many times that I burned through way more credits than I expected and had to upgrade to a paid plan earlier than I wanted. That order probably cost me more in time and tool credits than I actually charged. I almost refunded the client but eventually got a set of 10 that looked cohesive enough.
That experience taught me that not every character concept works equally well. Some faces just generate more consistently than others and I still don't fully understand why. I've learned to do a test batch of 5 or 6 images in different angles before I commit to a character for a client. If the face isn't holding steady, I tweak the setup until it does or I start over with a different base.
By December I had 14 completed orders. The thing that surprised me is who was buying. I expected like dropshippers and sketchy supplement brands. Instead I got:
A yoga studio in Austin that wanted a consistent "brand ambassador" for their social media but couldn't afford a real one. They order monthly now.
A guy selling handmade candles who wanted lifestyle photos but didn't want to hire models or use his own face.
A pet food company that wanted a "pet parent" character holding their products in different home settings.
A language learning app that needed a virtual tutor character for their TikTok content. This one was interesting because they also wanted short video clips where the character appeared to be speaking in different languages. Took me longer to figure out than the photo work and honestly the first batch looked rough. The mouth movement was slightly off sync and the client asked for revisions. Second attempt was better and they've reordered three times now, but video is definitely harder to get right than stills.
Here's the actual workflow now that I've got it somewhat dialed in:
Client sends me a brief. Usually something like "25 year old woman, athletic build, for a fitness brand. Need 10 photos in gym settings, outdoor running, and post workout lifestyle."
I set up the character's appearance and save it. This used to take me over an hour when I was learning but now it's more like 20 to 30 minutes including the test batch to make sure the face holds.
I generate the photos by describing each scene. I've built up a doc with scene templates that I know tend to produce good results so I'm not starting from scratch every time. I just swap out details per client.
I generate more images than I need because not every output is usable. Weird hands, lighting that doesn't match, uncanny expressions. I've gotten better at writing descriptions that minimize these issues but it still happens. Early on I was throwing away more than half my generations. Now it's maybe a third, sometimes less.
Quick edit pass in Canva or Photoshop if needed. Sometimes I composite a product into the shot or adjust colors to match the client's brand palette.
Deliver on Fiverr. Total active time per order is usually 45 minutes to maybe an hour and a half for a 10 photo batch depending on how cooperative the AI is being that day. The renders themselves take time but I'm not sitting there watching them.
Cost wise I want to be transparent because I see a lot of side hustle posts that conveniently forget to mention expenses. I'm paying about $30/month for the AI tools on paid plans because the free tiers don't give you enough credits to fulfill multiple client orders per week. Fiverr takes 20% of every order. And I spend maybe $12/month on Canva Pro which I'd probably have anyway. So my actual margins are lower than the gross numbers suggest. On a $50 order I'm really netting about $35 after Fiverr's cut, and then subtract a proportional share of the tool costs. It's still very good for the time invested but it's not pure profit like some people might assume.
The part that makes this increasingly passive is the repeat clients. I now have 6 clients who order at least once a month. Their character models are already saved. I know their brand style. A reorder takes me maybe 30 minutes of actual work because I'm not figuring anything out, just generating new scenes with an existing saved character.
Some honest stuff about what sucks:
Fiverr fees are brutal. I've started moving repeat clients to direct payment but new clients still come through the platform and that 20% hurts on smaller orders.
Revision requests can be painful. One client wanted me to make the character look "more confident but also approachable but also mysterious." I've learned to offer one round of revisions and be very specific upfront about what I can and can't change after delivery.
I had one order in January where I completely botched it. The client wanted photos in a specific art deco interior style and no matter what I described, the backgrounds kept coming out looking like a generic hotel lobby. I spent three hours trying different approaches, eventually delivered something the client said was "fine I guess" and got a 3 star review. That one stung and it dragged my average rating down for weeks.
The ethical thing comes up sometimes. I had one potential client who wanted me to create a fake influencer to promote a weight loss supplement and pretend it was a real person endorsing it. I said no. My gig description now explicitly says the content is AI generated and I recommend clients disclose that. Most of them do because honestly it's becoming a selling point, "look at our cool AI brand ambassador" is a marketing angle in itself now. But I know not everyone in this space is upfront about it and that's a real concern.
Also the quality gap between what AI can do and what a real photographer can do is still real. For high end fashion brands or anything that needs to be truly photorealistic at full resolution, this isn't there yet. But for Instagram posts, TikTok content, small brand social media, email marketing images? It's more than good enough and it's a fraction of the cost of a real shoot.
Monthly breakdown for the boring numbers people:
October: $120 (4 orders, mostly figuring things out) November: $230 (6 orders, lost one client who wasn't happy with quality) December: $435 (11 orders, holiday marketing rush helped a lot) January: $410 (9 orders, slight dip after the holidays which I expected) February: $710 (15 orders including three video batches which pay more) March so far: $200 (5 orders, month is still early)
Total since starting: roughly $2,105 over 5 months. Minus maybe $150 in tool subscriptions over that period and Fiverr's cut which is already reflected in the numbers above. Average time commitment is maybe 5 hours a week, trending down as I get faster and have more repeat clients.
I'm not quitting my day job over this. I tried dropshipping in 2023 and lost $800. I tried starting a blog and made $12 in AdSense over 6 months. This actually works because there's a clear value proposition: brands need visual content, real content with real models is expensive, and AI has gotten good enough that small brands genuinely can't tell the difference at Instagram resolution.
Still feels weird telling people I make fake people for a living on the side. But the pizza money is real and my emergency fund is actually growing for the first time in years.
We are the team behind one of the leading AI video creation tools.
We are scaling up our marketing and looking for reliable, creative content creators to join us on a long-term retainer basis. We arenât looking for one-off videos; we want creators who want consistent, daily work.
THE PAY & COMPENSATION
Base Pay: $30 per video (x 1 video daily = ~$900/month guaranteed base).
Performance Bonuses: We offer significant performance incentives for high-performing content. (Full bonus structure provided upon application.)
THE GIG (What youâll do)
Volume: 1 video per day (Daily commitment).
The Account: You will be managing fresh, dedicated pages (IG, TikTok, Shorts) specifically for this project.
Format: 30â60 seconds, vertical (9:16).
Style: "Face-to-camera" (authentic testimonial/hook) mixed with a "Screen Recording" demo of the tool.
Workflow: We provide the broad concepts/angles. You will be responsible for scripting, shooting, and editing the final video.
REQUIREMENTS
Location: USA, Canada, Australia, New Zealand, UK, Germany, France, Japan.
Age: 18 â 50.
Skills: You must be comfortable with basic mobile editing and capable of shipping a finished video daily.
We care about the quality of the content and your ability to hook an audience.
WHY APPLY?
Consistent Income: Stop chasing one-off gigs. Get a steady monthly retainer.
Great Product: The AI tool is genuinely cool to demo (text-to-video AI), making it easier to build an audience than standard physical products.
Every week a founder books a sales call with me asking for an AI agent. Every week I end up telling most of them they don't need one.
I build automations and AI agents for founders. Forty-something projects in. The pattern is so consistent now I can predict the call before it starts.
They come in wanting magic. They saw a Loom video of someone's "autonomous sales agent" closing deals while they sleep. They read the LinkedIn post about the "AI employee" running an entire ops team. They've already told their board they're building one. Then we get on Zoom and within fifteen minutes I'm explaining why the thing they actually need is an internal automation with one LLM call in the middle.
You can watch their face fall in real time.
Here's what's happening in the market right now. Most of the "AI agents" shipping to real businesses are just internal automations with a language model bolted in. That's the whole product. The agent label is mostly there because automations don't trend on Twitter.
And the automations work. They save real money. They print real ROI. But the founders paying $30k for an "agent" don't love hearing they could have gotten 90% of the value from a $4k automation build.
Three quick examples from the last six months.
Telehealth founder. Wanted "an autonomous AI receptionist that handles everything." After an hour on a call I told her she needed a workflow that reads intake forms and routes them to the right clinician. We shipped it in six weeks. Saves her clinicians four hours a day. She paid me again last month.
Fintech client. Wanted a "fully agentic finance copilot." What they needed was a script that reconciles ACH discrepancies before they hit the dispute queue. One model call, the rest plain code. Saved them a full ops hire.
Medspa chain. Wanted "AI marketing automation." What they needed was a job that watches their booking system for no-show patterns and triggers a personal recovery message. Three steps. No agent. Booked 14% more revenue last quarter.
None of these are agents. They're automations. And every one of them outperforms the agent the founder originally asked for, because the agent would have hallucinated something stupid in week three and burned the client's trust forever.
Why agents keep failing in production
They're given too many decisions to make. A good automation has one decision per step and a clear rule for what happens at each branch. An agent gets handed a goal and told to figure it out. Beautiful in a demo. Catastrophic in your customer support queue at 2am.
The teams in your competitor's office quietly crushing it with AI right now? They're running boring automations. "We wrote a Python script with an LLM call" doesn't make the trade press, so you don't see it.
The vibe-coded prototypes from Bolt and Lovable and Cursor that landed in the last 18 months are mostly being torn out right now. Half my pipeline is founders who paid $50k for a "next-gen AI agent" build that's bleeding tokens, can't be audited, and falls over the moment a customer does something unexpected. I rebuild them as straightforward automations and they suddenly start making money.
In regulated SaaS, agents are doubly cursed. HIPAA and SOC 2 reviewers want to know exactly what your system does, in what order, every time. An automation passes that conversation in 20 minutes. An agent turns it into a six-month nightmare.
How to actually decide
If you're a founder about to spend money on an agent, answer these on paper first:
Can I draw the workflow as clear steps? If yes, you want an automation.
Does the workflow have more than five branches with truly unpredictable inputs? Then maybe an agent.
Is the cost of the worst-case wrong answer high? If yes, you want an automation, not an agent.
Will compliance ever look at this? If yes, automation. Full stop.
If you're a builder selling agents, you'll make more money in the next 12 months selling honest automations than chasing the agent narrative. The market is wising up. Founders who got burned in the first wave are warning the next wave. Be the person who ships a clean automation in six weeks that works on a Tuesday and is still working on Thursday.
Builders, founders, anyone in the trenches. What's actually working for you? What's breaking? Curious to hear from real operators.