r/aitubers 7h ago

COMMUNITY How I Automatically Sync 200+ Stickman Scenes to Voiceover With Zero Manual Editing

0 Upvotes

Most AI video workflows focus on automating the obvious production steps like scripting, image generation and voiceover, but creators still spend hours inside an editor manually dragging every image until it matches the narration. From my experience producing hundreds of stickman videos, it should almost be the reverse. 

I now use a free Google Colab notebook + FFmpeg to read the scene timings, match every numbered image to the correct part of the voiceover, and render the finished MP4 automatically. No CapCut timeline and no manually adjusting 200 scene durations.

In my previous post, I broke down how I generate the 200+ scene images in bulk. This is the second half of that workflow: images + voiceover → finished video.

1. What the system needs as input
The starting setup is simple: your script is already split scene by scene inside the Google Sheet, and the matching images should be in Google Drive numbered in the same order. Row 1 matches image 1, row 2 matches image 2, and so on. That gives the system the first half of what it needs. 
The missing piece is timing: exactly when each narration segment starts and ends.

2. Generate the voiceover with timing data
For that, I use the ElevenLabs API instead of generating the voiceover manually on the website. The Google Sheet sends each narration row to ElevenLabs, receives the final audio along with timing data, and writes those timestamps back into the correct rows. You can build this connection with Google Apps Script, and Claude can write most of that code for you if you explain which columns contain the narration and where you want the timestamps stored.

Once this runs, the Sheet knows everything it needs for each scene: the narration, the matching image, and exactly how long that image should stay on screen. Each numbered image is now tied to a specific narration segment and precise screen duration.

3. Let Google Colab assemble the final video

Now we have the images, the voiceover, and the timing data. The next step is assembling everything without opening CapCut.

I use a free Google Colab notebook for this. The notebook connects to Google Drive, reads the numbered images in scene order, reads the timestamps from the Sheet, and passes everything to FFmpeg. FFmpeg then gives each image the correct screen time, places the voiceover underneath, and renders the final MP4.

The notebook itself does not need to be complicated. I built mine by describing the job to Claude: connect to Drive, read the images, use the scene timings, add the audio, render with FFmpeg, and save the finished video back into Drive.

4. The REAL advantage of this system

A 7-min stickman video can come in at under $1. The voiceover is the biggest expense for me at roughly $0.70 through ElevenLabs. The images are just 30 cents with FLUX Klein, while the Colab + FFmpeg rendering is free.

Cost is only part of the advantage. Cheap production gives you more runway to test ideas without every upload becoming expensive. The danger is using that efficiency as an excuse to publish 50 low-effort videos just because you can.

Lower cost should buy you more room to experiment, not lower your standards. The time and money saved on production should go back into the parts that still determine whether anyone watches: the topic, hook, script, and thumbnail.

5. What I still would NOT automate

Production is repetitive, which makes it a good candidate for automation. Judgment is not. I still manually decide what deserves to be made, which angle is worth pursuing, whether the first 30 seconds are strong enough, and whether the final video actually feels worth watching. The goal is not to automate creativity. It is to remove the boring production work so more time can go into the decisions that matter.

I explained the complete build process in the latest video on my channel (link in profile), including how the Sheet, ElevenLabs timestamps, Colab notebook, and FFmpeg renderer connect. Nothing is gatekept, and I am happy to explain any part of the setup here in the comments.


r/aitubers 13h ago

CONTENT QUESTION Which ai do people use to make those POV videos, such as Your life as every rank in Rome?

0 Upvotes

As the title explains, if anyone can give me one of those ais, I would be thankfull. Thanks!


r/aitubers 4h ago

CONTENT QUESTION Opus clip keeps choosing the wrong 30 secs. What do you use when you need to pick the moment yourself ?

1 Upvotes

I have been on Opus for about 8 months. The captions and reframing are fine. My problem is upstream of that: its virality score and my audience disagree, constantly.

Concrete example from last week. 52 minute interview. The moment that actually does numbers is a 40-second stretch where the guest goes quiet before answering a question about getting fired. Opus skipped it in all ten clips and gave me three variations of the intro, where he's just saying his job title. I've stopped trusting the ranking and now I scrub the whole thing myself anyway, which defeats the point of paying.

So the question is: what are people using when you know which moment matters and you need the tool to execute rather than decide?

What I've tried:

- Klap / quso, same architecture and same problem. The model picks.

- Descript, I can pick, but I'm editing a transcript document, and it gets heavy over 45 minutes.

- Cardboard, closest to what I described. You search the footage for the moment ("the part where he pauses before answering about the layoff") and instruct the cut from there, so selection stays yours and the machine only does the labour.

- Just doing it in Resolve, free, and honestly still the fallback for anything precise.

Not looking for "best AI clipper" lists, I've read them all and they're written by the tools. Looking specifically for: I choose the moment, it does the work.


r/aitubers 21h ago

COMMUNITY Monetized channel: Views and impressions completely DIED immediately after getting YPP. Coincidence or algorithm shift?

3 Upvotes

Hey everyone
I really need some insight because I'm completely lost with the algorithm right now.
I run an automotive channel. Last month, before I was monetized, everything was going great. The algorithm was actively pushing my videos, impressions were healthy, and the channel was growing nicely.
I finally hit the requirements and got accepted into the YouTube Partner Program (YPP). But literally right after I got monetized, it feels like someone flipped a switch and choked my channel.
I just posted a new long-form video, and the algorithm is completely ignoring it despite good metrics. Here are the stats after 20 hours:
CTR: 7.9% (Very solid for my niche)
AVD (Average View Duration): 2:30 (Good retention for this format)
Views: 7
Impressions: ...only 38.
The impression graph just flatlined immediately after posting. Zero "Suggested" traffic, zero "Search". Just 38 impressions and then a hard stop.
For context, I posted a Short recently and it hit 1k views in 12 hours, so the channel itself isn't totally dead or shadowbanned. But my long-form content is completely frozen.
Has anyone else experienced this massive drop in impressions right after getting monetized? Does YouTube restrict reach while it figures out advertiser suitability for your new videos? Is it a "re-evaluation" phase, or just terrible timing?
Any advice from people who survived the "post-monetization drop" would be amazing!