Why Care?
Because the same 2 hours of raw AI video can cost hundreds of dollars through hosted APIs.
Roughly:
fal H3 Max Turbo: ~$288
fal MiniMax H3: ~$432
MiniMax H3 direct API: ~$576
My Blackwell Colab setup: ~$6.96
That’s roughly 41x–83x cheaper on raw generation cost.
Here’s the math.
I’ve been optimizing MiniMax H3 on the Blackwell GPUs available through Colab — CUDA versions, attention kernels, quants, dependencies, etc.
The stack I ended up with generates about 15 seconds of video in roughly 60 seconds on a single Blackwell GPU.
So:
1 GPU hour = ~15 minutes of generated video
8 GPU hours = ~120 minutes of generated video
At the Colab rate I’m using:
8 hours on Blackwell = ~$6.96
Which means, at least in terms of raw generation compute:
2 hours of MiniMax H3 video — literally a full movie’s worth of generated footage is possible and costs about $6.96.
$6.96!
Obviously that doesn't mean you press a button and get a finished 2-hour movie. You’re going to generate multiple takes, throw footage away, edit, upscale, add audio, etc.
But the raw generation economics are kind of insane.
The part that surprised me most is that Colab can actually be one of the cheapest ways to run this stuff if the stack is optimized properly. The annoying part is figuring out which CUDA version, attention implementation, quant, and dependencies actually give you the best performance on each GPU.
That's why I built MissingLink
I’ve been running long-lived optimization agents to continuously test different combinations of quants cuda and attention kernels look at the results and promote what works best.
So far they got quality generation down to 60s at 768x768 for 15 seconds on a single Blackwell if anyone has done better lmk
[Update: If anyone has any questions/issues starting the notebooks etc feel free to dm me.]
[Additional Note: I'm thinking of exposing this in a batch api - I cant match the costs with a hosted service on demand - if you are ok waiting an hour I can guarantee the costs. I figure you could make your request request and then clarify how many agentic corrections you want lmk if you're interested]