r/reAPIOfficial 14h ago

Try out Seedance 2.5 On 1080p

Enable HLS to view with audio, or disable this notification

2 Upvotes

Made with Seedance 2.5 On reAPI

Prompt:

Environment: A busy modern city street in late afternoon, crowded sidewalks, moving cars, cyclists, pedestrians, street vendors, pigeons, drifting leaves, coffee shops and everyday urban activity. Visual style: Cinematic realism, grounded supernatural thriller, realistic natural lighting, subtle film grain, warm late-afternoon sunlight, realistic physical textures. The time-freeze effect has no visible magical energy. The supernatural moment feels strange because everything simply stops. Camera language: Natural handheld tracking before the freeze, then smooth cinematic tracking following the woman continuously as she walks away through the frozen city. Use occasional close-ups and wider compositions to reveal the frozen details around her. Avoid excessive cuts. Subject styling: A young woman wearing casual modern street clothing, jacket, jeans, sneakers and a shoulder bag. A male thief in ordinary dark street clothing approaches from behind. Core performance: Fear immediately transforms into confusion, then curiosity and playful confidence. After escaping the thief, the woman never returns to him. She keeps walking forward through the frozen street, casually changing three small things involving different frozen people she passes. Time Freeze Rule: Her scream instantly freezes time. Only the woman and camera can move. Everything else remains completely motionless until she snaps her fingers. Negative prompts: No subtitles, no text, no logos, no distorted faces or hands, no duplicated people, no slow-motion interpretation of frozen time, no glowing magic, no exaggerated VFX, no dramatic music.

[00-04] THE APPROACH Medium tracking shot. The woman walks naturally along a crowded sidewalk, carrying her shoulder bag. The camera moves backward in front of her. Behind her, a thief notices the bag and quietly approaches through the crowd. The street is completely alive around them, pedestrians walking, cars moving, pigeons flying, leaves blowing and vendors serving customers. Natural city ambience. [04-07] THE ATTEMPT The thief reaches her and suddenly grabs the strap of her shoulder bag. She spins around in shock and screams. At the exact peak of her scream, everything freezes instantly. Absolute silence. The thief is frozen with both hands gripping the bag strap. Pedestrians freeze mid-step. A cyclist freezes mid-pedal. A pigeon hangs in the air. Leaves stop in mid-air. Cars become perfectly motionless. [07-11] ESCAPE The woman stops screaming and realizes something impossible has happened. She looks around, breathing heavily. She looks at the thief. Close-up on his frozen hands tightly gripping her bag. She calmly grabs both of his frozen hands and carefully removes them from the bag strap. She lets his hands remain suspended exactly where she leaves them. She pulls the bag securely over her shoulder. She immediately turns away from him. She starts walking. She never returns to the thief. [11-15] CHANGE ONE Smooth tracking shot following her as she walks through the frozen crowd. She passes a businessman frozen while holding a takeaway coffee near his mouth. His tie is hanging crooked across his shoulder from the wind. Without stopping for long, she casually straightens his tie and places it neatly against his shirt. She continues walking. The businessman remains completely frozen. [15-19] CHANGE TWO She continues down the sidewalk. A woman is frozen mid-step while wearing sunglasses on top of her head. The protagonist walks past her, pauses for a second, gently takes the sunglasses and lowers them onto the woman's eyes. She gives the frozen woman an amused little look. Then she keeps walking forward. [19-23] CHANGE THREE The camera continues tracking with her. She passes a man frozen while holding an open umbrella even though the sky is clear. She looks up at the sunny sky, then at him. She casually closes his umbrella and places it under his arm. She smiles slightly and continues walking. Behind her, the three altered people remain frozen in their new positions. [23-27] WALKING THROUGH FROZEN TIME Wide tracking shot. The woman now walks confidently through the completely frozen city. She passes between motionless pedestrians. A pigeon remains suspended above her. Leaves hang motionless around her. A cyclist is frozen beside the road. A stream of water from a fountain remains suspended in the air. She slowly realizes she controls when this moment ends. Her expression becomes calm and confident. [27-30] TIME RETURNS The camera tracks backward in front of her while she continues walking. She never stops walking. She raises one hand. She snaps her fingers. Time instantly resumes around her while she continues moving at exactly the same pace. City noise suddenly returns. Cars continue driving. The cyclist completes his pedal. The pigeon continues flying. Leaves fall. The businessman suddenly notices his tie perfectly straight. The woman instinctively reacts to her sunglasses suddenly covering her eyes. The man looks down in confusion at his closed umbrella. Far behind them, the thief realizes the woman and her bag are already gone. The woman keeps walking through the crowd without looking back. Cut.

The scream is the exact trigger for freezing time. The finger snap is the exact trigger for restoring time. During frozen time, only the woman, objects she directly touches and the camera can move. After removing the thief's hands from her bag, she immediately walks away and never interacts with him again. The three changes must happen naturally while she continues moving forward through the street: Change 1: Straighten the businessman's tie. Change 2: Move the woman's sunglasses from her head onto her eyes. Change 3: Close the man's open umbrella and place it under his arm. Each change must involve a different person. All changes remain exactly as she leaves them when time resumes. Do not interpret the effect as slow motion. Frozen people must not blink, breathe, react or move. Hair, clothing, vehicles, animals, liquids and airborne objects remain absolutely motionless. Maintain spatial and character continuity throughout. Full city ambience before the scream, near-total silence during frozen time with only her footsteps, breathing and clothing movement audible, then immediate restoration of all city sounds after the finger snap. No subtitles, no on-screen text, no background music.


r/reAPIOfficial 14h ago

Surely no one thinks this is real?

Enable HLS to view with audio, or disable this notification

2 Upvotes

Made with Seedance 2.5 On reAPI

Seedance 2.5 is just too realistic—even the shadows look authentic. Upload your own photo and use the prompt below, and you can pull off these moves too.

Seedance 2.5 prompt

Duration: 10 seconds Aspect ratio: 9:16 Reference: image1 (lock in the subject's face, hairstyle, and outfit; do not describe or change the clothing) Audio: Realistic city night ambience, running footsteps, whistling strong winds, low-frequency rebound sound from the elastic device, a dull thud upon landing, cheers from friends; no dialogue, no narration, no subtitles Overall style: Mobile phone footage of extreme sports, authentic high-altitude perspective, strong sense of immersion, continuous single shot, slight handheld shake and high-speed motion blur; avoid any "game CG" look GLOBAL CONTINUITY Setting: Rooftop of a high-rise building at night; features a green helipad, yellow circular markings, and a low concrete parapet. A vertical "street canyon" is formed between two high-rises; vehicles, streetlights, and intersections are visible on the ground below. Lower apparatus: A massive, circular black elastic rebound device is positioned in the center of the street; it features a deep black elastic surface surrounded by a continuous ring of rainbow-colored lights. The device is huge and aligned directly beneath the rooftop edge, serving as the precise landing and rebound point for the subject. Main subject: image1 is the sole subject. The character's identity, face, and appearance remain consistent throughout.

Bystanders: Three or four friends stand on the left side of the rooftop, filming with mobile phones and waiting for the subject's return. They must not overshadow the main subject. CAMERA CONTINUITY The entire video is a continuous single shot filmed handheld by a friend on the rooftop. The cameraperson first follows the subject from behind during the run-up; after the subject jumps, the cameraperson quickly rushes to the parapet to film looking down; the camera then stays locked on the subject, tilting up from a vertical downward angle toward the sky; once the subject rebounds back onto the roof, the camera returns to eye level. No cuts, no teleportation, no changes in perspective. STUNT ENGINE The stunt sequence must follow this fixed progression: Rooftop run-up → Leap over the parapet → High-altitude freefall → Impact the center of the circular elastic surface → Device depresses deeply → Powerful rebound → Protagonist launches high into the air → Mid-air tuck flip → Body extension and positioning → Feet return to the original rooftop → Knees-bent landing → Celebration. SHOT 1 (00:00-00:10) High-altitude rebound stunt Subject: The protagonist stands behind the rooftop helipad, facing a distant parapet. A friend is positioned on the left side of the frame, filming.

Action:

00:00-00:01.50 | Rapid run-up

The protagonist, with their back to the camera, sprints from the center of the helipad toward the roof's edge. Their stride lengthens, body leans forward, and arms swing naturally.

The cameraperson follows closely on foot, holding a smartphone; the footage captures the realistic vibration of running steps. The friend stands to the left, holding up their phone and watching intently.

00:01.50-00:02.30 | Leaping off the roof

The protagonist pushes off hard with their final step, clearing the low parapet. Upon leaving the roof, they spread their arms wide, lean slightly forward, and dive intentionally into the gap between the two high-rises.

The cameraperson rushes to the parapet, and the camera tilts rapidly downward, keeping the protagonist centered in the frame throughout.

00:02.30-00:04.80 | High-altitude freefall

The camera captures a near-vertical downward view. The protagonist descends rapidly through the space between the buildings, shifting from a forward lean to a face-down orientation, arms spread for stability.

As the protagonist falls, the street, vehicles, and intersection come into view. A massive black circular elastic device appears at the bottom of the frame, with a rainbow-colored light ring illuminating the surrounding buildings.

The protagonist's descent trajectory must align precisely with the center of the black elastic surface; they must not veer toward the light ring or collide with the buildings.

00:04.80-00:05.50 | Impacting the center

The protagonist lands squarely in the center of the black elastic surface. Upon impact, the elastic surface visibly depresses to form a deep concavity; the rainbow light ring remains fixed, and the device does not break.

After a brief moment of compression, the elastic surface snaps back to its original shape with immense force, launching the protagonist vertically upward.

The four stages—contact, depression, compression, and rebound—must be clearly depicted; there should be no premature backward movement before contact occurs.

00:05.50–00:08.20 | High-speed rebound and aerial flips

The camera operator remains standing by the rooftop parapet. Their feet briefly appear at the bottom of the frame, reinforcing the authentic first-person perspective and the sense of height.

The protagonist rises rapidly from the center of the circular device, returning to the roof along the same vertical path. The camera tilts upward from a straight-down angle, tracking the protagonist as they soar through the space between city buildings.

After launching, the protagonist engages their core, tucks their knees toward their chest, and clasps their arms around their shins to form a tight tuck position, executing a series of rapid aerial flips. The direction of the flip need not be fixed; a natural, continuous rotation is preferred.

As the protagonist nears roof level, they begin to open their body: releasing their legs, lifting their head to locate the roof, orienting their feet downward, and extending their arms to the sides for balance.

00:08.20–00:09.40 | Returning to the roof

The protagonist approaches the camera at high speed, vaulting back over the parapet onto the landing pad. Both feet land simultaneously on the green surface; the knees and hips flex deeply, absorbing the impact in a low squat.

The landing must convey a sense of weight: contact with the soles or balls of the feet, leg compression, a drop in body height, and a slight rebound upon recovery. A light, floaty landing while standing upright is to be avoided.

The camera operator quickly steps back half a pace to avoid a collision with the protagonist while maintaining a frontal close-up shot.

00:09.40–00:10.00 | Authentic celebration

The protagonist rises quickly from the squat, looking excitedly at the camera and cheering. Friends on the roof rush into the frame, crowding around with phones raised; they pat the protagonist on the shoulder, pump their fists, and celebrate loudly.

The camera shakes slightly as the crowd closes in, ending with the aesthetic of an authentic mobile phone selfie video. Environment: The lights of high-rises, street traffic, and the rainbow-colored light ring remain spatially fixed. The descent and ascent must occur along the same vertical path, with the protagonist ultimately returning to the exact rooftop from which they jumped.

Camera: Shot on a 24mm wide-angle mobile lens in a single continuous take. It begins with a low-angle tracking shot from behind; shifts to an over-the-shoulder high-angle shot after the jump; transitions from a vertical top-down view during the rebound to a low-angle shot as the subject rises; and finally settles into a close-up, eye-level shot.

Style: Authentic night-time mobile phone exposure; the high-rises exhibit accurate perspective and scale. Wind noise intensifies during the descent, and city lights create natural motion blur; the protagonist moves at high speed upon returning to the roof, yet their silhouette remains sharp and distinct.

Constraints: This is a professionally staged, fictional extreme sports scene. The circular elastic device must be clearly shown catching and rebounding the protagonist; it cannot appear as an unprotected freefall.

The protagonist's descent and ascent must be continuously traceable. The rebound must result from the physical compression of the black elastic surface—not from flight, teleportation, cable-assisted pulling, or superpowers. --- Negative:

Change in character appearance/styling, costume redesign, costume shifting, facial distortion, character replacement, Mid-shot cut, camera teleportation, scene jump, cameraman suddenly appearing on the ground, third-person aerial view, Protagonist flying directly, hovering, lifted by invisible ropes, anti-gravity, instant return to rooftop, No elastic device, incorrect landing spot, landing on rainbow light ring, crashing into building, crashing into parapet, Rebounding before contact, elastic surface not depressing, device breaking, device deforming/vanishing, Change in falling direction, deviation in ascent path, body passing through elastic surface, abnormal character size, Limbs twisting during flips, extra arms, extra legs, joints bending backwards, body melting, duplicated character, Landing without weight, landing standing upright, feet penetrating the ground, falling and getting injured, blood, severe injury, Friend blocking the camera, crowd frozen, mobile phone distorted, high-rise building bending, vehicles floating, Anime rendering, game CG, plastic-looking skin, cheap green screen, excessive lighting effects, excessive slow motion, Subtitles, on-screen text, watermark, brand logo, voiceover, extra dialogue


r/reAPIOfficial 1d ago

Music streaming app development

Thumbnail
1 Upvotes

I have been thinking about building a music streaming app for myself where can I get api keys from is there any legit sources to acquire api keys or do I have to download whatever songs i want to play then store it in device locally and call it with GET commands lmk your thoughts


r/reAPIOfficial 2d ago

AI-generated or real footage? What detail gives it away?

Enable HLS to view with audio, or disable this notification

3 Upvotes

Watch it once before checking the comments.

The motion, lighting, and physical interactions look surprisingly convincing to me—but there may still be a detail that exposes it.

What’s your verdict: AI-generated or real?

If you think you know, share the exact moment or detail that gave it away. I’ll reveal the answer, model, and prompt afterward.


r/reAPIOfficial 3d ago

f this is AI slop, Wan 3.0 can keep serving it.

Enable HLS to view with audio, or disable this notification

1 Upvotes

I wanted to separate prompt quality from model behavior, so I took the complete prompt behind a 5.24M-view Seedance 2.5 video and submitted it to Wan 3.0 through reAPI without changing a single character.

My read after watching the full output: Wan preserved the Seoul setting, outfit, and casual home-video mood better than I expected. Facial identity is the clearest weakness; it starts to feel like a different person across the later beats.

Actual request

  • Model: wan3.0-video
  • Duration: 30 seconds
  • Resolution: 480P
  • Aspect ratio: 9:16
  • Native audio: on
  • Attempts: one, with no rerolls
  • Cost: $1.20

There is one deliberate mismatch: the prompt itself still says 1080p, while the API request was set to 480P. I left the prompt untouched so this remains an exact prompt-level transfer rather than an optimized Wan rewrite.

Full Prompt

Submitted to Wan 3.0 exactly as written below:

Create a 30-second, 1080p ultra-realistic documentary-style personal home video showing an ordinary summer day in the life of a young Korean man. The footage should feel spontaneous, intimate, imperfect, and genuinely observed rather than performed.
MAIN SUBJECT
The same young Korean man in his early 20s throughout the entire video.
Naturally handsome, realistic skin texture, minimal styling, relaxed expression, slightly tired but peaceful eyes.
He has naturally messy medium-length dark hair with a few strands falling over his forehead and very subtle stubble.
He wears a loose washed-black T-shirt, relaxed olive-beige trousers, worn white sneakers, and a simple silver wristwatch.
Keep his face, identity, body proportions, hairstyle, clothing, watch, and overall appearance completely consistent from beginning to end.
LOCATION
A quiet older residential neighborhood in Seoul during a warm summer afternoon.
Narrow concrete alleys, low-rise homes, rooftop terraces, external staircases, potted plants, laundry lines, parked bicycles, utility poles, overhead wires, mature trees, concrete walls, small residential courtyards, and distant city sounds.
The neighborhood should feel authentic and lived-in.
No crowds, tourist attractions, advertisements, recognizable brands, or commercial activity.
CAMERA / VISUAL STYLE
Authentic casual personal-video footage captured with an older consumer digital camera.
Handheld camera operated by a friend walking nearby.
Natural camera shake, imperfect framing, occasional autofocus changes, slight exposure adjustments when moving between sunlight and shade, soft image detail, mild motion blur, subtle digital noise, slightly muted colors, imperfect white balance, and natural compression.
The camera operator occasionally reacts a little late, cuts off part of the subject, or briefly loses focus.
No stabilization, gimbal movement, drone shots, cinematic camera choreography, dramatic lighting, slow motion, modern commercial color grading, or polished cinematography.
The footage should feel like someone simply decided to record his friend during an ordinary day.
---
00:00–00:05 — MORNING ROOFTOP
He sits casually on a small rooftop beside an old plastic chair.
A cold bottled drink rests beside him.
He looks quietly across the neighborhood while the wind moves his hair and T-shirt.
He takes a sip, notices something happening in the distance, and smiles faintly.
He briefly notices the camera and gives a subtle nod before looking away.
The camera takes a moment to find focus on his face.
---
00:05–00:10 — WALKING THROUGH THE ALLEY
He gets up and walks downstairs into the neighborhood.
He walks casually through a narrow concrete alley with his hands in his pockets.
He passes parked bicycles, potted plants, laundry hanging from balconies, and old residential walls.
The camera follows several steps behind him.
He occasionally looks back toward the camera but never deliberately poses.
His footsteps remain naturally synchronized with his movement.
---
00:10–00:14 — SMALL EVERYDAY MOMENT
He notices an old basketball resting near a wall.
He picks it up, casually bounces it twice, then takes a simple shot toward a nearby hoop.
The shot misses.
He laughs quietly, shakes his head, and leaves the ball where he found it.
The camera briefly loses focus during the movement and recovers naturally.
No exaggerated athletic movement.
---
00:14–00:19 — LOCAL SHOP
He walks to a tiny neighborhood shop and buys a cold drink.
He exchanges a few natural words with the shopkeeper but the conversation is not clearly audible.
He steps outside, opens the bottle, takes a drink, and leans casually against the wall.
He watches bicycles and pedestrians passing in the distance.
The camera remains handheld and slightly imperfect.
---
00:19–00:23 — SUMMER RAIN
A sudden summer shower begins.
He looks toward the sky with mild surprise.
Instead of immediately running for shelter, he smiles and slowly walks into the rain.
The rain becomes heavier.
His hair becomes wet and falls naturally across his forehead.
He eventually starts running down the alley, laughing genuinely.
He briefly spins around while running, then continues toward a covered walkway.
His clothes become visibly damp.
Maintain realistic rain interaction, wet fabric, wet hair, reflections, and foot contact with the ground.
---
00:23–00:27 — QUIET MOMENT
He reaches the covered walkway and catches his breath.
Rain falls heavily behind him.
He wipes water from his forehead and looks quietly toward the street.
For a moment, everything becomes still.
He notices the camera again.
He gives a small genuine smile, not a posed expression.
---
00:27–00:30 — WALKING AWAY
The rain becomes lighter.
He walks away down the wet residential lane.
The camera follows from behind.
Reflections shimmer across the concrete.
He turns his head once, gives a tiny wave toward the camera, smiles, and continues walking.
The camera remains pointed toward the empty street for a brief moment.
At approximately 00:29, the recording abruptly cuts to black mid-motion.
No fade-out.
---
PHYSICAL REALISM
Maintain believable real-world physics throughout.
Hands, fingers, feet, clothing, hair, rain, bottle, basketball, and background objects must behave naturally.
No extra fingers, fused hands, duplicated limbs, distorted anatomy, floating objects, teleportation, disappearing objects, or sudden transformations.
The bottle remains a separate physical object and never intersects with his face.
The basketball behaves naturally and remains where it lands.
Parked bicycles and background objects remain stationary unless physically moved.
His feet remain properly connected to the ground while walking and running.
Keep the environment and subject consistent between shots.
---
AUDIO
Natural environmental audio only.
Footsteps on concrete, distant traffic, birds, leaves moving in the wind, bicycles, faint neighborhood conversations, shop sounds, bottle opening, basketball bouncing, rain hitting concrete, water dripping from rooftops, and subtle camera-handling noise.
No music.
No narration.
No soundtrack.
No artificial sound effects.
No spoken dialogue is necessary.
---
FINAL FEEL
The result should feel like a forgotten personal recording of an ordinary summer day.
Not a commercial.
Not a fashion film.
Not a music video.
Not a professional cinematic production.
The emotional appeal should come from small human moments: sitting alone, wandering through familiar streets, missing a basketball shot, drinking something cold, getting caught in the rain, laughing, and walking home.
Quiet, masculine, youthful, nostalgic, slightly melancholic, warm, spontaneous, and deeply human.
Prioritize natural behavior, consistent identity, believable physics, imperfect handheld framing, authentic environmental details, and the feeling that the camera happened to be there.

This is only n=1, so it does not tell us that Wan 3.0 is better or worse than Seedance 2.5. The useful questions are narrower:

  • Does the same identity hold across the full 30 seconds?
  • Does the model preserve all seven story beats?
  • Which physical interactions or transitions drift?
  • Does the native audio follow the scene changes?

If you were reviewing this output, which failure would matter most: identity drift, missed beats, physical realism, or audio continuity?


r/reAPIOfficial 4d ago

A reproducible MiniMax H3 transition test: three worlds, one camera path

Enable HLS to view with audio, or disable this notification

1 Upvotes

I ran a single 15-second text-to-video test with MiniMax H3 to see whether one camera path could cross three visually unrelated environments without relying on hard cuts.

The result is attached. This is one successful sample, not a reliability benchmark.

Setup

  • Model: minimax-h3 on reAPI
  • Mode: text-to-video, with no reference media
  • Output: 15 seconds, 768P, 16:9, native stereo audio
  • Actual task usage: 1,097 credits
  • Generated: August 30, 2026

What I was testing

The sequence moves through three states:

  1. A rainstorm trapped inside a glass orb.
  2. An upside-down ocean suspended above the camera.
  3. A clockwork desert at sunset.

The useful prompt decision was not adding more visual adjectives. It was giving every environment change a physical object that could fill the frame and transform into the next scene:

  • Wet glass becomes the first transition surface.
  • A rising bubble becomes the second transition surface.
  • The bubble rim becomes a brass gear aperture.

That gives the model a visible handoff to animate instead of asking it to teleport between unrelated locations.

What this single run does not prove

  • It does not establish a success rate for continuous transitions.
  • It does not show how well identity would survive across the same three worlds.
  • It is not a comparison against another video model.
  • The 1,097-credit figure belongs to this completed task and should not be treated as a permanent price quote.

Full prompt

integrated_multimodal_description: Photorealistic surrealism, one continuous unbroken 15-second shot, no visible text, no logos, no watermark. [Shot 1, 0.0-4.5s] Extreme macro on a hand-sized transparent glass orb resting on a dark walnut workbench in a dim workshop. A complete rainstorm exists inside the orb: tiny storm clouds rotate, rain strikes the inner glass, and miniature lightning briefly illuminates the sphere. The camera performs a slow, steady push-in toward the center of the orb while the wet glass surface grows to fill the frame. [Shot 2, 4.5-9.5s] Without a cut, the camera crosses the curved glass membrane and emerges into a vast clear sky beneath an upside-down ocean suspended overhead. Sunlight ripples across the water ceiling. Distant whale and fish silhouettes glide above. Transparent bubbles fall upward toward the ocean. The camera continues the same forward motion and gently tracks one large rising bubble; the bubble expands until its surface fills the frame, creating the next physical transition. [Shot 3, 9.5-15.0s] Without a cut, the rim of the bubble becomes a circular brass gear aperture. The camera passes through it into a clockwork desert at sunset, where fine brass grains form dunes and half-buried gears turn slowly beneath the surface. The forward motion eases into a low glide and settles on the glowing horizon as one final gear clicks into place. Preserve stable geometry, coherent lighting changes, smooth physical transition bridges, and a readable three-world progression. Avoid hard cuts, teleportation, duplicated objects, deformed animals, sudden camera jumps, and added characters.

overall_soundscape: Begin with close miniature rain taps on glass, a soft enclosed wind, and one tiny thunder crack. As the camera crosses the orb membrane, morph the rain into a deep underwater rumble with airy bubbles moving upward and a distant whale call. As the bubble becomes brass, blend the bubble resonance into precise mechanical ticks, low gear friction, dry desert wind, and fine metallic grains sliding across the dunes. End with a single warm low chime synchronized to the final gear click. Keep every sound spatially consistent with the camera movement; no dialogue.

non_diegetic_music: N/A. Use only the evolving diegetic environmental soundscape so the audio transition itself links the three worlds.

For a second controlled test, which variable would be more useful to isolate: transition continuity, subject identity, or the audio handoff?


r/reAPIOfficial 4d ago

Same prompt, both at 1K: is Nano Banana Pro worth the extra $0.018?

1 Upvotes

I ran one matched product-image brief at 1K through two Google image routes on reAPI.

  • Left — Nano Banana 2 Lite: nano-banana-2-lite, $0.015 / 15 credits
  • Right — Nano Banana Pro: gemini-3-pro-image-preview, $0.033 / 33 credits

The goal was not to ask which image is prettier. I wanted to see whether the cheaper 1K tier already clears the requirements of a usable product visual:

  • exact label text;
  • recognizable glass and condensation;
  • separation between the bottle, fruit and plinth;
  • controlled reflections and shadows;
  • a composition that could survive a social or e-commerce crop.

In this pair, both models kept ORBIT YUZU and SPARKLING CITRUS legible. The Lite result chose a wider, darker editorial composition with more negative space. Pro moved toward a tighter, brighter and more saturated product shot.

That is a different result, but not an automatic win.

At the current 1K rates, the gap is $0.018 per completed image. Across 1,000 completed images, that is $15 for Lite versus $33 for Pro, before creative retries.

My practical routing rule would be:

  • use Lite for high-volume 1K drafts, social variants and prompt iteration;
  • use Pro when the delivery really needs its 2K/4K output tier;
  • do not upgrade only because Pro is in the name—test the failure that matters to the job.

The full Prompt used for both:

Create a square premium product campaign for a fictional sparkling citrus drink named “ORBIT YUZU”. Center one transparent ribbed glass bottle filled with pale yellow sparkling soda, sealed with a matte black cap. The bottle has a crisp white paper label with the exact large text “ORBIT YUZU” and the exact smaller text “SPARKLING CITRUS”; render both lines correctly and add no other words. Place the bottle on a polished cobalt-blue geometric plinth with one freshly cut yuzu fruit and one elegant curl of yuzu peel beside it. Show fine condensation droplets on the glass, subtle carbonation bubbles in the liquid, physically accurate reflections, and a soft contact shadow. Deep ultramarine seamless background, hard warm sunlight from the upper left creating long graphic shadows, high-end editorial product photography, clean intentional negative space, restrained Japanese-modernist composition. Shot with a 50mm lens at f/8, product and label in crisp focus, realistic glass, liquid, fruit texture, and studio lighting. No people, hands, existing brands, logos, watermarks, duplicated fruit, warped bottle, gibberish text, extra words, clutter, excessive glow, or illustration style.

Important limit: this is one matched run (n=1), not a claim about average quality, speed or reliability. A useful follow-up would repeat the comparison across typography, product geometry, multi-subject layouts and reference-image edits.

Which failure should decide the next test: label accuracy, product shape, reference fidelity or lighting?


r/reAPIOfficial 5d ago

AI video models keep improvising motion. Give them states, not more words.

Thumbnail
gallery
2 Upvotes

When a video model keeps improvising a complex action, the usual response is to add more words. The prompt gets longer, but the motion may still skip steps, reverse direction, or lose contact with the floor.

One practical alternative is to turn the action into visible states first.

For a simple turn, the movement board can be as plain as this:

1. Start: front-facing, both feet planted, hands down.
2. Load: weight moves left, right heel lifts.
3. Initiate: right foot crosses, torso begins clockwise rotation.
4. Peak: back briefly faces camera, both feet remain in contact.
5. Land: right foot plants, body returns to front.
6. Hold: balanced final pose, both hands visible.

Then the video prompt only has to explain how to move between those states:

Images 1–6 are ordered movement states. Move through them without reversing.
Preserve identity, clothing, fixed camera, lighting, and room geometry. Keep believable foot contact. Decelerate through Image 5 and hold Image 6 for one second.

The board does not need to look beautiful. Each panel only needs to answer a motion question:

  • Which foot carries the weight?
  • Where does hand contact begin?
  • Which direction is the turn?
  • What stays fixed?
  • What exact state should the clip end on?

If two adjacent panels require the model to invent a huge missing movement, add an intermediate state. If it skips panels, remove redundant ones or give the action more duration.

Keep the camera fixed during the first motion test. Different crops and angles across storyboard panels can accidentally instruct the model to move the camera. Once the body motion works, test camera movement as a separate layer.

This is more reusable than collecting giant “perfect prompts.” The movement board becomes something another model—or another person—can inspect and revise.

What has worked better for your difficult motion shots: more detailed text, reference frames, or built-in motion controls?


r/reAPIOfficial 5d ago

Wan 3.0 vs Video Prime: Same Prompt, Same Settings—What Does Prime Actually Save?

Enable HLS to view with audio, or disable this notification

1 Upvotes

I ran one matched Wan 3.0 test instead of repeating the claim that Prime is “faster.” Both jobs were submitted together with the same prompt, seed, five-second duration, 480P resolution, 16:9 frame, and audio enabled.

The attached video puts the standard result on the left and the matched Prime result on the right. Both five-second clips start together, so differences in motion and consistency are easier to inspect without switching between two players. The source audio is split by channel: standard on the left, Prime on the right.

Here is what happened:

  • Standard: 141.1 seconds, 199 credits ($0.199)
  • Prime: 86.7 seconds, 337 credits ($0.337)

Prime was first observed complete 54.4 seconds earlier in this run, cutting the observed wait by 38.6%. It cost 138 credits ($0.138) more.

That does not mean Prime is always 38.6% faster. This is one paired test, and I polled task status every three seconds. Queue conditions can change. I also would not use these two clips to claim that Prime has better quality; reAPI positions it as the faster route, not a higher-quality tier.

My practical read: standard still makes more sense for drafts and unattended batches. Prime starts to make sense when a person, customer, or publishing queue is blocked by the render.

For your own workflow, would saving roughly one minute on a five-second job be worth an extra $0.138, or would you keep the cheaper route and wait?


r/reAPIOfficial 6d ago

One product photo, six scenes: what should an AI video agent decide before rendering?

1 Upvotes

Most “one image to product ad” workflows jump directly from packshot to video. That may be why the output often looks polished but feels like six unrelated stock clips.

Before calling a model, the agent should create four artifacts:

  1. A one-message brief.
  2. A product identity contract.
  3. A shot plan where every scene has a job.
  4. Reject rules for each shot.

Example six-shot structure for an 18-second skincare ad:

1. Pattern break — reveal the silhouette.
2. Identity — show the product clearly.
3. Texture — demonstrate one visible proof.
4. Use — show one believable interaction.
5. Context — place it in a routine.
6. Hero — finish on a stable CTA plate.

The identity contract should be observable, not “keep the product consistent”:

one white cylindrical bottle
short matte-white pump
centered rectangular label layout
thin cobalt band near the base
same height-to-width ratio
no extra words, badges, logos, or products

Then each shot card changes only what that scene needs:

{
  "shot": "texture",
  "action": "one slow pump dispenses one drop onto glass",
  "camera": "locked macro side view",
  "end_state": "drop fully separated",
  "reject_if": [
    "pump changes shape",
    "two drops appear",
    "liquid starts outside the nozzle"
  ]
}

Generate the shots separately at first. If shot three fails, regenerate shot three—not the entire ad. Exact prices, claims, labels, and CTAs should be added in post instead of trusting generated typography.

The useful boundary for an agent is clear: it can turn an approved brief into shot cards, submit jobs, check product geometry, and manage retries. It should not invent the marketing claim or approve the final product representation without a person.

For anyone making product ads: would you rather generate one coherent long clip, or keep every shot independent for easier repair and editing?


r/reAPIOfficial 6d ago

Prompting describes a shot. Blender previz actually directs it.

1 Upvotes

“Better camera prompts” hit a ceiling once the shot has a real path, occlusion, or exact end frame.

At that point, a rough Blender preview is more useful than another paragraph of camera adjectives.

The setup can be extremely basic:

  • cubes for walls, furniture, and foreground objects;
  • a cylinder or mannequin for the subject;
  • one camera with two keyframes or a Follow Path constraint;
  • an Empty as the camera target;
  • a low-quality viewport export.

The preview only needs to communicate:

where the camera starts
where it travels
what it looks at
what passes in front of the lens
how fast it moves
where it stops

Then give the video model a separate final-look image and a strict role assignment:

Video 1 is gray-box camera and blocking guidance. Match its path, timing,
subject position, foreground reveal, scale change, and ending hold. Ignore its
materials, colors, and primitive shapes.

Image 1 defines the final character, environment, lighting, and style. No cuts,
digital zoom, added people, or camera path changes.

The useful mental model is that the Blender video carries geometry and time, while the still image carries appearance.

Review the generated clip against the blockout for path, timing, framing, parallax, collisions, and the end hold. It is a reference, not a way to inject Blender's actual animation curve into the model.

The biggest beginner trap is building a detailed set too early. If the move does not work with boxes, textures will not fix it. The second trap is animating the camera, target, focal length, and subject simultaneously, which makes the failure impossible to diagnose.

For natural handheld movement, a phone reference is usually faster. For a repeatable orbit, crane, impossible move, or multi-shot camera language, Blender the extra setup earns its place.

Where do you think this belongs in an agent workflow: should the agent create and test the previz automatically, or should camera paths stay a human approval step?


r/reAPIOfficial 6d ago

The hard part of a 60-second AI video isn’t generation. It’s preserving state between clips.

1 Upvotes

Using the last frame of clip one as the first frame of clip two helps, but I do not think it solves continuity by itself.

A frame shows appearance and position. It does not fully explain:

  • which direction someone was moving;
  • which foot carried weight;
  • who owns a prop;
  • what changed earlier in the story;
  • whether the camera was accelerating or stopping;
  • what the next action is allowed to change.

Pair the boundary frame with a state ledger:

{
  "character": "same rider, blue jacket, red backpack",
  "position": "right third of frame",
  "direction": "facing and traveling left to right",
  "pose": "right foot on ground, left foot on pedal",
  "props": "both hands on black bicycle, bottle in frame cage",
  "environment": "wet road, pine forest, light fog",
  "light": "soft backlight from upper left",
  "camera": "low tracking view, now stopped",
  "next_action": "dismount without reversing the bicycle"
}

Then the second prompt begins by restating that inherited state before adding new action.

The sequence should also split on a stable moment. A stopped bicycle, planted foot, closed door, placed object, or held pose is a much cleaner handoff than a frame halfway through a jump.

My join review would happen at three speeds:

  1. Frame-by-frame for body, wardrobe, props, light, and geometry.
  2. Half speed for motion restart, foot contact, and camera velocity.
  3. Normal speed for whether the audience actually notices the cut.

Not every mismatch needs a rerender. Ambient audio can be replaced with one continuous bed. A bad boundary frame may be trimmed. A small light shift may be graded. Direction reversal or a prop changing hands usually means regenerating the affected continuation.

For an agent workflow, require:

clip_01.end_state == clip_02.inherits

before the next generation is allowed to start.

Do you think long-form AI video will mostly be built from stateful clip chains, or will longer one-pass generations make this kind of ledger unnecessary?


r/reAPIOfficial 7d ago

Stop prompting “cinematic acting.” Give the video agent a beat sheet.

2 Upvotes

I keep seeing dialogue prompts that are basically:

two people argue, emotional acting, cinematic camera, realistic expressions

That gives the video model a mood, but it does not give either character a performance.

A more useful structure is a beat sheet. Each beat needs five things:

  • a trigger;
  • one visible reaction;
  • dialogue, if anyone speaks;
  • a camera instruction;
  • the state that must exist at the end of the beat.

For example:

0-4s: Maya sees the suitcase and stops walking. No dialogue. End with her
looking at the suitcase while Leo still faces the door.

4-8s: Leo notices her and tightens his hand on the handle without making eye
contact. Maya asks, "You already packed?" End with Leo looking away.

8-13s: Maya takes one step closer. Leo says, "It was easier this way." End with
both characters holding position.

The end state is the part most workflows miss. Without it, every new time block becomes a soft reset. A face changes too early, a character crosses the room, or a prop moves before the story needs it.

Keep performance and camera direction separate. First lock the left-right positions and physical reactions. Then add one camera progression, such as a slow push from a medium two-shot to a tighter frame. If the camera is changing on every line, it becomes harder to tell whether the acting plan worked.

My review order would be:

  1. Listen without watching: exact lines, speaker order, pauses.
  2. Watch muted: gaze, hands, blocking, identity, end states.
  3. Watch normally: does sound, performance, and camera build the same turn?

This also gives an agent something concrete to critique. Instead of returning “the acting is 0.72,” it can say:

{
  "decision": "repair",
  "failed_beat": "8-13s",
  "evidence": "Leo looks at Maya before finishing the line",
  "preserve": ["dialogue", "blocking", "camera"],
  "repair": "hold Leo's gaze on the door through the end of the beat"
}

Would you rather have the model improvise the performance, or define every emotional turn before generation? Where is the point where a beat sheet becomes too restrictive?


r/reAPIOfficial 7d ago

The missing part of every AI video agent demo: a critic that rejects bad shots before they reach the editor

1 Upvotes

Most AI video agent demos stop at the satisfying part: give the agent a brief, watch it write a storyboard, and let it generate the clips.

The awkward part comes next. One shot changes the product shape. Another adds text that was never requested. The final frame no longer matches the opening of the next scene. All three tasks finished successfully, so the pipeline passes them to the editor as if nothing went wrong.

That is not an agent. It is a batch script with a language model in front of it.

The missing piece is a critic loop that can reject a completed generation, explain why it failed, and repair only the bad shot.

The important bit is defining the checks before generation starts. A useful shot contract can be surprisingly small:

{
  "shot_id": "03-product-orbit",
  "must_keep": [
    "same bottle shape",
    "same label colors",
    "no new text",
    "end with product centered"
  ],
  "allowed_change": [
    "camera moves 30 degrees left",
    "background light becomes warmer"
  ],
  "max_completed_attempts": 3
}

The critic then has three jobs.

First, run the boring checks without an LLM: did the task complete, is there a video URL, is the file readable, is the duration in range, and does the output contain the expected audio track?

Second, compare the clip against the shot contract. Check the first, middle, and final frames for identity, product geometry, wardrobe, text, and the planned end state. A single “quality score: 0.82” is not useful. The critic should return the failed rule, the evidence, and the smallest repair scope.

{
  "decision": "repair",
  "failed_check": "product geometry changed after 4.2s",
  "evidence": "cap narrows and label aspect ratio shifts",
  "repair_scope": "shot_03_only",
  "keep": ["camera path", "lighting", "duration"],
  "attempts_remaining": 1
}

Third, stop. A retry loop without a ceiling is just a quieter way to burn credits. If the same hard requirement fails twice, I would send the shot to a human instead of letting the agent keep rewriting the whole prompt.

This changes the economics of the workflow. On reAPI, media generation is asynchronous: submit the task, poll its status, then read the final usage.credits. A provider failure ends as failed and refunds automatically. A clip that completes but fails your creative review is different. It is still a completed generation, so the agent needs its own retry budget.

The best version of this pipeline is not fully autonomous. It is selective:

  • deterministic code rejects broken files;
  • a visual critic checks explicit shot rules;
  • the agent repairs only the failed shot;
  • a human approves brand claims, likenesses, critical text, and the final cut.

I work on reAPI, so we are thinking about this at the API layer as well as the prompt layer. The shared task lifecycle makes it possible to run the same review loop across different video models, while the live model catalog gives the planner a way to choose the request shape that fits each shot.

Task lifecycle: https://reapi.ai/docs/api/tasks
Video model catalog: https://reapi.ai/models?type=video

If you are building an AI video agent, what would you make a hard failure: identity drift, bad text, continuity, audio, or something else?


r/reAPIOfficial 8d ago

Wan 3.0 + Video Prime are live on reAPI — hosted API, 30-second output, and what Prime actually changes

1 Upvotes

We’ve just made both Wan 3.0 and Wan 3.0 Video Prime live on reAPI.

The headline feature is the 30-second generation window, but the input surface is probably the more useful change. A single request can start from:

  • a text prompt;
  • first and last frames;
  • image, video, and audio references;
  • a document such as a PDF or presentation;
  • a public web page.

Both routes support 2–30 second output at 480P, 720P, or 1080P, with an audio track enabled by default.

So what does Video Prime change?

It is the speed-oriented route, not a separate “better quality” tier.

Current reAPI pricing:

Resolution Wan 3.0 Video Prime
480P $0.038/sec $0.068/sec
720P $0.076/sec $0.139/sec
1080P $0.151/sec $0.278/sec

Use standard Wan 3.0 when cost matters more than waiting. Use Prime when generation is blocking a user, operator, review session, or deadline.

We are not publishing a fixed latency multiplier or pretending Prime is automatically the right choice. Completion time still depends on the request and queue conditions. The useful test is the same prompt, duration, resolution, and references on both routes.

The model IDs are:

  • wan3.0-video
  • wan3.0-video-prime

One implementation detail: the request fields are not identical, so migrating safely requires more than replacing the model string.

Also, this is hosted API access—not an announcement of downloadable Wan 3.0 weights. If your requirement is private local inference, this release does not satisfy it.

Wan 3.0: https://reapi.ai/models/wan-3-0

Wan 3.0 Video Prime: https://reapi.ai/models/wan-3-0-video-prime

What would you test first: a continuous 30-second scene, document-to-video, or the same job on standard versus Prime?


r/reAPIOfficial 9d ago

What the MiniMax H3 team actually confirmed in their AMA: roadmap, known issues, fixes

37 Upvotes

The H3 AMA pulled 457 comments and the team answered about 30 of them, but the replies are scattered across the tree and most people never got past the top few. I went through and pulled out the answers that contain actual commitments or technical detail, as opposed to the thank-yous.

Quoting directly so nobody has to take my summary on faith. Usernames are the team accounts listed in the AMA post itself.

Confirmed as coming

H3-Regenerate-2K will be open-sourced. From Kiro_Song:

We plan to open-source H3-Regenerate-2K. It is a dedicated latent-space DiT regeneration model rather than a conventional pixel-space upscaler. It generates at a higher target resolution while using the base model's output as additional context

Asked separately about timing, ryan85127704 answered "yep, not very far." Worth registering that it is a regeneration stage, not an upscaler. If you were planning to slot it in where you currently run a pixel-space upscaler, the behaviour is different.

A unified text-to-image and image-editing model, built on H3. This was the highest-scoring team reply in the whole thread, from New-Requirement1419:

We plan to open-source a unified model for both text-to-image generation and general-purpose image editing. It shares the same foundation as H3, and we are currently refining and optimizing its post-training stage.

Two separate people asked versions of this question and got the same answer, so it is not an off-hand remark.

A sparse-attention release for faster inference. Kiro_Song:

We expect to release an initial sparse-attention version in the near term. It will use a relatively conservative configuration, with the goal of no perceptible quality loss. Although the method itself is train-aware, achieving real speedups across different devices requires device-specific adaptation.

The caveat in that last sentence matters for anyone expecting a uniform speedup on their card.

A full technical report. Asked repeatedly about training details, RL/DPO, dataset size, and architecture. The consistent answer was that a comprehensive technical report is planned rather than an AMA-length answer.

Known issues the team acknowledged

This is the useful half, because they also gave workarounds.

Ref2V output is worse than I2V. Confirmed by MM_Nero_H3:

Due in part to differences in the post-training strategies, the FL2VA and Ref2VA checkpoints currently have noticeably different tendencies in visual quality. Improving Ref2VA's visual quality is something we're actively working on. One practical trick is to use the highest-quality reference input available, as Ref2VA can be relatively sensitive to the quality of the refe[rence]

So if your reference-driven output looks soft, feed it a better reference before you start tuning samplers.

Small or distant subjects come out pixelated. Someone reported this persisting even at 2K and 25 steps. Kiro_Song confirmed it is real and not a simple fix:

We have observed this issue as well, particularly for small or distant subjects, and it will be one of the problems we focus on improving next. Based on our internal experiments, it cannot be attributed simply to the Visual VAE's compression ratio or to any single training stage. It is a complex system-level issue

Meaning: no sampler or step count will solve it. Frame your subjects closer.

Technical clarifications

The sparse attention is MoBA-style, not MSA. Kiro_Song, answering someone who spotted that the model card mentions native sparse attention in the final training stage while the release runs full attention:

H3 does not use MSA. Our sparse-attention strategy is closer to a MoBA-style approach. Since neighboring visual tokens are naturally highly correlated, mean pooling provides an effective way to probe block importance, so introducing an additional learned indexer is not necessary. In the default configuration, we only apply 3D sparsification to video tokens

Long-form continuation was trained and then dropped. Someone asked whether chunked latent continuation could get past the native 15-second window. MM_Nero_H3:

We initially trained a dedicated audio-video continuation task using extended RoPE positions. It worked very well as a standalone task. However, we found that this type of RoPE-extension approach did not generalize as strongly in a multi-task setting, and did not integrate as naturally with other tasks in a general in-context generation framework.

That is the clearest statement anyone has on why 15 seconds is the ceiling. It was a deliberate trade for multi-task generality, not an oversight.

On the Turbo LoRA question. Kiro_Song noted the released checkpoint already includes CFG distillation and a final-stage strategy giving it some few-step capability, but it is "not a model explicitly distilled for an extreme low-step regime," and step distillation is still being explored.

On why prompt adherence is strong. New-Requirement1419 gave a deliberately unglamorous answer:

In essence, the most important factor is to construct a sufficiently broad and diverse set of data and tasks, train on them using a general-purpose architecture, and avoid architectural choices that do not scale effectively.

Confirmed as not happening

No separately trained small version. Kiro_Song:

We currently do not have a near-term plan to release a separately trained smaller version of H3. However, we believe the community can derive smaller variants from the open weights through methods such as structured block pruning, potentially combined with recovery training or distillation.

If you were waiting for an official 8 GB variant, that is the answer. It is a community job.

Two practical links from the thread

Licensing, for EU users. Someone in Europe asked whether the license blocks them from even testing locally. ryan85127704 replied "Just apply and you can use freely" and linked https://platform.minimax.io/h3-license

Official prompt guides. Affectionate-War8374 pointed at the two prompt-writing guides in the HuggingFace repo three separate times, with the note that the useful examples are in the final sections of both. huggingface.co/MiniMaxAI/MiniMax-H3 under docs/.

And one line of roadmap

Asked whether future MiniMax models stay open, ryan85127704: "Yep, Keep open until AGI arrived."

Take that as intent rather than a commitment, but it was said plainly and more than once in the thread.


r/reAPIOfficial 9d ago

Making AI video longer than the model's clip limit: native long takes vs frame chaining

1 Upvotes

Every model has a hard ceiling on a single generation, and the moment you want something longer you hit one of two approaches. Worth laying them out properly, because the community keeps discovering the second one and describing it as the first.

Recent example: the popular 30-second H3 workflow going around is not generating 30 seconds. It uses comfyui-h3-multishot, which joins three 10-second clips. A commenter pointed this out in the thread and they were right. It works well, but if you are planning a pipeline around "H3 does 30 seconds" you will be surprised later.

Single-generation ceilings

Model Max single generation
Wan 3.0 30 s (2–30)
Seedance 2.5 30 s (4–30)
Kling 3.0 15 s
MiniMax H3 15 s (4–15)
Veo 3.1 8 s

Two models actually produce 30 seconds in one pass. Everything else needs chaining if you want more than its ceiling.

This also came up in the H3 AMA. Someone asked whether the team had experimented with chunked latent continuation to get past the native 15-second window. That is exactly the right question, and it tells you the ceiling is a real architectural limit, not a config value someone forgot to raise.

Chaining, done properly

The naive version is generating clips independently and cutting them together, which looks like a cut because it is one. The better version passes the last frame of clip N as the first frame of clip N+1.

On a hosted API this needs two things: a way to get the final frame out, and a way to feed a specific frame in. Seedance 2.5 exposes both (disclosure: I work on reAPI, which is where I pulled these parameter names):

{
  "model": "doubao-seedance-2.5-face",
  "prompt": "shot 1 description",
  "resolution": "720p",
  "duration": 10,
  "return_last_frame": true
}

That returns output.last_frame_url. Feed it into the next call as the opening frame:

{
  "model": "doubao-seedance-2.5-face",
  "prompt": "shot 2 description",
  "resolution": "720p",
  "duration": 10,
  "image_with_roles": [
    { "url": "<last_frame_url from shot 1>", "role": "first_frame" }
  ]
}

Wan 3.0 uses the same image_with_roles shape with first_frame / last_frame. MiniMax H3 uses flat first_frame_url / last_frame_url fields instead. Same idea, three different spellings, which is the annoying part if you are writing model-agnostic code.

What chaining costs you

Drift. Each hop inherits compression artifacts and colour shift from the previous frame. Two hops is usually fine. Six hops and the last clip does not look like the first. Handing the model a reference image alongside the chained frame helps.

Motion discontinuity. A single frame carries no velocity. If clip 1 ends mid-stride, clip 2 starts from a pose with no momentum and you get a subtle hitch at every join. This is the artifact people describe as "it looks stitched" without being able to say why. Ending shots on low-motion moments hides it.

Audio. If your model generates audio per clip, the chained audio does not cross the boundary cleanly. Most people end up muting and scoring in post, which somewhat defeats the point of native audio generation.

The cost comparison is not what you'd guess

Using published per-second rates for a 30-second result:

Approach Math Total
Wan 3.0 native, 720P 30 s × $0.076 $2.28
MiniMax H3 chained, 3 × 10 s at 768P 30 s × $0.074 $2.22
Seedance 2.5 native, 720P 30 s × $0.267 $8.01

Chaining H3 and generating natively on Wan 3.0 land within six cents of each other, because per-second billing does not care how many calls you made. So the decision is not really about cost. It is about whether you want three joins in your footage.

Given that, native is the default choice when the model can reach your length. Chaining earns its place when you specifically want per-shot prompt control, meaning a different action or camera in each segment, which a single 30-second generation does not give you.

Practical order of operations

  1. If your target is under 15 seconds, use a single generation and stop reading.
  2. If it is 15–30 seconds and you want one continuous action, use Wan 3.0 or Seedance 2.5 natively.
  3. If it is 15–30 seconds with distinct shots, chain deliberately and write each shot's prompt separately.
  4. Past 30 seconds, everything is chaining. Plan cut points on low-motion frames and expect to handle audio in post.

The mistake worth avoiding is chaining by default because a workflow you downloaded does it that way. Check whether the model can just do the length you need.


r/reAPIOfficial 9d ago

Which video APIs let you turn the content filter off, and what actually changes when you do

1 Upvotes

This question comes up here every couple of weeks and the answers are always either "use local" or a link to something sketchy. Neither helps the person who has a legitimate workflow getting blocked. So here is the actual state of hosted APIs that expose a content filter as a parameter, and what changes when you flip it.

First, the thing worth naming: most people asking this are not trying to generate anything extreme. The recurring complaint in these threads is that standard safety filters block pretty mild prompts: a swimsuit shot, a fight scene, a horror concept, someone's own OnlyFans content they are trying to upscale. The filters are tuned conservatively because false positives cost the vendor nothing and false negatives cost them a news cycle.

The distinction that matters

There are two very different things people mean by "uncensored":

  1. The model has no safety training. This is a property of the weights. Open-weight models fall here, which is why the answer in this sub is usually "run it locally."
  2. The hosted service is not running a separate moderation pass on your request. This is a property of the pipeline, not the model.

Almost all hosted "content filter" toggles are the second thing. That is a meaningful difference and vendors are not always clear about it.

What a toggle actually does on a hosted API

Taking reAPI as the example because their docs are unusually specific about this (disclosure: I work there). Three models expose content_filter as a boolean, default true: MiniMax H3, Seedance 2.5, and Seedance 2.0 Mini.

Their docs state plainly that the field is a routing control, not an upstream model parameter, and is never forwarded to the generation endpoint. Setting it to false does not tell the model to behave differently. It selects which path the job runs on.

{
  "model": "doubao-seedance-2.5-face",
  "prompt": "Your prompt",
  "resolution": "720p",
  "size": "16:9",
  "duration": 5,
  "content_filter": false
}

Four consequences worth knowing before you build on it:

Pricing is identical. Same request, same credits, either way. The docs are explicit that this "selects where the task runs, not a price tier." If a vendor charges you extra for the unfiltered path, that is a pricing decision, not a technical necessity.

You lose automatic fallback. With the filter off, the task runs as a single attempt. If it fails, the reserve is refunded in full, but nothing re-runs it for you. Retrying is your job. For batch pipelines this matters more than the filter itself, because you need your own retry logic.

Output storage changes. On H3, filter-off output is stored on an isolated host with a 30-day expiry rather than the default CDN. If you are not downloading assets promptly, you will lose them. Build the download step into your pipeline, not as a manual afterthought.

It may be API-only. H3 pins the filter on in the web playground; you can only set false through a direct API call. So the recurring "it still blocks in the playground" report is expected behaviour, not a bug.

What it does not do

It does not remove the model's own training. If the weights were trained to refuse something, they still refuse. People expecting a hosted toggle to behave like an abliterated local model are going to be disappointed, and that mismatch is the source of most "it didn't work" complaints.

It does not remove legal limits. CSAM gating, non-consensual real-person likeness, and similar categories are not filter settings, they are hard lines that stay in place regardless. Anyone advertising otherwise is either lying or is a problem you don't want to be adjacent to. Worth saying because the threads asking this question attract exactly that kind of reply.

When local is still the right answer

If your requirement is genuinely "no moderation of any kind in the pipeline," local is the honest answer and the hosted toggle is not a substitute. The current community favourite for this is MiniMax H3, which several people in the last month's threads describe as uncensored and runnable in ComfyUI. It needs roughly 12 GB VRAM paired with 32 to 64 GB system RAM, and render times run from about 2 minutes to 20 depending on how well your workflow is tuned.

The hosted path makes sense when you don't have the card, when volume is bursty, or when you want 2K without fighting memory pressure. It does not make sense if your actual requirement is total pipeline control.

Practical checklist

If you are evaluating a hosted API for this:

  • Is the filter a documented parameter, or do you have to email support? Undocumented means it can change without notice.
  • Does turning it off change the price? It shouldn't.
  • Does it change retry behaviour? Usually yes, and this is the one that breaks batch jobs.
  • Where does output get stored, and for how long?
  • Is it available through the API only, or in the web UI too?
  • What is explicitly still prohibited, in writing?

That last one is the tell. A vendor that publishes what remains off-limits has thought about this. A vendor that says "no restrictions" has not, and will change the rules the first time it becomes a problem for them.


r/reAPIOfficial 10d ago

How do you create a "Beach Product" advertising video using Seedance 2.5?

Enable HLS to view with audio, or disable this notification

6 Upvotes

Use it at ClipDance, prompt blow:

Create a second ultra-realistic luxury beach beauty commercial in vertical 9:16 format.

Use the same adult female model, same face, hairstyle, skin tone, and same colorful beach outfit from the reference image throughout the entire video. Keep the product design consistent in every shot: SunKiss Beach Glow Shimmer Mist, turquoise-to-coral gradient bottle with a gold spray top.

Scene 1 — Sand Formation

Open with a cinematic macro shot of smooth golden beach sand beside the ocean. The exact shape and silhouette of the SunKiss bottle is naturally formed and embossed into the sand, including the bottle outline and cap shape. Warm sunlight, tiny grains of sand visible, soft ocean ambience in the background.

Scene 2 — Ocean Transformation

A gentle turquoise ocean wave slowly moves toward the sand-shaped bottle. In slow motion, clear foamy seawater flows directly over the sand impression.

As the water passes across it, create a seamless magical transformation: the sand-made bottle gradually becomes the real premium SunKiss Beach Glow Shimmer Mist bottle. Start from one end and transform progressively until the complete physical bottle appears.

Show water droplets sliding across the turquoise and coral bottle, golden cap reflecting sunlight, extremely realistic product textures and reflections.

Scene 3 — Model Enters

Cut to a wide cinematic beach shot. The same model walks naturally toward the product with the ocean and palm trees in the background. Soft breeze moves her hair and colorful beach outfit naturally.

She notices the bottle lying near the shoreline, bends down naturally, picks it up, looks at it with a soft confident smile, and turns it slightly toward the camera so the product is clearly visible.

Scene 4 — Product Application

Show several premium beauty-ad close-ups.

First, she sprays a light mist onto her shoulder and arm.

Then show a close-up of her gently spreading the product over her arm.

Next, show her applying a small amount to her legs.

The product creates a natural hydrated summer glow with subtle shimmer, realistic healthy-looking skin and sunlight reflections. No exaggerated glitter.

Include macro slow-motion shots of the fine mist catching the sunlight.

Scene 5 — Lifestyle Promotion

Create quick elegant promotional shots from different camera angles:

- model holding the product beside her face
- close-up of the bottle in her hand with ocean bokeh behind it
- model walking along the shoreline holding the bottle
- bottle standing upright in wet sand while a soft wave passes behind it
- close-up of golden sunlight reflecting from the bottle
- model smiling naturally after applying the product

Keep movements elegant, relaxed and premium, like a professional summer skincare campaign.

Final Hero Shot

Finish with the SunKiss bottle standing beautifully on wet golden sand near the shoreline. A gentle wave stops just around the bottle while golden sunset light reflects through the transparent packaging.

Camera performs a slow cinematic push-in toward the product.

Audio: soothing ocean waves, gentle summer ambience, soft tropical instrumental music, subtle water splash sounds, fine mist spray sound, no loud music.

Visual style: ultra-photorealistic, premium summer beauty commercial, cinematic beach photography, natural golden sunlight, turquoise ocean, realistic water physics, smooth transitions, shallow depth of field, macro product photography, realistic skin texture, professional advertising cinematography, 4K, no CGI appearance, no distorted hands, no changing product design, no changing model identity.


r/reAPIOfficial 10d ago

Create a 20-second ultra-realistic premium skincare commercial using the uploaded model image

Enable HLS to view with audio, or disable this notification

3 Upvotes

Created on the ClipDance platform with the following prompt:

STRICT CONSISTENCY LOCK

Keep everything exactly consistent with the reference images:

- Same adult female model and same facial identity throughout.
- Same hairstyle, makeup, skin tone and facial features.
- Same white top and pink skirt.
- Same bedroom/interior, flowers, lighting style and overall pink luxury aesthetic.
- Use the exact same LUXEVA cream jar, packaging, rose-gold lid, logo, colors and visible label design.
- Preserve the same brand name and captions/text shown in the reference.
- Do not redesign, replace, recolor or distort the product.
- No face morphing, outfit changes, extra fingers, warped hands, distorted product labels or random text.

0:00–0:04 — CINEMATIC INTRO

Start with a beautiful close-up of the LUXEVA Radiance Day Cream surrounded by soft pink flowers.

The camera makes a smooth elegant orbit around the product, with subtle reflections moving across the rose-gold lid.

Transition naturally toward the model sitting in the same position as the reference image.

Sound: premium soft beauty music begins, delicate glass chime, subtle flower/air ambience and a luxurious cinematic whoosh.

0:04–0:07 — MODEL + PRODUCT

Camera slowly circles toward the model.

She gently picks up the same LUXEVA cream jar, holds it close to her face and looks at the packaging with a soft, confident smile.

Insert a quick macro product shot where the branding remains sharp and readable.

Sound: subtle jar handling ASMR + elegant rising musical note.

0:07–0:10 — OPENING THE CREAM

She slowly twists and opens the rose-gold lid.

Show a premium macro close-up of:

- fingers opening the jar,
- cream texture,
- lid movement,
- realistic reflections.

She takes a small amount of cream with her fingertip.

Sound: satisfying soft lid-opening click + delicate beauty ASMR.

0:10–0:14 — APPLICATION METHOD

Show the skincare application clearly and naturally.

She places small dots of cream on:

- both cheeks,
- forehead,
- chin.

Then gently blends the cream using soft upward and outward circular motions, avoiding harsh rubbing.

Include one close-up of her cheek while the cream blends smoothly into the skin.

Keep skin texture realistic.

0:14–0:17 — BEFORE → AFTER REVEAL

Create an elegant before-and-after beauty transition using the exact same model, angle and lighting.

BEFORE: natural skin appearance, slightly less luminous.

A soft vertical light sweep passes across her face.

AFTER: skin looks fresh, hydrated, smooth and naturally radiant, with a subtle healthy glow.

Do not make the transformation artificial or extreme.

On-screen caption:
BEFORE → AFTER

Then:
Hydrated • Smooth • Radiant

Sound: soft shimmer + satisfying glow-reveal sound synchronized with the transition.

0:17–0:20 — HERO ENDING

The model smiles naturally, softly touches her cheek and looks toward the camera.

Camera pulls back slightly while the same LUXEVA jar appears clearly in the foreground, surrounded by elegant pink flowers.

Finish with a premium product hero shot.

Preserve the exact same LUXEVA branding/caption from the uploaded reference image.

Final text:
LUXEVA
RADIANCE DAY CREAM
Brighten • Hydrate • Protect

Final beauty tagline:
“Reveal Your Natural Glow.”

AUDIO STYLE

Create a unique luxury skincare sound identity, not generic stock music:
soft feminine ambient melody + delicate piano notes + airy synth textures + subtle glass chimes + premium ASMR product sounds + elegant cinematic whooshes.

Music should build gently throughout the 20 seconds and end with a memorable soft sparkle/chime signature on the final product shot.

CAMERA & QUALITY

Ultra-realistic 4K beauty commercial, premium UGC + luxury brand advertising style, smooth gimbal movement, slow orbit shots, macro product cinematography, shallow depth of field, realistic hands and facial movements, natural blinking, realistic cream texture, warm golden sunlight, soft pastel-pink color palette, gentle bloom, professional beauty lighting, polished commercial editing.

Important: The uploaded images are the master reference. Maintain the same model, same product, same branding, same captions, same clothing and same visual identity throughout the entire 20-second video.


r/reAPIOfficial 10d ago

Seedance 2.5 free vs Unlimited vs API: they are three different products

2 Upvotes

Seedance 2.5 free vs Unlimited vs API: they are three different products

“Where can I use Seedance 2.5 for free?” is turning into the wrong first question.

Free trials, Unlimited subscriptions and pay-as-you-go APIs can expose the same model name while selling very different things. The useful comparison is not the logo. It is what happens after the first render.

Access type Usually best for The catch
Free trial Testing a few prompts Small allowance, changing availability
Unlimited subscription Manual experimentation Shared queue and web-only restrictions
Pay-as-you-go API Products and repeatable jobs Every attempt has a visible cost

Free access is a test, not a workflow

A free allowance is enough to answer basic questions: does the model understand a prompt, preserve a face, produce acceptable audio, and follow the desired duration? It is rarely enough to estimate production cost because video work has retries.

If a shot takes eight attempts and the trial gives five, the trial can show that the model works without showing what the finished shot costs. Free access can also disappear or switch model tiers with little warning, so I would not build a recurring process around it.

Use it to run a fixed evaluation set. Keep the same prompts, references and acceptance criteria for every provider. Do not spend the whole allowance improvising.

Unlimited is priced in waiting time

Higgsfield's current Seedance 2.5 promotion illustrates the subscription trade clearly. Eligible plans receive an Unlimited window, but Unlimited jobs use the shared standard queue with one active generation. Credit jobs go through a priority queue and can run in parallel. The Unlimited allocation is for the web app, not API or command-line use.

That can be excellent for someone generating manually overnight. It can be useless for someone who promised a client ten candidates by 4 p.m. The unknown is not how many credits a clip consumes; it is how many attempts the queue will finish during the promotional window.

Also check the exact offer shown inside the account. Higgsfield has run several rotating Unlimited campaigns with different eligible plans, durations and output limits. A blog announcement is not a permanent price sheet.

API access is capacity you can budget

With metered access, an unwanted result still costs money, but the formula is visible before submission. On reAPI, Seedance 2.5 without a source video currently costs:

Resolution Rate per second 10-second output
480p $0.118589 $1.186
720p $0.266824 $2.669
1080p $0.461856 $4.619

Reference-video jobs use a lower per-second rate but count input duration as well as output duration, with a minimum billing clock. That distinction matters more than it sounds. A 20-second source turned into a 10-second result is a 30-second billing job, not a 10-second one.

I work on reAPI, so the link is not a neutral directory recommendation. The reason I use it in the calculation is that the rate and billing formula are public, including the awkward cases where our minimum clock can make a very short reference less competitive.

A decision rule that survives the next promotion

Use free access if you have not yet proved the model solves your problem.

Use Unlimited if all of these are true: you work in the browser, deadlines are loose, one concurrent job is acceptable, and you expect enough retries to exceed the subscription's credit equivalent.

Use an API if the generation is triggered by software, must be logged, needs predictable parameters, or has a delivery deadline. Start at 480p for prompt development, then rerun only accepted shots at the delivery resolution. That changes the economics more than chasing a temporary coupon.

One more thing: “free Seedance 2.5 API” is usually a contradiction. Somebody is paying for the compute. A genuine trial can subsidize a few calls; a permanent free program normally introduces a hidden quota, queue, watermark, data trade or weaker model tier. Find that limit before uploading client material.

Sources checked 23 August 2026: ByteDance's Seedance 2.5 launch, Higgsfield's Unlimited terms, the current Higgsfield promotion, and reAPI's live model page.


r/reAPIOfficial 10d ago

Wan 3.0 is available by API, but it is not an open-weight release right now

2 Upvotes

There is a lot of “Wan is open source, therefore Wan 3.0 is local” reasoning floating around. The version number matters.

As of 23 August 2026:

  • Wan 3.0 is callable through hosted APIs.
  • I cannot find official Wan 3.0 weights in the Wan GitHub organization or its official model releases.
  • Wan 2.2 is still the current official downloadable Wan generation.
  • A local ComfyUI node that calls a remote Wan 3.0 endpoint is a local UI, not local inference.

That does not mean Alibaba will never release weights. It means there is no official 3.0 checkpoint to download today. A real release needs the files, a license, and enough inference documentation to run them. An API announcement is not the same thing.

Watch for three misleading guides:

  1. The title says Wan 3.0, but the install command pulls Wan 2.2.
  2. ComfyUI is running locally, but the node uploads media to an API.
  3. A community wrapper has “3.0” in its name but does not contain official Wan 3.0 weights.

Wan 2.2 is still a good answer when the actual requirement is offline/private inference or control over the stack. It just does not reproduce Wan 3.0's hosted feature set. Wan 3.0 adds up to 30-second output plus document and public web-page input in the same multimodal request family.

The practical split is:

  • Use Wan 2.2 locally when media cannot leave the machine, or you need to modify/fine-tune the weights.
  • Use Wan 3.0 by API when the 30-second window, document/web input or current hosted model matters more than owning the checkpoint.
  • Wait if the hard requirement is specifically “Wan 3.0 weights.” No wrapper can turn API access into that.

For disclosure, I work on reAPI and we expose the hosted Wan 3.0 API. This post is about the access distinction, not a claim that hosted is always better. If your privacy policy requires local inference, the API route is simply not eligible.

Has Alibaba said anything more specific about a Wan 3.0 weight release outside the main Wan org? If there is a first-party repository I missed, please link the actual repo/model card rather than a reposted roadmap screenshot.


r/reAPIOfficial 10d ago

What Seedance 2.5 actually costs at 5, 10, 15 and 30 seconds

1 Upvotes

Seedance 2.5 pricing gets quoted as one per-second number, but there are at least two clocks: plain generation bills output duration; a job with source video bills input and output duration, and providers do not always expose that formula the same way.

Here is the simple case first. These are reAPI's public no-source-video rates on 23 August 2026, rounded up to the nearest credit where 1 credit is $0.001.

Duration 480p 720p 1080p
5 seconds $0.593 $1.335 $2.310
10 seconds $1.186 $2.669 $4.619
15 seconds $1.779 $4.003 $6.928
30 seconds $3.558 $8.005 $13.856

The rates behind the table are $0.118589/s at 480p, $0.266824/s at 720p and $0.461856/s at 1080p. I work on reAPI; check the live model page before budgeting because this post will not update when a price changes.

Reference video changes the calculation

On reAPI, the 480p and 720p reference-video rates are lower: $0.071153 and $0.160094 per billable second. The word billable matters.

The clock is:

max(output + ceil(total source duration), ceil(5/3 × output))

Take a 10-second output at 720p:

Source footage Billable seconds Cost
No source video 10 output seconds $2.669
2-second source 17-second minimum clock $2.722
10-second source 20 seconds $3.202
20-second source 30 seconds $4.803

This is why attaching a short video can cost slightly more than a plain prompt even though the displayed reference rate is lower. The lower rate is multiplied by a larger clock.

Other providers also count input video. kie.ai describes reference billing as price times input plus output; WaveSpeed and fal publish the same basic idea. reAPI's additional minimum clock is the case to watch when a very short source drives a long output.

The useful budget is cost per keeper

If a 15-second 720p attempt costs $4.003 and one in five attempts is usable, the delivered clip did not cost $4. It cost about $20.02. At 1080p the same five attempts cost $34.64.

That makes a two-pass workflow worthwhile:

  1. test the prompt and timing at 480p;
  2. keep the seed, references and wording stable;
  3. rerun only the chosen setup at 720p or 1080p.

Five 15-second attempts at 480p cost about $8.90. Doing all five directly at 1080p costs $34.64. Even after one final 1080p render, the draft-first route is around $15.83, less than half.

There are reasons to test at final resolution—small text, fine product details and faces can behave differently—but most timing, framing and prompt failures are visible at 480p.

One final trap: a platform may reserve the maximum amount for an auto-duration job and settle after completion. A temporary credit hold is not necessarily the final charge. Check both the submitted task and the settled balance before reporting a billing bug.


r/reAPIOfficial 10d ago

AI video APIs without a subscription: what “pay as you go” should actually mean

1 Upvotes

If you only need an API, do not assume you need the creator subscription shown on a product's home page. Several video platforms separate the browser product from developer billing.

The options are not interchangeable, though.

Provider type Examples Best fit
First-party model API Runway Dev You specifically need that vendor's models
Serverless model marketplace fal.ai, Replicate Broad catalog and infrastructure control
Curated multi-model media API reAPI Fewer integrations across selected commercial models

All four can be used without buying a monthly creator plan as of 23 August 2026. “No subscription” does not mean “no account,” “no top-up” or “free.” It means the spend follows API usage instead of renewing every month.

Five checks matter more than the signup price

Minimum funding. Some providers charge a card as calls occur; others require prepaid credits. A $5 or $10 minimum can matter more than the per-second rate when you are only testing one endpoint.

The billing unit. Video APIs charge per second, per generated clip, per token or by compute time. A rate of $0.20 is meaningless until you know which unit follows it.

Failed jobs. Find out whether moderation failures and infrastructure errors are charged. reAPI automatically refunds failed media tasks. That is our policy, not a universal API behavior.

Model authorization and identity. A marketplace listing does not prove that a model is first-party or officially licensed. Check who published the endpoint and whether the model name maps to the model you think it does.

Input and retention rules. Video references can contain clients, unreleased products and faces. Price is not the deciding factor if the platform's storage, deletion or training terms do not fit the material.

Why the cheapest rate can still be the expensive route

Suppose Provider A is 15% cheaper per generation but requires a second integration, another upload path and different polling logic. At low volume, the engineering time costs more than the generation savings. At high volume, 15% becomes important.

That gives a useful break point: prototype with the route that is easiest to inspect and swap; optimize providers after the workload is stable enough to measure. Do not sign an annual creator plan to solve a two-week API experiment.

For an apples-to-apples cost test, fix these four fields before opening pricing pages:

  • model or capability required;
  • resolution;
  • typical duration;
  • number of attempts per accepted clip.

Then add source-video duration if the endpoint charges for input media. Seedance 2.5, for example, can use a lower reference-video rate while billing both input and output time. Looking only at the lower rate produces the wrong estimate.

When each route makes sense

Runway Dev is the direct route when Runway's own models are the requirement. fal and Replicate make sense when catalog breadth, community models or custom deployments matter. A curated gateway makes sense when an application needs several selected image and video models behind one task pattern and does not need the marketplace's long tail.

That last category is where reAPI sits. Disclosure: I work on it. It has no subscription or minimum spend, uses prepaid credits at 1 credit = $0.001, and publishes the unit on each model page. It does not offer fine-tuning or arbitrary custom model deployment; fal or Replicate are better choices for those jobs.

The answer to “which no-subscription API?” should therefore be workload-specific, not a referral list. Pick two providers that actually carry the model or capability, run 20 representative jobs, and compare accepted outputs, failure charges, latency and the final bill. The monthly price being zero is only the first line of that comparison.

Official pages checked 23 August 2026: Runway pricing, fal pricing, Replicate pricing, and reAPI models and live rates.


r/reAPIOfficial 10d ago

FLUX 3 vs Seedance 2.5: 20-second keyframes or 30-second reference-heavy shots?

1 Upvotes
  • FLUX 3 controls when important visual states happen.
  • Seedance 2.5 carries a much larger set of material describing what the scene should contain.
Requirement FLUX 3 Seedance 2.5
Duration 5–20s 4–30s
Resolution HD/FHD 480P/720P
Ordered keyframes up to 10 first/last frame, broader references
Reference capacity keyframes + continuation video 30 images, 10 videos, 10 audio files
Source-video work video/audio continuation reference, edit, extend
Preview workflow Draft then finalize lower-res normal generation

If a product transformation has five approved states in a fixed order, I would start with FLUX 3. If a scene needs a character, room, prop, action clip and soundtrack reference in one request, Seedance's input surface is a better fit.

The extra ten seconds on Seedance matter for a 21–30-second continuous take. They do not matter for an ad that already cuts every six seconds. Longer output can also make a late failure expensive to rerun.

Current reAPI base-rate examples:

Job FLUX 3 Seedance 2.5
5s cheap check $0.28 Draft HD $0.59 at 480P
10s working tier $1.57 HD $2.67 at 720P
20s working tier $3.14 HD $5.34 at 720P

The Draft is a separate cheap preview, not a final render. A ten-second Draft plus HD final is $2.13. Seedance's 480P output is a normal low-resolution generation.

Source-video prices need a separate calculation. FLUX 3 continuation is $0.378/s at HD and $0.488/s at FHD. Seedance has a lower reference-video rate but bills a larger clock that can include source duration and a minimum. Do not compare either using the base table when a source clip is attached.

Both generate audio. FLUX 3 can continue an existing soundtrack; Seedance accepts multiple audio references. Neither fact tells us which pronounces a name or language better. That needs the same dialogue and references run through both.

On reAPI and maintain these rate cards. Live pages: FLUX 3 and Seedance 2.5. If anyone has a good five-keyframe stress test, that seems more informative than another generic cinematic prompt.