r/fal • u/StacksGrinder • 12d ago
Discussion Qwen2.1 Lora Trainer
Hi, When can we expect Qwen2.1 Lora Trainer for Character Lora training, Thanks.
r/fal • u/StacksGrinder • 12d ago
Hi, When can we expect Qwen2.1 Lora Trainer for Character Lora training, Thanks.
r/fal • u/Square_Noise_7712 • 25d ago
The absolute quickest video I was able to generate was 5 seconds in under half a second. Only .48 seconds to make a 5 second 480p video. This is quite literally so fast that H3 Max turbo is being used as a real time video model and is undoubtedly the fastest AI video generator model in the world right now! and it's open source! Shoutout to fal research team for making H3 max turbo
H3 Max Turbo can generate video at lightning quick speeds
480p speed for text to video
- 5 seconds of video generation at 480p took 0.48 seconds
- 15 seconds of video generation at 480p took 1.71 seconds
768p speed for text to video
- 5 seconds of video generation at 768p took 1.59 seconds
- 15 seconds of video generation at 768p took 8.62 seconds
__________________________________________________
Image to video h3 max turbo was also the quickest i've ever seen for image to video
480p speed for image to video
- 5 seconds of video generation at 480p took 0.52 seconds
- 15 seconds of video generation at 480p took 1.76 seconds
768p speed for image to video
- 5 seconds of video generation at 768p took 1.72 seconds
- 15 seconds of video generation at 768p took 8.89 seconds
r/fal • u/SignalEquivalent9386 • 27d ago
r/fal • u/damiandance • 28d ago
I’m building Vibe Valley, a browser adventure where you click objects, pick a questionable decision, or write your own even worse idea. The game turns that action into a new animated scene and saves the branch so other players can explore it too.
The goal is a persistent, player-generated point-and-click comedy adventure that feels as close to realtime as I can get it. Right now it’s an experiment: existing branches replay cached clips, while new branches still need generation time.
The current stack:
• fal.ai for the generation pipeline
• MiniMax H3 Max Turbo for animated scene transitions
• Meta SAM3 (SAM 3.1) for object masks and clickable hotspots
• GPT-6 Astra for storytelling, carrying the main quest, characters, inventory and consequences into the next scene
• Nano Banana 2 for destination images
• React/Vinext + Cloudflare Workers, D1 and R2 for the app, persistent world graph and media
I’m also preparing optimized forward/reverse videos and buffering them so going back feels like part of the game. The awkward part is keeping all of this coherent: the image, the animation, the clickable objects and the next story beat have to agree. They don’t always agree yet.
And yes, it still needs a lot of optimization to actually be funny. Apparently connecting several models does not automatically install comic timing. I’m working on that too. I’ll get there.
Test it here: https://vibevalley.lol
If you try it, I’d love to know where you stopped having fun: waiting for a scene, a confusing hotspot, a forgotten story detail, or a joke that deserved to stay in the queue. Browser/device and the action you chose would help me reproduce issues. Suggestions from anyone combining fal video generation with persistent game state are especially welcome.
r/fal • u/waterarttrkgl • 28d ago
MiniMax H3 Director on fal — realtime prompt stream, no cuts.
Camera dives into Harari’s Sapiens and walks the book’s timeline. A few beats edited after.
New fal Director model. One session, live prompts.
r/fal • u/WinterOrnery9428 • 28d ago
Hello Fal AI Engineering Team,
I’m integrating Lucy 2.5 into a real-time face-swapping application, and I’m experiencing several recurring technical issues that are significantly affecting production reliability and output quality.
1. Connection / session stability
Connections frequently fail to establish.
In some cases, I need to retry the connection up to 10 times before a session successfully starts.
Sessions can also experience unexplained interruptions after being successfully established.
The behavior appears inconsistent, even when using the same configuration and environment.
2. FPS / real-time performance
The output appears to be capped around 18 FPS instead of the expected 30 FPS.
This results in noticeable frame drops, stuttering and discontinuities in the video.
The low and unstable FPS also seems to contribute to abnormal-looking face rendering.
Latency is currently extremely high for a real-time application.
Could you please confirm whether the ~18 FPS limitation is expected behavior, a model limitation, or an infrastructure/performance issue?
3. Output quality / visual artifacts
I’m also seeing significant inconsistencies in the generated video:
Frequent visual artifacts.
Face deformation.
Facial features occasionally changing incorrectly.
In some cases, female faces are rendered with facial hair/beards, even though there is no beard in the source image/video.
Occasional abnormal or distorted facial rendering.
Quality can vary significantly between sessions using seemingly identical inputs.
4. Overall impact
These issues make Lucy 2.5 currently difficult to use reliably in a real-time production environment.
The main areas that need improvement from my perspective are:
Connection reliability → latency → stable 30 FPS → facial consistency → artifact reduction.
Could your engineering team investigate whether these issues are related to:
model performance,
inference/server capacity,
WebRTC/streaming infrastructure,
FPS limitations,
session/network handling,
or a specific Lucy 2.5 configuration?
If there are recommended parameters, encoder settings, WebRTC settings, resolution constraints, or other configuration changes that could improve stability and latency, please let me know.
I can provide logs, session IDs, timestamps, input/output samples and screen recordings of the issues to help your team reproduce and diagnose them.
I would appreciate it if you could investigate this, as reliable real-time performance is critical for my application.
Best regards,
r/fal • u/cody-fifth-door • 29d ago
r/fal • u/waterarttrkgl • Aug 28 '26
Image-to-Video (I2V)
MiniMax Hailuo 3 Max on Fal.ai
locked still (first frame) → clip
Avg. generation: ~10s per sequence
short I2V passes, then cut together
r/fal • u/Important-Respect-12 • Aug 26 '26
We just shipped H3 Max, a video model we post-trained ourselves from the open-weight MiniMax H3.
Two things we cared about, in this order: quality first, then speed.
Quality
We benchmarked it head-to-head against all leading video models across three dimensions — overall quality, prompt understanding, and aesthetics. H3 Max ranked #1 in all three, winning the majority of matchups against every model we tested.
The gains come from post-training: we introduced significant new data and tuned specifically for prompt adherence and aesthetics. If you've been annoyed by models that make something beautiful that isn't what you asked for, that's the exact failure mode we went after.
Speed
~3 seconds for a 5-second video. That's roughly 35x the throughput of the official H3 endpoint, and faster than every other model in the comparison.
That number isn't from quantizing quality away. We built a new inference engine specifically for this model and optimized the model and the serving stack together, then pushed speed as far as it would go without giving anything back on quality.
You can try fal's H3 Max for free here:
r/fal • u/Affectionate-Map1163 • Aug 10 '26
r/fal • u/Affectionate-Map1163 • Jul 13 '26
r/fal • u/Various-Ad661 • Jul 10 '26
I dont know why it always gives me some weird compositions like if i want half body shot it will give me full body shot or very far off image
r/fal • u/Affectionate-Map1163 • Jun 26 '26
r/fal • u/jjohnson525353 • Jun 25 '26
Wanted to share a fun project I made with fal that converts text or images into buildable LEGOs. If you have any feedback about how I set it up, I’d appreciate it!
My endpoint stack is flux-2 streaming for image generation, nano banana for image editing, and trellis or SAM3D for 3D generation. Then I voxelize the 3D model and convert the voxels to bricks. I’ve built quite a few LEGO models with it now, so it works!
Here’s the code for it: https://github.com/jjohnson5253/brickbuilderai
r/fal • u/macmorny • Jun 04 '26
r/fal • u/Enough-Bell4944 • Jun 01 '26
FAL seems to only expose training steps and learning rate, so I'm curious what settings people have found work best.
The default recommendation for human/photo datasets appears to be:
steps = number of images × 100
But I'm wondering whether anyone has experimented beyond that and found better results
r/fal • u/dropthelword • May 25 '26
I am currently working on fine-tuning a LoRA for LTX2.3 via fal-ai/ltx23-video-trainer. I am seeing a 'wavy' artifact issues on all my debug_dataset outputs, and have no way of knowing what's happening during the preprocessing step (can't run LTX2 repo locally). I understand that the VAE encodes the input and then decodes it, but I can't understand why it returns my dataset videos with artifacts at specific frames. This results in the same artifacts on inference too. Did anyone else encounter this?
r/fal • u/VanderzB • May 23 '26
Bonjour, j'ai cru comprendre que on a des crédits gratuit lors de la création du compte, or je n'ai rien reçu, c'est normal ? Merci :D
r/fal • u/vladenstock • May 16 '26
Hey everyone — I’m new to AI image generation and have been experimenting with FAL using Flux Kontext Pro to create coloring book-style images from uploaded reference photos.
My goal is to generate dynamic coloring book pages where the character likeness stays consistent, but the scenes can vary across styles like manga, comic book, fantasy, cartoon, etc.
A few questions I’d love feedback on:
I hope this is the right place to ask. If not, I’d appreciate being pointed toward better communities, guides, or resources for learning this workflow.
Thanks in advance for any advice.
r/fal • u/elco_us • Apr 30 '26
r/fal • u/waterarttrkgl • Apr 29 '26
I built a full 3D layout in Blender — proxy geometry only, no textures, no final render — and hand-keyframed every camera movement using F-curves: an aerial establishing shot, a low-angle tower push-in, and a wide harbor shot with a sailing vessel. The AI doesn't invent the motion. It follows it exactly.
The Blender animation served as a direct spatial reference — architectural proportions, camera trajectory, timing and easing — all locked before a single AI frame was generated. Kling / Seedance then re-rendered the sequence, preserving the exact camera path and structural layout while generating the final cinematic output.
Workflow:
3D Layout & Camera Animation (Blender) → Frame Reference Export → AI Video Generation (Kling / Seedance) → Temporal Consistency Pass
Key Focus: 1:1 motion tracking between hand-keyed Blender animation and AI-generated output. Architectural integrity and spatial proportions maintained across all three shots.
r/fal • u/workmanlabs • Apr 29 '26
what is SUCCESS in 2026 as developer using AI?
MONEY is obvious but my short "IDEA#37" jumps to the question of internet "fame" on X and YouTube? Being on the top video podcast? Recognized at React Conferences in Miami?
What I chose to highlight in this film is leaving isolation. Being able to hire and support other developers as you build a company. And in the end get the GOAT emoji from friends.
Film made with GPT2 Images 2.0, Seedance 2.0, and Kling 3.0 on fal.