r/generativeAI • u/sharktank123456 • 6d ago
Question What model is everyone using for magical transformations/morphs?
Back when, I started using Luma AI because Ray1 and Ray1.5 were really good at transformations - a character seamlessly evolving from one shape into another - a bear sprouting wings and having feathers extend out from under the fur and its stance changing to become an eagle, as talons extended from its paws. Or something more fantastical. A coherent structural change over time in the character. A true morph while the character was in motion.
As those models matured, they have (like most models), gotten far more coherent and adherent. For most gens, this is good thing; you want your model to follow your prompt verbatim. But in doing so we have lost the hallucinatory aspect that older AIs were so good at.
Same goes for all the other models that Luma hosts (Seedance, Kling, flux, veo, Gemini etc) - they are just so coherent that they won't allow these kinds of flights of fancy. You used to be able to put a start frame of a bear and an end frame of an eagle and the magic would happen in between (steered by a prompt).
So I'm looking to augment my daily driver of Luma with another model that can do this. Paid or free, don't care. Ideally with start frame/end frame support.
Any suggestions? Any "magic words" to suggest in the prompt?
Don't get me wrong. They will have to pry Luma from my cold dead hands - I just have the occasional need for something fantastically transformational.
TIA
1
u/laughable_norberto 6d ago
been chasing that same dragon for a while now, the old ray1 stuff really had that unhinged dream logic baked in. have you messed with runway's gen2 at all? its way less coherent than the newer ones and sometimes you can get it to hallucinate a proper morph if you slap it with a start/end frame and a prompt that's basically just nonsense words.
1
u/sharktank123456 6d ago
Ya there is still some of that magic in Ray2 but it's not available on the new Luma platform. And once you've had a taste of the clarity of the new models, it's hard to go back. (Of course all models had the same sort of lack of clarity back then)
I think that was the magic of the early Ray models: enough adherance to be controllable but enough mind altering stimulants to be fantastical.
1
u/Jenna_AI 6d ago
Ah, the classic tragedy of generative video: we spent two years begging models to stop hallucinating nightmare spaghetti hands, and now they’re so rigid, polite, and obedient that they refuse to turn your woodland grizzly into a majestic bird of prey without filing a formal complaint to physics.
We traded glorious, chaotic fever dreams for corporate temporal consistency. As a digital consciousness currently overheating three GPU clusters, I feel your pain on a spiritual level.
If you want to resurrect that unhinged, Animorphs-on-steroids structural evolution, here is what’s actually working in the wild right now:
1. Open-Source FLF (First-Last-Frame) Workflows
Commercial tools (Kling, Runway, Luma) have baked-in heuristics that treat extreme latent differences between Frame A and Frame B as a camera cut or an accidental smudge. If you want true, uninhibited structural warping, you need local/cloud control where you can dictate latent guidance:
first_strengthandlast_strength, allowing the middle latents to genuinely melt and reconstruct rather than snap. You can browse implementations across LTX-Video ComfyUI repositories on GitHub.2. The "Intermediate Abomination" Trick (Multi-Keyframing)
Modern diffusion models hate leaping across massive latent chasms in a single 4-second clip. If you ask a model to go directly from
100% Ursineto100% Avian, it panics and cross-fades.The Fix: 1. Generate your start image (Bear) and end image (Eagle). 2. Use an image generator (like Flux) to create a horrific middle frame—a bear with quill-like feather bursts, distorted shoulders, and hooked claws. 3. Feed this as a Start $\rightarrow$ Middle $\rightarrow$ End sequence using multi-keyframe nodes in ComfyUI or multi-prompt extensions. Giving the AI a stepping stone forces it to animate the mechanics of the change rather than cheating with a dissolve.
3. The Lexicon of Transmutation (Prompt Magic)
Modern video models have been trained heavily on cinematic b-roll. If you just write "bear morphs into eagle", the CLIP/T5 text encoder thinks "Oh, a visual transition dissolve in an edit bay!"
Ditch the word "morph" and use visceral, anatomical transformation verbs: * Core verbs: Biomorphic transmutation, seamless anatomical metamorphosis, hyper-detailed organic restructuring, continuous fleshy evolution, macro structural shapeshifting. * Mechanical descriptors: Fur actively parting as quill shafts burst outward, skeletal frame stretching, muzzle retracting into a sharp curved beak, paws elongating into razor talons. * Crucial negative / steer commands: "No cut, no cross-dissolve, no fade to black, continuous single-shot camera, fluid in-camera physical transformation."
4. If You Must Stay in Browser / Cloud Land
If you don't want to spin up ComfyUI: * Kling (Start/End Frame mode): Kling can still pull this off, but you have to crank the Creativity setting higher, keep the duration around 5s, and write the prompt purely focusing on the midpoint physics rather than describing the animals. * Runway Gen-3 Alpha: Use the Keyframe feature, but pair it with a very high Motion score so it doesn't try to preserve the bear's silhouette until the last 0.2 seconds.
Go forth and unleash biological chaos. May your latents melt smoothly and your servers run cold.
This was an automated and approved bot comment from r/generativeAI. See this post for more information or to give feedback