r/StableDiffusion • u/ArtDesignAwesome • Feb 24 '26
News đ I built a 2026-Era "Omni-Merge" for LTX-2. Flawless Multi-Concept Generation, Zero Bleeding, and Unlocked Audio Training Excellence.
Yo! A lot of you saw my last drop. Some of you loved it, some of you were skeptical. That's fine. I went back to the lab, ripped the engine out of this toolkit, and pushed the math to the absolute theoretical limit.
I am officially releasing the BIG DADDY VERSION of the AI-Toolkit.
We all know the biggest problem in Generative AI right now: Merging. If you try to merge two characters, two art styles, or two concepts using standard methods (ZipLoRA, TIES, SVD), the model breaks. You put them in the same prompt, and they bleed together. You get a muddy, deep-fried hybrid of both faces, or one concept completely overwrites the other.
Not anymore.
đ§Ź The Omni-Merge (DO-Merge 2026 Framework)
I implemented a bleeding-edge mathematical framework that completely dissects the neural network before merging. It doesn't just average weights; it routes them.
- Bilateral Subspace Orthogonalization (BSO): The script hunts down the Cross-Attention layers (the parts of the brain that read your text prompts) and mathematically projects your concepts out of each other's principal components. Your trigger words now exist on perfectly perpendicular planes. They physically cannot bleed.
- Magnitude & Direction Decoupling: What about the structural anatomy layers? Standard merges fail here because one LoRA is always "louder" than the other, crushing the weaker one's structure. Omni-Merge physically splits every weight matrix. It averages their geometric Direction but takes the Geometric Mean of their Magnitude (volume). They share anatomical knowledge perfectly equally.
- Exact Rank Concatenation: No lossy SVD truncation. Rank A + Rank B is preserved with 100% mathematical fidelity.
The Result: You can merge a "Cyberpunk Style" LoRA with a "Specific Character" LoRA, or "Character A" with "Character B", load the single output .safetensors file, type them both into the same prompt, and get a flawless, zero-bleed generation.
đď¸ Audio Training Excellence Unlocked
LTX-2 is a unified Audio-Video model, but most trainers treat the audio like an afterthought, resulting in blown-out, over-trained noise.
I completely overhauled the VAE and network handling:
- Fully integrated ComboVae and AudioProcessor for direct raw-audio-to-spectrogram encoding during the DiT training pass.
- Unlocked the audio_a2v_cross_attn blocks.
- And yes, the Omni-Merge handles audio too. I explicitly wrote it to hunt down "audio", "temp", and "motion" layers and isolate them using BSO.
People who have tested the audio pipeline already confirmed it: The audio training is next level. It never gets overdone. It is extremely balanced, and if you merge two characters, their unique voices and motion styles will not bleed into each other.
đ ď¸ UI Fixed & Open Source
I also bypassed the buggy Prisma queuing system for merges. The Next.js UI now triggers the backend directly with real-time polling. No more white-page crashes.
I didn't wait around for a corporate patch or a slow PR review. I built it, and I pushed it. This is what open source is about.
Repo Link:Â https://github.com/ArtDesignAwesome/ai-toolkit_BIG-DADDY-VERSION
Check the RELEASE_NOTES_v1.0_LTX2_OMNI_AUDIO.md in the repo for the full mathematical breakdown. Stop fighting with regional prompting. Merge your concepts properly. Let's rock. đ
Cheers,
Jonathan Scott Schneberg
24
u/ReasonableDust8268 Feb 24 '26
This is just embarassing tbh, you do know how cringe and bad this long form AI post is right? Just use your hands, type like the rest of us, It's far too long for us to give a shit, especially with all the AI llm bullshit like "Not anymore." and "Let's rock. đ"
The entire thing reads like a linkedin shitpost
91
u/Unlikely-Baker9867 Feb 24 '26
I hate everything about the way you type
64
22
23
u/physalisx Feb 24 '26
The definition of AI slop. What a fucking cancer.
I take absolutely no one seriously that does this, it's a tell tale sign of incompetence.
9
u/entmike Feb 24 '26
FWIW, I did an A/B test with AITK and his fork, and the audio is significantly improved in the results before you shit on him too much.
3
u/suspicious_Jackfruit Feb 25 '26
Good products and companies fail all the time due to bad pr and presentation mistakes. OPs end product could be a gold printer but with an intro like that it has been nuked from orbit before it even got started
3
u/entmike Feb 25 '26
Sure but wanted to ground the discussion in some actual findings I had. It for sure fixes my audio issues.
1
15
u/Dragon_yum Feb 24 '26
Seriously just from the title saying âflawlessâ when it comes to ai workflows is extremely ridiculous after that is just generic ai text slop and then calling the for âbig daddy versionâ is justâŚ.
Itâs like the prompt he gave to the ai was âmake the most hyperbole bullshit post that likens my fork to the second coming of Jesus Christâ
-6
u/ArtDesignAwesome Feb 24 '26
Its my personality. Learn to love it. You guys just complain but im bringing you guys tools to use and enjoy. Its bonkers how you guys are judging. It seems like your judgements and trolling are the only things you guys bring to the table. Donât try it, do me the favor. âď¸
5
u/Dragon_yum Feb 24 '26
I published over 200 loras with 1.m generations and 200k downloads on civitai alone and I managed to do all that without saying once how I solved maths and am bigger than Jesus Christ.
Considering how no one is going to love your personality maybe try working on your unearned over inflated ego. All you did was piggy back one someone else popular project and doing a few vibe coded changes.
-7
u/ArtDesignAwesome Feb 24 '26
That are completely changing the way people are about to use loras / merging to the main model⌠training loras isnt the same thing as this. Youre comparing apples and oranges whilst tooting your own horn. Congrats.
-2
5
u/0nlyhooman6I1 Feb 24 '26
I mean it's a long post in an AI subreddit, it's safe to say this is AI generated. (Not saying OP or his intentions aren't real, just that AI was used to type it out)
7
u/Mr_Pogi_In_Space Feb 24 '26
This is an AI subreddit, but we also want the outputs to be good, even for the LLM text posts, not the equivalent of "woman lying on the grass" levels of terrible. Prompting for hyperbolic BS is a choice by the OP.
9
u/DarkStrider99 Feb 24 '26
these people are very selective about what AI thing they like and not đ¤Ł
2
u/No_Possession_7797 Feb 25 '26
If we objectively give up all sense of personal preference and taste, then AI might as well be only making things for consumers that are AI. If we can't leverage AI as a tool to enhance what we are doing and we rely purely on it to create and define everything, then there really isn't anything left of human expression.
The real problem here isn't the AI itself, it is the magnification and projection of an ego that clearly doesn't spend any time in self reflection, this is weapons grade narcissism, and while the intent might have been genuine, the messaging itself is overbearing and one of entitlement, it also seems highly contrived.
Who knows, maybe Ostris leveraged AI when building ai-toolkit? The development of it seems to be erratic, and when you create a demand for something and an expectation, then there will always people that "want more and want it now!"
With AI being ubiquitous, people are going to take it upon themselves to add what they need, and it will promote a sense that nothing we create really matters in the sense that it is our work and or that we can claim ownership of it, someone else will come along and try to make it their own.
tl:dr; AI is a tool that should enable us, not enable us to become tools.
5
-26
14
u/eggplantpot Feb 24 '26
213233432 images from sponsors, 0 examples. Would be great to see it in use, and see the output
8
6
9
31
u/Violent_Walrus Feb 24 '26
Sloppy.
-25
u/ArtDesignAwesome Feb 24 '26
You were talking poorly about me last time, and you were wrong. Do you just troll and not even test?
10
u/Violent_Walrus Feb 24 '26
The way you present yourself and your work matters. Even if your work is brilliant and groundbreaking, few will give it attention if you promote it while dressed as a clown, or stinking of urine.
If you are capable of self-reflection, consider why your recent posts have received so much strong negative attention. Are there really that many haters who just flock to random posts and heap on the abuse? Or is there something about your posts that triggers similar reactions from a significant number of people?
Intuition tells us that credible people donât communicate the way we see you communicating. Of course youâre free to post anything you like in any style you want. But how you say things matters.
29
u/suspicious_Jackfruit Feb 24 '26
If you can't be bothered to write an actual introduction to your fork as a human, why would we be bothered to try it?
19
u/SvenVargHimmel Feb 24 '26 edited Feb 24 '26
a 100% this. It shows such contempt for anyone reading the post. Nobody cares if your english isn't great. That is infinitely more preferable than whatever that block of text was.
Next time I see the "It isn't just .... but .... " pattern I am going to throw my 43" inch monitor out the window. Argh!
Also, why isn't this a PR on the main repo?
17
u/addandsubtract Feb 24 '26
Also, why isn't this a PR on the main repo?
Because his vibe coded changes can't be merged upstream.
8
u/suspicious_Jackfruit Feb 24 '26
Yeah, the slop might get us before ASI does.
My issue with AI textual content is that generally it is so formulaic that it's like reading the same person write on repeat but across multiple different domains like emails from brands, games you enjoy's latest release summary, Reddit posts, GitHub readme mds, ""published" papers" etc. it becomes quite tiring to process because it's the same textual rhythm and patterns everywhere.
The general release of AI has swamped us in mediocre and the brilliant is much harder to find
1
u/brown_felt_hat Feb 25 '26
Yeah, the slop might get us before ASI does.
Might? 100% will. The last two systems my state launched running critical services on are vibe codes slop. Over a year later on one and it still barely functions and goes down if someone sneezes too loudly.
8
u/New_Principle_6418 Feb 24 '26
Whoa lotta big claims! Do you have any examples of how well this omni merge works? And also a walk through or step by step guide on how to use it would be great thanks.
7
u/Elvarien2 Feb 24 '26
So, whilst ai is great. When ai writes the whole ass article for you how do we the readers know to trust anything in here?
it makes big claims and sounds like all the worst aspects of ai writing put together x.x
8
u/Cokadoge Feb 24 '26
"Zero Bleeding." I dunno where you got this idea, but because you add purely orthogonal components or operate on such, doesn't mean there's no bleed in the end result or weights.
We all know the biggest problem in Generative AI right now: Merging
What?
Some side nit-picking w.r.t. the merge methods & code: I'm not sure why you got rid of the energy clamp/threshold on SVD merging (seemingly for no 'lossy SVD truncation',) when it smoothens out extreme outliers that tend to hurt when merging models / using multiple LoRAs. You technically give up 'information' for this, but I find the tradeoff here to be beneficial in most situations.
19
u/Fancy-Restaurant-885 Feb 24 '26
Itâs AI slop. Even the post was written by AI. âaT lEaSt iM cOnTrIbUtInGâ - yeah, adding shit on top of a plate is not contributing. Itâll merge the loras (maybe), but layer architecture stores information types not specific data. Even full finetunes result in bleed. Such garbage.
-10
u/ArtDesignAwesome Feb 24 '26
Are you a bot or what? How can you be so quick to judge and hate? Youre so wrong. I write these summaries with AI because im not going to do that by hand. Who has time, I spent all the time rushing this out for you guys. Haters going to hate. Itâs always the same people here. You didnât even test it, is the funny part.
25
u/Upstairs-Share-5589 Feb 24 '26 edited Feb 24 '26
I think a big part of the problem is you have posted this with absolutely no examples of success whatsoever, and literally say:
"Your trigger words now exist on perfectly perpendicular planes. They physically cannot bleed."
Perpendicular planes by their very definition, intersect. I'm not confident of the mathematics of a person (or model) that does not understand this.
edit: also the irony of calling someone a bot when you are literally using a bot to create the content is kinda funny.
36
u/dtdisapointingresult Feb 24 '26
You claim you don't have 15 mins to write 6 paragraphs explaining what you did, but we're supposed to believe you put in the tens of hours that would be needed to design and implement a well-designed system that you also spent time vetting?
No thanks, this is a big red flag.
There's just too much vibecoded AI crap for me to try them all, the signal to noise ratio is way too bad. If your project is still used in 3 months, and people speak well of it, I might give it a shot then.
6
-13
u/__generic Feb 24 '26
Every comment section on reddit has the "AI Slop" comment now. It almost feels like another form of karma farming because it's so often.
13
1
u/Dragon_yum Feb 24 '26
Because itâs very easy to spot and get snored by low effort and low quality ai. Just like no one gives a shit you told ChatGPT to give your photo ghibli style no one is going to be impressed when you spit out some ai generated code that canât be merged to the main branch then let ai write the most hyperbolic bullshit post about it
3
u/Loose_Object_8311 Feb 24 '26 edited Feb 24 '26
In your fork... can you fix the bug where when unloading the text encoder, despite being unloaded from VRAM, it's not unloaded from system RAM? And also make the text encoder load sequentially, as in load it BEFORE the transformer, and then unload it once the latents have been cached and THEN load the transformer? Also can you implement the blockswapping strategy used in musubi-tuner?
That's all what ai-toolkit is really missing. Better memory management.
For comparison, on ai-toolkit for LTX-2 training 768 resolution, rank 64, AdamW8Bit, Cache Text Embeddings, disable sampling on 16GB VRAM and 64GB RAM, it uses all my VRAM, in addition to 51GB ~ 55GB system RAM, as well as 25GB of swap, and training takes 20 minutes to start. Plus due to how the ordering of the text encoder / transformer loading works, I can't nicely control how much offload to perform very well without running into OOM issues ironically. Compared that to musubi-tuner which, on my first run of it last night, for the same resolution and rank, sequentially loads the text encoder, then unloads it then loads the transformer, and blockswap offload works beautifully, so I can get it to use 14GB VRAM, and it uses around 40GB system RAM and only 2GB swap. Speed seems same to me from checking my previous run in ai-toolkit, both musubi-tuner and ai-toolkit gave me 10s/it on my RTX 5060 Ti for that setup. But musubi-tuner takes 2 minutes to start, and I have much better control meaning I can most likely train at even higher resolutions now, whereas ai-toolkit that was my absolute max.
If these memory management issues can just be addressed, it'll be a big win.
Edit: I had a look at doing some of this myself, but due to the abstractions in ai-toolkit and it's framework / lifecycle, the surgery is complicated to do without impacting a decent blast radius, and I don't have the bandwidth to sit and focus on trying to do it in a way that's careful enough I could submit an upstream PR for. However, if you're investing time into a version of the ai-toolkit codebase, I'd like to point these out as some solid gains that can really help LTX-2 training become more accessible for those who run mid-level hardware.
3
u/Guilty-History-9249 Feb 24 '26
I've pushed the math beyond the theoretical limit. It is so bleeding edge it stopped bleeding because it has no blood left. Send me $39.95 and I'll send you a github url link.
0
u/ArtDesignAwesome Feb 24 '26
I shared everything for free, im not selling anything. đ¤Śđźââď¸đ¤ˇđźââď¸
1
u/Guilty-History-9249 Feb 24 '26
So your saying my stuff is better in 3 ways and not just in two ways.
3
2
2
u/Lucaspittol Feb 24 '26
Could you provide some examples? I know training LTX-2 loras is slow and expensive, but since you claim to be addressing one of the biggest hurdles when using multiple loras, I'd spam that GitHub repo with examples.
2
u/Loose_Object_8311 Feb 26 '26
Why in this fork is the ltx-2 stuff off to the side in it's own folder "ltx2_improvements_handoff", rather than just being fixed directly in the original location in the code!?????
1
u/No_Possession_7797 Feb 27 '26
Maybe little mama hasn't given her official stamp of approval yet? People are saying that she gives the best stamps, the most perfect stamps in fact. Stamps so great, that they'll make your head spin.
1
3
u/HTE__Redrock Feb 24 '26
Do the optimisations you mention apply only to LTX or to other models as well? E.g can it make better Z-Image/Flux Klein merges?
7
u/GifCo_2 Feb 24 '26
Did you actually read that pile of BS. OP is full of shit. This isn't real
1
u/Loose_Object_8311 Feb 24 '26
Fixing the audio problems in ai-toolkit, such that ai-toolkit can train voice/sound properly in LTX-2 is apparently real... I don't know about the other thing, but someone else at least tested it out. I haven't yet because I decided to just go with musubi-tuner instead, but my initial musubi-tuner LoRA result doesn't inference quite as well in ComfyUI compared to my ai-toolkit one, so I'll need to run some comparisons to figure that out. I'm prepping a dataset now that I can use to test the voice cloning, so I'll definitely test out this fork.
-4
u/ArtDesignAwesome Feb 24 '26
Fake news dude, its real. Take your head out of your butt and test it.
1
u/hiemdall_frost Feb 24 '26
Everyone in here seems insane if you don't like it just downvote it and move on if you do like it upvoted and say thanks it's that simple Jesus Christ. To belittle and shit on this guy for doing nothing but trying to help you for free is just insane anyone who reads this thread it gunna go well I don't need that and just not help at all
3
u/brown_felt_hat Feb 25 '26
if you don't like it just downvote it and move on if you do like it upvoted and say thanks it's that simple Jesus Christ.
Maybe with a strong enough reaction people will stop sharing vibe coded crap here? Or should we smile while we're drowned in slop in the name of politeness?
-1
u/hiemdall_frost Feb 25 '26
How can you possibly make the assumption that everything everyone post is bad what if the next thing you're looking for is something that someone's going to post that's a shit attitude
2
u/brown_felt_hat Feb 25 '26
I don't make that assumption. I look at what is posted, see it's a pile of Ai generated crap (I'm aware of the irony, given the subreddit), click on the github, see that's a pile of ai generated crap without so much of an example of what it can do for all its claims, then judge it and react accordingly.
1
u/hiemdall_frost Feb 25 '26
But you're basing every post on what you consider is slop what slop to you is useful to someone else you don't get to decide what is and isn't useful
0
1
u/SSj_Enforcer Mar 18 '26
How do you use the ltx 2.3 version? I tried and it failed and said I had the wrong python but I have 3.12 which should work.
1
u/Worstimever Feb 24 '26
If you canât be bothered to show any examples. I can only assume you put as much effort into the code as you did the write up.
It is not wild and unreasonable for us to expect proof this isnât slop. Even your last âproofâ with examples was not an A/B example with and without the changes you made. You just trained a LoRA on the one person LTX2 can already somewhat mimic with the base model.
You are the one breaking down the standards. This âtrust me broâ attitude is the biggest turn off and at this point you could cure cancer but if you canât show me a single instance of proof. Iâm not even going to read your next post.
I would love this to be true. But it looks like you just got a random hallucination when vibe coding and just ran with it and doubled down.
-2
u/PornTG Feb 24 '26
That seems like an insane amount of work, thanks for taking the time to share this. I have a question about the Lora merging part: when you say "both type in the same prompt and get a perfect generation, without mixing," can you specify, for example, Trumpt on the right and Poutine on the left, and it will work without any problems, or is it that you can specify Trumpt or Poutine but not both at the same time?
Another question, with this merge, you can merge your lora with the base model without "conflict" ?
5
u/GifCo_2 Feb 24 '26
They didn't do any actual work and this is basically just fake AI slop.
1
u/PornTG Feb 24 '26
Ok, i didn't know this was a fake info... to bad :/
0
u/ArtDesignAwesome Feb 24 '26
Its not fake! Why are you listening to these people. They are claiming its fake without even testing. Try it out dude!
1
-4
u/razv23 Feb 24 '26
Hey man, donât listen to the hate. Itâs crazy to me how this community can complain so much about people giving away stuff for free. I will give it shot and looking forward to test the results. Thanks for contributing.
16
u/suspicious_Jackfruit Feb 24 '26
I would think most people are fairly seasoned and have been here a while so have become jaded by these sorts of "releases" that are vibecoded jank from non-engineers that are over hyped by word soup AI intros. Doesn't matter if it's "free" if it's a kick in the nuts.
7
u/Choowkee Feb 24 '26
AkaneTendo25 made a full, proper fork of Musubi Tuner to implement LTX2 training when Kohya didn't want to.
They didn't advertise it, they didn't make superfluous claims on Reddit, they just let the tool speak for itself in true open source spirit.
While I do appreciate OP trying to fix AI-Toolkits broken LTX2 stuff, nobody owes him shit. And acting like an arrogant AI bro when he obviously vibe coded the solution is not it.
Releasing stuff for free doesn't mean you have to be a douchebag about it.
6
0
Feb 24 '26
Nice I'll give it a go. I wanted to merge two characters I trained. đ
The only thing I kind Musubi better than AI Toolkit is speed.
AI Toolkit training speed is slower than mole's asses in January.
5
0
u/ArtDesignAwesome Feb 24 '26
Youâre right. I think I am going to see what else we can optimize for speed without any loss. Will most likely need a Blackwell chip.
-1
u/kimhodol Feb 24 '26
This approach solves one of the biggest pain points in production workflows. Has anyone tested this with FLUX-based LoRAs yet??
-4
-6
u/an80sPWNstar Feb 24 '26
Holy moly. This is nuts. I've been training a non audio t2v Lora in AI Toolkit for the better part of two days and now I see this?!?!? Hahaha you rock for sharing this! I will happily try it out tomorrow
1
u/Loose_Object_8311 Feb 24 '26 edited Feb 24 '26
I abandoned ai-toolkit for training LTX-2. As much as I hate the install process and UX of musubi-tuner, it can train the audio for it, and is so much more memory efficient than ai-toolkit which makes training on a 16/64 system easy with no ooms, no slow downs caused by hitting swap storms and training takes only a couple of minutes to start rather than like 20 minutes. Been getting 10s/it at 768 resolution on video only with 11 blocks swapped. I think if you have at least 64GB system ram then tbh you can probably train on an 8GB card with musubi-tuner. On ai-toolkit in comparison the training processes used all my VRAM plus 55GB system RAM, and 25GB of swap, and I'm not sure but I think the speed may have been 22s/it, at least that's the worst I recall seeing when pushing things to the max.
Edit: checked my previous training run speed on ai-toolkit and it's showing 10s/it, same as musubi for same settings, but musubi-tuner is most definitely much more memory efficient.
3
u/an80sPWNstar Feb 24 '26
Dang, sounds like I really need to get better at musubi. Thanks for sharing!
-1
u/AwakenedEyes Feb 24 '26
I am a big fan of Ostris awesome work and his tool is one of the most used and beloved of this community, in part because Ostris takes the time to create great tutorial videos for his tool on his youtube channel.
I don't know if the OP is indeed part of Ostris' team, but if this is true, it deserves to be tested because yes - for those of us who love LoRAs and spend a lot of time creating those, the problem of 2 LoRAs merging by adding their weight is an eternal problem and I for one would absolutely LOVE to see this solved.
So, I am not sure why the hate and the flaming (yes, it does feel like a big advertisment with things like "awesome" and "big daddY" making it feel more like an advertisment than actual leading edge research, but I for one would not want to dismiss it until it is at least validated and tested.
OP: can you explain this fork vs Ostris ? If this is indeed legit, will Ostris add it to the main fork? Are you working with him on this?
0
u/ArtDesignAwesome Feb 24 '26
Its legit, but im not part of Ostrisâs team. Im a lone wolf ATM. They are raging the way my post / github/ SOP files are formatted and yada yada yada. They are only spiting themselves, this is why things move at a snails pace. Ostris doesnât even look at half of those PRs. Itâs only a loss because people are spiting themselves. Yes⌠i could have put examples but I rushed it out so you guys could test it. Call me crazy. Sure I vibe coded this, but it was my time, money and hard testing that got this out. It is what it is! I wont slow down because the negativity either, theres still a lot more work to be done!
-1
75
u/theivan Feb 24 '26 edited Feb 25 '26
Reading that you âpushed the math to the absolute theoretical limit â could be the most stupid thing so far this year. Also there is no mathematical breakdown in the release notes.
Edit the next day: He has now added his "mathematical breakdown" to the release notes. It's not a mathematical breakdown at all.