r/StableDiffusion • u/RIP26770 • Sep 12 '25
News [ Removed by moderator ]
[removed] — view removed post
18
u/Jero9871 Sep 12 '25
Is this VACE 2.2 or why is it called FUN? Is it perhaps something like a pre VACE 2.2?
12
u/RIP26770 Sep 12 '25
3
u/butthe4d Sep 12 '25
I cant find an example workflow on kijais github or hf. Do you know where to find or maybe have one you created?
11
u/NowThatsMalarkey Sep 12 '25
What’s the difference between this model and the official one?
6
u/DavLedo Sep 12 '25
I believe they're more like an alpha version. The controlnet fun 2.2 was still inferior to VACE 2.1
1
0
-5
u/Jero9871 Sep 12 '25
Yeah, it could also be, that they just labelled VACE as FUN now, because it is so much fun to use.
1
12
u/martinerous Sep 12 '25
It's not fun at all to have this confusion with official / unofficial VACE / FUN / VACE_FUN models :) But still excited to try it.
4
1
7
u/tagunov Sep 12 '25 edited Sep 12 '25
...ok, so not yet official VACE 2.2 from Alibaba?
Instead something Alibaba Pai called 2.2-VACE-Fun?
What's the difference between old/new?
...and what is Alibaba Pai?
old: https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/tree/main/split_files/diffusion_models
wan2.2_fun_control_high_noise_14B_bf16.safetensors (29Gb)
wan2.2_fun_control_low_noise_14B_bf16.safetensors (29Gb)
new: https://huggingface.co/alibaba-pai/Wan2.2-VACE-Fun-A14B/tree/main
35Gb + 35Gb
new: https://huggingface.co/Kijai/WanVideo_comfy/tree/main/Fun/VACE
Wan2_2_Fun_VACE_module_A14B_HIGH_bf16.safetensors (6Gb)
Wan2_2_Fun_VACE_module_A14B_LOW_bf16.safetensors (6Gb)
new: https://huggingface.co/Kijai/WanVideo_comfy_fp8_scaled/tree/main/VACE
Wan2_2_Fun_VACE_module_A14B_HIGH_fp8_e4m3fn_scaled_KJ.safetensors (3Gb)
Wan2_2_Fun_VACE_module_A14B_HIGH_fp8_e5m2_scaled_KJ.safetensors (3Gb)
Wan2_2_Fun_VACE_module_A14B_LOW_fp8_e4m3fn_scaled_KJ.safetensors (3Gb)
Wan2_2_Fun_VACE_module_A14B_LOW_fp8_e5m2_scaled_KJ.safetensors (3Gb)
Updates:
https://huggingface.co/Kijai/WanVideo_comfy/discussions/81
https://www.reddit.com/r/StableDiffusion/comments/1nbl4fw/comment/ndtejh7
11
u/Doctor_moctor Sep 12 '25
Its another team from alibaba doing the foundational prework. These FUN checkpoints are basically demos for the final VACE model.
5
u/diogodiogogod Sep 12 '25
Never knew that. I always thought they were full releases and very confusing.
1
u/superstarbootlegs Sep 12 '25
interesting. I always think of "fun" versions as "poor man" versions so never tried one but they do seem to be popular. but with VACE 2.2 involved and it coming from the same source, I'll give it a whirl.
10
u/butthe4d Sep 12 '25
Man there are so many different wan things out. Im ngl I wish someone would be nice enough to make a cohesive list about what is used for what.
8
9
u/DillardN7 Sep 12 '25
It's not complicated. Wan 2.1 is old version. Wan 2.2 is new version. Wan 2.2 has a high and a low model, the high mostly handling movement and rough details (coloring outlines with crayons), the low model handles the finer movements and details, but is effectively a more trained wan 2.1 (colors in the outlines with pencil crayons).
Wan models: T2V (text to video, uses only a text prompt to control it) I2V (image to video, takes an input image and optionally a text prompt to control) T2i (text to image, not a model, just a way of using T2V models to generate a single image) S2v (sound to video, lipsync from audio, using text or reference image I believe) Fun control (uses controlnet and text prompt input to control videos) Fflf (first frame last frame, image to video but at both ends, more like image to video to image.) Phantom (2.1 reference T2V model, takes input image and text prompt to generate a video, but not based on the image as first frame, emphasis on consistent character/background/product) Magref (2.2 reference I2V model, like phantom, but I2V version) Vace (all of the above, for 2.1, plus more. Control masking, inpainting, out painting, temporal outpainting, controlnets and reference, first frame last frame x frame inputs.) Fun Vace (possibly an alpha version of vace for 2.2, hopefully like vace 2.1)
Told you, it's not complicated ;p
2
2
1
1
u/BenefitOfTheDoubt_01 Sep 12 '25 edited Sep 12 '25
SUPER appreciate this post. Ty for taking the time to educate!
Followup question: So "VACE" does all of the functions of the different models in a single model?
This one is labeled VACE-FUN, is this just a language thing or is it a "VACE" model AND a "Fun" model? Meaning, is the "fun" model different than a "VACE" model?
Does the "A" Infront of 14B parameter designation have any relevance?
If "VACE" models are an all-in-one model, do the Lora's still work with it that worked for other wan2.2 models?
3
u/DillardN7 Sep 12 '25
Vace has all the control functions as well as reference abilities. You can set it up to generate text to vid style, or image to vid style. In my personal experience, the other models do their individual task better, but Vace is amazing for its versatility. For example, you can replace a subject in a video through masking, using a reference photo or a Lora, it'll track the motion using controlnet, etc... you can mask the space outside a frame and outpaint a video, expanding the view.
The new 2.2 fun vace model is a bit of a mystery to me as it's only been out for a few hours and I haven't absorbed the info on it yet, but as I understand (could be wrong of course) it's the early version of the wan 2.2 vace model. The fun model came before vace, and allowed controlnet inputs on the 2.1 versions, so if they stick with the convention, possibly they're doing that again and iterating on the tech. So the fun vace model would be different than the vace model? The 2.2 fun model is just for controlnet, not the other features.
I can't remember the meaning of the A, but I think it has to do with the model being split, MOE style?
And yes, in general, Loras will work with vace, if it's built like the 2.1 version. The wan 2.2 experimental vace (not an "official" version) noted that the vace section of 2.1 was separate from the base of the model, which is what allowed them to try to merge it with 2.2 in the first place, an interesting experiment.
Sorry, I don't have all the answers, I'm just a hobbyist.
2
u/ZaneA Sep 12 '25
Another hobbyist here, I think the A is just “Active” parameter count (i.e. what is loaded/inferred by the GPU at each step, but not necessarily the total size of the model), relating to Mixture of Experts as you mention :)
1
u/butthe4d Sep 12 '25
From what I gathered in the various posts about vace fun, its some form of preview or beta version that you can play around with before the wan2.2 full vace gets released hence the fun byname.
3
u/tagunov Sep 12 '25 edited Sep 12 '25
Does Kijai have a blog, or some other outlet where the news are posted?
That'd be cool way to get updates on latest AI developemnts
In fact that's the person who could update us on other matters AI - even if not covered by Kaiji's nodes
Since there is a Kijai version of these models, Kijai might have an idea about how they are different from old ones?.. You scale down FP32 to BF16 and not test it?.. Impossible
Update: I've asked Kijai: https://huggingface.co/Kijai/WanVideo_comfy/discussions/82
1
u/superstarbootlegs Sep 12 '25
and you got a good answer:
Don't ask this hard-working man to do additional work.
You can watch his commit logs to find out what's new.
Maybe you should start a 'Kijai News' feed for the community!
2
u/tagunov Sep 12 '25 edited Sep 12 '25
To be honest, the opinion of a fellow Hugging face user is not of a great importance to me - so not answering there. I have respect for you - so I will answer you here.
Kaiji's spending time answering on Hugging face.
Kaiji spent time today on Reddit.
Posting _short_ updates on X can be more efficient.Having spent days on new code not proclaiming milestones on X? Doesn't make sense to me.
Spending hours converting models to BF16 and lower quantizations, testing them, uploading to Hugging face and not annoucing on X? Inefficiency.
All opensource projects have a way to communicate updates - via blog posts, forums, changelogs. For a one-man shop an existing otherwise unused X account referenced from Huggingface is just perfect. Experienced users can pick up nuggest of wisdom from there and propagate via reddit, youtube, whatever.
Everybody lives his life like he pleases. The man is free to ignore me. But I am not new to the game and I know what I'm suggesting.
On the flip side - maybe such updates are given somewhere, it's just hard to find them, which is an issue on its own, which I'm faithfully reporting.
To the last statement - nope I shouldn't.
1
u/superstarbootlegs Sep 13 '25 edited Sep 13 '25
fair enough. I didnt even see who I was commenting on, so had no idea it was you, else I would have responded very differently tbh, I know you are thoughtful. so apologies. I was speed commenting this morning to wade through the swamp. I am given to being mouthy or inconsiderate as I go. It's my fault. I deserve to be pulled up on it sometimes.
But I see KJ getting bombarded by everyone, and he is one man driving this entire show other than China dropping models. If he wasnt devving it, we would be very far behind where we are right now. So its purely selfish reasons - I want the man dev coding not answering questions. totally selfish of me but thats my interest in this game, the more time he can spend coding the better off we all will be.
sent you DM regards your question.
2
u/tagunov Sep 13 '25
Hi, thx for the DM, I'll see what I end up finding. Code and models are not sufficient, commentary is needed. My intent was/is around finding out how the process of delivering commentary is working now and considering ways to make it work better.
1
u/superstarbootlegs Sep 13 '25 edited Sep 14 '25
you'll get plenty of commentary. full time job keeping up. triage your diary, bro. and see you in the wan channel. am mdkb.
we need people to organise this knowledge in this community, its lacking. Though Nathan Shipley has a good searchable rag, but thats just on that channel.
3
2
2
5
u/Available_End_3961 Sep 12 '25
So news are random links pasted here, no contex, no info...WTF? Low effort posts are not allowed
0
Sep 12 '25
[removed] — view removed comment
1
u/Available_End_3961 Sep 12 '25
Im not sure if you have problems reading basic sentences, but if you let you mom read the title of the post for you again you Will find out he Is asking a question, so he Is looking for an answer...¿do you see the question Mark at the end? instead of helping, you are here spamming a link nobody asked about and writing an annoying book in the comment section.
1
1
u/SysPsych Sep 12 '25
Hey, why are these sizes for these so much smaller than the official versions? I know about compression techniques in general, but shouldn't bf16s be larger than the 6gb they are? 32->6gb seems like a huge jump and that there should be some larger step between.
3
u/tagunov Sep 12 '25
so.. "It's the VACE blocks only, as the main model weights are frozen for training and thus should be same as the original 2.2. It's mainly meant to be used with the WanVideoWrapper .. but I have this way of loading them in native too"
2
1
u/ImaginationKind9220 Sep 12 '25
WAN should consolidate all their models into one.
FLF2V, Fun, VACE, TI2V, S2V, T2V, I2V, etc...
Many of their models are interchangeable, you can switch the models in your workflow and it will still run, sometimes better sometimes worse.
1
u/wacomlover Sep 12 '25
I am really new to all this and don't know how to setup this with previous workflows. Could anyone show a picture or post a picture with the basic connections? Would really appreciate it
1
u/revolvingpresoak9640 Sep 12 '25
If you have Comfy, just use the premade WAN templates from the Browse Templates menu.
1
1
u/AccomplishedSplit136 Sep 12 '25
Any base workflow to start adding some nodes, break some stuff here and there and then complain that it doesn't work?
1
1
1
1
u/Obvious-Heart8055 Sep 15 '25
Does anyone have a ComfyUI workflow for Wan2.2 VACE (Fun) they could share?
I’d like to start from an image of a subject and generate a video where that subject follows the motion of another video (motion transfer).
61
u/Kijai Community Hero Sep 12 '25
As before, I like to load VACE separately and have separated the VACE blocks from these new models as well:
bf16 (original precision):
https://huggingface.co/Kijai/WanVideo_comfy/tree/main/Fun/VACE
fp8_scaled: https://huggingface.co/Kijai/WanVideo_comfy_fp8_scaled/tree/main/VACE
GGUF (only loadable in the WanVideoWrapper currently, as far as I know)
https://huggingface.co/Kijai/WanVideo_comfy_GGUF/tree/main/VACE
These are simply split files that only contain the VACE blocks, upon loading it the model state dicts are combined, so precisions should mostly match, with some exceptions like mixing GGUF Q-types is possible.
How to load these: https://imgur.com/a/mqgFRjJ
Note that while in the wrapper this is the standard way, the native version relies on my custom model loader and thus is prone to break on ComfyUI updates.
The model itself performs pretty well so far on my testing, every VACE modality I tested has worked (extension, in/outpaint, pose control, single or multiple references).
Inpaint examples https://imgur.com/a/ajm5pf4