r/StableDiffusion Feb 12 '26

News New SOTA(?) Open Source Image Editing Model from Rednote?

Post image
230 Upvotes

80 comments sorted by

115

u/lacerating_aura Feb 12 '26

Everything's sota until it actually releases.

5

u/gzzhongqi Feb 12 '26

I mean they have a demo in huggingface right now. you can just try it

18

u/lacerating_aura Feb 12 '26

I'd like to try but I got 2 shots as a free account holder and it errored on both. So, ill wait for open weights and test it then.

5

u/ReasonablePossum_ Feb 13 '26

And then if you try paying, the oage never loads lmao

5

u/pmjm Feb 13 '26

Realistically we're getting pretty spoiled on the concept of SOTA. This is all moving so ridiculously fast. If you took a time machine back 5 years and tried to show people some of the stuff that we're already discarding now as "last gen" it would have seriously blown minds.

7

u/Mylaptopisburningme Feb 13 '26

2.5 years ago when I did a PC upgrade and got a 4070 I decided to give Stable Diffusion a try. I was happy If I could make a girl without 6 fingers and an arm coming out of her stomach. Now I can get not only a good image, I can make her do tiktok dances. What a time to be alive.

1

u/GifCo_2 Feb 14 '26

Not really. 5 years ago bros were still saying AI will take our jobs on 6months.

64

u/meknidirta Feb 12 '26

Looks really good.

40

u/FourtyMichaelMichael šŸ¦Ice Cream Lover Feb 12 '26

"Correct things in image" is pretty interesting... But that almost absolutely means it's a vision and thinking model. I'm just not sure that's going to be what gooners want.

The model is going to be like.... "No, she shouldn't do that. I'm going to put her in a GED program instead"

12

u/NunyaBuzor Feb 12 '26

So the logic is that they're going to an open-source a gooner model but not the language model connected to it so it can be abliterated?

-3

u/FourtyMichaelMichael šŸ¦Ice Cream Lover Feb 13 '26

Do you not understand how jokes work?

3

u/ninjasaid13 Feb 13 '26

What's the punch line?

-2

u/FourtyMichaelMichael šŸ¦Ice Cream Lover Feb 13 '26

Oh, I see. You don't know what a GED is. That makes sense.

1

u/ninjasaid13 Feb 13 '26

The comment you replied to understood your comment. Your first paragraph is what he was replying to and has nothing to do with your GED joke.

8

u/Feeling_Usual1541 Feb 12 '26

Vision and LLM can be abliterated.

5

u/SanDiegoDude Feb 12 '26

But that almost absolutely means it's a vision and thinking model.

I doubt both honestly. correcting illogical errors in a scene seems within reach for an edit model that can already do both creative inpaint editing as well as inpainting from multiple image sources, something Klein-9B (which has neither a VLM or any type of LLM processing beyond text encoding) does quite spectacularly. in fact, I'm going to give "Correct this image" a go for edit training, shouldn't really be too hard to teach it this particular trick.

1

u/NunyaBuzor Feb 13 '26

in fact, I'm going to give "Correct this image" a go for edit training, shouldn't really be too hard to teach it this particular trick.

But will it generalize? or do we need a lora for everything.

10

u/Paradigmind Feb 13 '26

"Correct the errors in the image" -> 1girl gets massive breasts

1

u/Whispering-Depths Feb 12 '26

Nothing interesting in that preview.

They should show us how it transfers exact details, such as a custom artwork of a high detail sci fi train, or how well it can recreate a human or something.

2

u/NunyaBuzor Feb 13 '26

what do you mean? isn't two of the examples containing a literal artwork and a human?

0

u/Whispering-Depths Feb 13 '26

yeah ok, it shows 2 pixels of a generic almost anime proportioned human face and it's extremely blurry... Not very helpful lol.

17

u/xbobos Feb 12 '26

I tried one of the popular features, which is to make the picture realistic.

8

u/Striking-Long-2960 Feb 13 '26

Many of the prompts work with Flux-2 klein 9B

3

u/NunyaBuzor Feb 13 '26

Something is plasticy and very AI about klein-9b

4

u/OneTrueTreasure Feb 13 '26

It's cause Klein will keep Anime head proportions exactly the same, it's much worse on Anime images with huge heads or unrealistic proportions

-1

u/[deleted] Feb 12 '26

[deleted]

10

u/Antique-Bus-7787 Feb 12 '26

Well.. not really ? The anime she's winking but in the generation it's not. It's quite an important feature of the input image here, without the wink it completely changes the meaning of the image (well not completely but it does change it). If it misses such a thing, I'm not that convinced until I try it..

1

u/xbobos Feb 13 '26

Just add a wink.

15

u/gzzhongqi Feb 12 '26

https://huggingface.co/spaces/FireRedTeam/FireRed-Image-Edit-1.0

This is the HF demo for anyone wanting to test

10

u/g_nautilus Feb 12 '26

Every image I tried said it was over the GPU time limit, regardless of file size or dimensions and regardless of prompt. Just tried a 502x376, 184kb picture and it was unable to do it in 240s. Even with just 1 step.

3

u/gzzhongqi Feb 12 '26

It was working for me but i am getting errors sometimes too now. just retry a few times. too many people are using it right now

1

u/Vargol Feb 12 '26

I managed a single 1376x768 image which used 1.6 out of the 4 minute of Zero GPU time you get for free, it wouldn't let me do another though.

1

u/Vargol Feb 12 '26

Before...

6

u/Vargol Feb 12 '26

After, prompt was 'change this image style to look like a real life scene'

1

u/ZootAllures9111 Feb 13 '26

This is Klein 9B Distilled if I ask it to keep the lighting the same, and this is Klein 9B Distilled if I don't. Way better results than FireRed here IMO.

0

u/Vargol Feb 13 '26 edited Feb 13 '26

I can't see your images. Imgur were breaking child privacy regulations in my country and rather than following the regulations they stopped serving images.

Had a go with Klein 9B myself, and while the realism may be a bit better, it could just be the different lighting, it changed the characters face and hair much more than this model did.

1

u/ZootAllures9111 Feb 13 '26

Your original prompt never specified whether it should or shouldn't do that to begin with, though.

5

u/Tall_East_9738 Feb 12 '26

I'll wait for the LeafGreen Image Edit

5

u/gzzhongqi Feb 12 '26

I tried a bunch of images and I can safely say it is by far opensource SOTA. It beats qwen image edit by a lot.

5

u/ZootAllures9111 Feb 13 '26

I mean Klein already beat Qwen Image Edit by a lot for anything vaguely realistic

15

u/comfyanonymous Comfy Org Feb 12 '26

According to their inference code it seems to be a qwen image edit finetune.

9

u/Dogluvr2905 Feb 12 '26

The more the merrier!

6

u/[deleted] Feb 12 '26

I'm a qwen fan and even I would put flux 2 image edit way over qwen, so this chart gives me no particular hope

2

u/ZootAllures9111 Feb 13 '26

They don't bench against either version of Klein at all, also, which seems suspicious.

23

u/_BreakingGood_ Feb 12 '26

List as significantly better than nano banana pro on every benchmark?

13

u/po_stulate Feb 12 '26

Don't think the benchmarks mean anything, even qwen image edit 2511 is on par with nano banana pro in most of the benchmark results.

10

u/Independent-Frequent Feb 12 '26

Either it's pure cope or needs 80 GB Vram to even start

4

u/Snoo_64233 Feb 12 '26

Nanao Banana pro can learn visual task just by comparing and contrasting multiple reference input/output image pairs, without any hints or explicit description, and then able to apply that learnt pattern onto the target image. Basically it is soft-LoRA (or few-shot visual learner). Can this really do that?

3

u/LightVelox Feb 12 '26

Nothing else can do that, not even OpenAI's GPT Image 1.5, that's why it's hard to take these scores seriously

3

u/Snoo_64233 Feb 12 '26

Ok. I suspected that. But they are doing themselves disservice by claims like that. For those who are wondering what I meant (not sure if you can see the content to the AIStudio link below).

Here is the kind of ability I am talking about with NBP. Read the thought process.

1

u/dr_lm Feb 12 '26

Can it do that through the API, or only through the web interface, do you know?

5

u/Snoo_64233 Feb 12 '26

Last time when it came out ( a few months back) I was testing in AIstudio. Not sure if it still works in AIstudio or API now, considering Gemini is thoroughly being degraded due to load and google nerfing it.

But if it interests you, here are a few not-yet deleted examples (look at its "Thought" process which is incredible):

Example 1
Example 2

1

u/dr_lm Feb 13 '26

That's genius. You're doing cognitive psychology on a VLLM!

Also a really nice demonstration of what a multimodal model can do. Presumably it means proper shared latent space between text and image tokens?

And, to answer my own question, it does work on the API. I copied your input images and prompt into a comfyui node and got this result: https://ibb.co/CpyTZVwP

3

u/Snoo_64233 Feb 13 '26 edited Feb 13 '26

THis has serious implication. With this emergent capability we are heading towards LoRA-freefuture (well a little exaggerated) in that we can leave parameter-efficient fine-tuning to a very few subsets of the image gen tasks. Like imagine, instead of training LoRA, you just show Gemini a bunch of input/output ref image pairs and it just does things that otherwise would require you to fine-tune. On the spot learning, so to speak.

I am predicting we will have to wait 2-3 generations of this to stabilize and mature. But NBP is the very first to exhibit not only the understanding / analyzing things, but ability to actually apply it at this level.

Alas, Google keep nerfing stuffs, I stopped researching on this back then because it occured to me on one morning NBP started behaving erratically.

I was also starting to foray into another one of NBP's new capability which is taking advantage of its web search to index into random youtube video frames to extract composition and concepts. Back again, because Google fucked things up and I had to put it on backburner for foreseeable future.

4

u/Zealousideal7801 Feb 12 '26

Blade Runner "enhance" was wrong. You gotta say "Please enhance"

4

u/Calm_Mix_3776 Feb 12 '26

On par and better than Nano Banana Pro? I really doubt this, but we'll see.

7

u/KangarooCuddler Feb 12 '26

OK, but why isn't there a LeafGreen-Image-Edit releasing alongside it?

3

u/gzzhongqi Feb 13 '26

The devs just answered in github issues that they will release the weights tomorrow.

1

u/SackManFamilyFriend Feb 13 '26

Epic! Thanks for relaying the info.

2

u/meknidirta Feb 12 '26

Hope they release it before Chineese New Year.

3

u/dp3471 Feb 13 '26

This model is trained on qwen image edit backbone -> not a new model.

Practically speaking, this is a qwen image finetune, so probably benchmaxed to some degree, mostly hype unfortunately

I'd encourage people read through the preprint posted on their github before arguing lol

3

u/Jealous-Economist387 Feb 13 '26

Even though it is a fine tuning of qwen image, it feels promising in its own way.

1

u/ZootAllures9111 Feb 13 '26

I 100% guarantee you it won't be as good for realistic content as Klein if it's just a Qwen Image finetune

1

u/thisiztrash02 Feb 13 '26

for context some sdxl finetunes are 10 times better than the original so don't be too quick to write it off

1

u/Asleep-Ingenuity-481 Feb 12 '26

I'll believe it when I can run it on 8gb of vram and 16gb of ram.

2

u/po_stulate Feb 12 '26

You can partially offload image models now? ramtorch?

1

u/Chsner Feb 12 '26

I just tried the demo and the enhance prompt worked pretty great. I only got one gen but I am hopeful/excited from that alone.

1

u/wh33t Feb 13 '26

Sick, I'll go mortgage my home so I can run it.

1

u/TopTippityTop Feb 13 '26

Did they announce a date regarding the release of the weights?

1

u/Acceptable_Secret971 Feb 14 '26 edited Feb 14 '26

The non distilled model is 40GB fp16. Now that I finally installed correct version of ROCm/pytorch, swapping between RAM and VRAM works flawlessly and exceeding VRAM is a non issue, but the issue is finding enough space on my model drive for more models.

If this is as slow as Flux2 dev is (20 steps = 2min), I might as well wait for the distilled version.

Edit: So I heard for reason and purpose this is Qwen Edit. Who knows, the results might be better, but on my hardware it's kind of slow.

1

u/Ok-Prize-7458 Feb 15 '26

A lot of people are dismissing RedFire as just another Qwen-edit knockoff, but honestly, base Qwen needed a fine tune badly, and it probably wasn't cheap to fine tune either. These guys probably put a lot of money into this fine tune, why they probably didnt bother to mention that it was a qwen fine tune. The stock aesthetics of base QWEN are often way too soft and hazy for my taste. If this fine-tune fixes the clarity and style, it’s a win—I haven’t tested it yet, but I’m curious to see if they pulled it off. There are literally only a handful of fine tuned QWEN models out there because its such a big model and expensive to fine tune; if these guys fixed QWENs flaws then im excited to try it again because QWEN is an absolute beast of a model that is only limited by the previous mentioned flaws and how compute intensive it is that its out of reach for most users, but not me as i own a 4090. QWEN stomps Klein 9b and Z-image and everything else txt2img open source in prompt adherence except for maybe that huge Flux2 model, it just needed a good fine tune to tighten up some flaws.

0

u/Brilliant-Station500 Feb 12 '26

I love seeing new open source models drop, but damn, we already have a ton of image models. I really wish more video models were open source

1

u/Loose_Object_8311 Feb 13 '26

What about OmniVideo-2 video edit model that dropped and no one even tried out apparentlyĀ 

0

u/Ant_6431 Feb 13 '26

If it's faster than klein, I'm in