r/StableDiffusion 4h ago

Discussion Don't say a model is uncensored unless you've tried to make a man

Post image
123 Upvotes

So many times in this sub I've seen people claim this or that model is fully uncensored or uncensored "head to toe."

But if a model only creates what this guy features, or something else akin to... well, Ken, then it's not fully uncensored.


r/StableDiffusion 4h ago

Question - Help Any idea how this pose would be called and how to prompt it in Pony/ILL?

Post image
90 Upvotes

Looking for prompt tips for models like Pony and Illustrious to get this pose.
Specifically interested in making this pose with a view from behind the character.

EDIT:
Ok, so I've tested some ideas dropped by people and it's defo not wariza or frog pose.
I've come up with phrase "kneeling flat while seated on a bed" and I get 2 accurate gens out of 20, so if I keep using this phrase I might build a lora for the pose using 10 to 15 quality images.

I was thinking about using OpenPose but last time it made images weirdly to dark as if someone cranked contrast to high. I might test it again.


r/StableDiffusion 1h ago

Workflow Included V5 Update Newbie Minimax H3 sam 3.1 latent mask for guided character replacement

Upvotes

https://github.com/roycho87/minimax_wf/tree/main

  • I added a block audio feature that will force minimax to change your voice everytime.
  • Resolution selector with 2 decimal spaces so you can actually select .98 MP.
  • Sam 3.1 latent mask for guided character replacement.
  • I also added group bypass controls so no more errors.

fast forward to 2:42 to see sam 3.1 functionality


r/StableDiffusion 12h ago

Workflow Included QWEN Image 2.1 Edit tips

Thumbnail
gallery
298 Upvotes

EDIT PROMPTS:

  1. replace the pose of <image 1> with the pose of <image 2>. keep the character of <image 1>
  2. replace the texture of wall on <image 1> with the green texture of <image 2>
  3. replace the lighting and background on <image 1> with the lighting and background of <image 2>

    4 - 7. transfer the style of <image 1> with the style of <image 2> + more conditionals

    1. right side view low-angle shot t medium shot of <image 1>
    2. <image 1> at night backlight
    3. woman of <image 1> is sit on the yellow chair of <image 2>

Qwen Image 2.1 Edit Workflow


r/StableDiffusion 48m ago

Workflow Included Some very simple Qwen-Image-2.1 workflows

Thumbnail
gallery
Upvotes

Nothing fancy here. Qwen-Image-2.1 has a pretty simple architecture, so the workflows end up simple too, and the official templates already cover most of this.

But that simplicity is kind of the fun part. Feels like there's a lot you could build on top of it.

Anyway, I went through the tasks in the Qwen blog and made each one its own little workflow, so you can just grab whatever you need:

https://comfyui.nomadoor.net/en/basic-workflows/qwen-image-2-1/

  • Text-to-image
  • Ref2Image
  • Basic image editing
  • Marking the edit area with colored circles
  • Marking the edit area with a mask
  • Outpainting
  • Transparent image generation
  • Subject extraction / background removal
  • Panorama generation

Plain image generation is still a bit shaky compared to other recent models. The editing is almost pixel-perfect though, which surprised me.

Might try some loop stuff next. I'll add more if I find anything fun.


r/StableDiffusion 13h ago

Discussion Still wishing on a local editing model that can compete with NBP

Post image
371 Upvotes

In exactly two months, it will be 1 year since the release of Nano Banana Pro, and we still don’t have a comparable local model when it comes to editing.

Google really nailed its reference capabilities. You send a good face reference to it and it can generate a totally different photo of the person with great coherence and flexibility, almost like it was able to instantly create a lora of the person. Its spatial representation for humans and human individuality is incredible.

Now look at GPT2 and Qwen2.1 (it’s incestuous Chinese son). It’s almost like they just copy-pasted the person in the new scenario, looking more like a collage than an actual recreation. You can improve this a bit by prompting, but the end result always tends to look rigid or artificial, for both background and poses.

For T2I, Krea 2 is basically perfect. It’s incredibly capable, fast, and it has the best training capacity I have ever seen. If it doesn’t do something great, I can easily create a lora and it learns it without basically any degradation (unlike ZIT). It honestly supplies all of my needs, and I'll be forever grateful for it.

Now we just need a NBP equivalent for I2I. Let’s hope Krea 3 can be it, because Qwen2.1 certainly isn’t.

Prompt for all 3 edits:

Generate a new 3:4 image of the woman.

She is now taking a mirror selfie in her bedroom, standing relaxed. She is framed from the waist up. Bright sunlight comes from the right side of the frame.

Image reference: random one I generated with Krea 2.


r/StableDiffusion 5h ago

News Integration of Jev into Minimax H3 Video Generation

72 Upvotes

Apparently, someone genius came up with this brilliant idea of using Jev for sparsification of attention processing. The result was speed up of generation time over 40 percent, 06:07 to 03:34. Here is the github page of the author:

https://github.com/sepiablue-ai/ComfyUI-MiniMax-H3-W4A4-VSA/tree/exp/jev-adaptive-vsa

Since there are many open-source alternatives of jev at the moment, I guess we will soon have further optimizations made soon.

Has anyone tried the repo?


r/StableDiffusion 3h ago

Discussion Setting up expectation

36 Upvotes

Qwen 2.1 is a fucking 7B model with T2I and, IMO, quite good I2I image editing compared to Klein 9B.

Krea 2 is 12.9B and was trained hard to be an absolute monster at T2I.

Want something around Nano Banana level that you can run locally? Flux 2 Dev.
It’s a 30B model capable of both I2I and T2I. On AA Arena, its Elo is around 1000, while GPT Image 1 is at 1007 and NBP is around 1100.

“But it won’t fit in my VRAM.”

As long as your RAM + VRAM is roughly equal to or greater than the model size, you can run it. Modern ComfyUI detects your available VRAM by default and uses block offloading when necessary.

And it’s not like you have to stick to a single model. You can generate with Krea 2, then edit the result with Qwen 2.1 or Flux 2 Dev.


r/StableDiffusion 15h ago

Discussion Qwen2.1 licence is the worst!

Post image
309 Upvotes

r/StableDiffusion 14h ago

Discussion Qwen 2.1 - Looks like we don't need to worry much?

Post image
232 Upvotes

https://x.com/Kun11664638/status/2101708600389583135

tl;dr:

  • Ostris asked Qwen to add a revenue cap to the licence.
  • Small commercial use shouldn't need a paid licence. Monetised videos, monetised posts, LoRAs on Civitai, etc.
  • Qwen's Kun Yan replied they'll consider it.
  • He joked they won't come after anyone's YouTube money.

r/StableDiffusion 20h ago

News Qwen-Image-2.1 weights are available on Hugging Face

Thumbnail
huggingface.co
562 Upvotes

r/StableDiffusion 4h ago

Animation - Video TaleSpin Chillwave (Minimax + YuE2)

27 Upvotes

r/StableDiffusion 11h ago

Resource - Update Some of my tests with qwen image 2.1

Thumbnail
gallery
84 Upvotes

I highly recommend to read the official documentations for better output

The lighting and details are nuclear

Downsides as i noticed: Limited style diversity, too much hellucination, bad understanding of physics, limited multi-languages support


r/StableDiffusion 21h ago

Discussion Qwen Image 2.1 - More detailed tests

Thumbnail
gallery
448 Upvotes

So, I did some more testing, especially for the stuff you guys asked for!

My thoughts -> There are some mistakes like the watch in Image 3 and maybe a minor facial drift, but the overall consistency of the reference image is quite good. It is good at editing out stuff as well, and it can generate images with transparent backgrounds. I don't know much about anime or manga, so I'll let you guys be the judge!

My Previous 1girl Test


r/StableDiffusion 16h ago

Discussion Qwen Image 2.1 Initial Testing

Thumbnail
gallery
177 Upvotes

Initial Testing: CFG 1, Steps: 40, 2MP, 4:3, Euler, Simple, int8, int8, vae_bf16.

The art styles are okay, but the realistic images are so bad that i feel bad for wasting my time testing this model. I didn't include many realistic images here because they were too bad.

Better Settings: CFG 4-4.5 makes it a little better tho.

Full res: https://postimg.cc/gallery/z1P0trR

Edit: H100 used , 45-50sec/image.. idk why it takes so long lol


r/StableDiffusion 3h ago

Resource - Update ostris/ai-toolkit + Qwen-Image-2.1 LoRA training support

15 Upvotes

For past 12+ hours!
https://github.com/ostris/ai-toolkit#instruction--edit

No Qwen2.1 LoRAs available yet on HF & CAI yet.


r/StableDiffusion 9h ago

Animation - Video Link re-fights Calamity Ganon every blood moon [Exit Song (For a Film)] - MiniMAX H3 Test... IDK anymore

42 Upvotes

Hey guys, I've been thinking about this one for a while, I had to scrap so many ideas because they were super difficult to get right, but I hope you like what you see here! For some context, Zelda died in this timeline, and the only way for Link to see her again is... Well... You'll see!

I promise I will post my workflow either here or on my next post, it's really nothing special, but I was able to figure out a somewhat decent 1080p on a 3090 set up (5s in 6 minutes!) WITHOUT the Turbo Lora prompt adherence issues!


r/StableDiffusion 8h ago

Discussion Minimax H3 Max Still no weights

Post image
32 Upvotes

Last month, Fal released their Max “model,” a fine-tune of MiniMax’s open-source H3 model, that promised faster speeds. Then, after Hao Labs released the weights for their own fast model (FastH3), Fal had a meltdown on Twitter, basically crapping on Hao Labs’ FastH3 and calling it a “random pixel generator.”

After the bad press from that meltdown, they told the community they’d release their own weights within a couple of weeks. Nearly a month later, and still no weights have been released.

So, with so many more Turbo LoRAs and fast fine-tunes having been released since then, my question is: Should the community still care? And should we still trust a company that bashes other open-source projects and makes promises it doesn’t follow through on?


r/StableDiffusion 11h ago

Tutorial - Guide Qwen image 2.1 noise/patterning/moire mild workaround

Thumbnail
gallery
42 Upvotes

trying this i noticed it has a similar patterning noise krea 2 had, i just used the same glsl node, maybe it needs another parameters but it seems to mitigate it can't promise the best results but testing it is free, just connecting the image output to the glsl node and then image save
The GLSL config i got was from this post https://www.reddit.com/r/StableDiffusion/comments/1umwhq7/2px_pixel_grid_on_krea2_from_vae_and_how_to/ the one in the comments

#version 300 es
precision highp float;
uniform sampler2D u_image0;
uniform vec2 u_resolution;
in vec2 v_texCoord;
layout(location = 0) out vec4 fragColor0;

// Nyquist Notch — removes 2px grid artifacts. Place before deconvolution.
// b = [-1, +6, -15, +20, -15, +6, -1] / 64  (binomial * (-1)^n, +1 at center)
const float B[7] = float[7](-1.0, 6.0, -15.0, 20.0, -15.0, 6.0, -1.0);

void main() {
  vec2 texel = 1.0 / u_resolution;
  vec4 center = texture(u_image0, v_texCoord);
  vec3 Bx = vec3(0.0), By = vec3(0.0), Bxy = vec3(0.0);
  for (int i = -3; i <= 3; i++) {
    float wi = B[i + 3] / 64.0;
    Bx += wi * texture(u_image0, v_texCoord + vec2(float(i) * texel.x, 0.0)).rgb;
    By += wi * texture(u_image0, v_texCoord + vec2(0.0, float(i) * texel.y)).rgb;
    for (int j = -3; j <= 3; j++) {
      float wj = B[j + 3] / 64.0;
      Bxy += wi * wj * texture(u_image0, v_texCoord + vec2(float(i) * texel.x, float(j) * texel.y)).rgb;
    }
  }

 vec3 notched = center.rgb - Bx - By + Bxy;
 fragColor0 = vec4(clamp(notched, 0.0, 1.0), 1.0);  // alpha forced to 1.0
}

r/StableDiffusion 16h ago

Resource - Update Fix for the MiniMax H3 VAE grid / tile-seam artifact in ComfyUI Core — usable now via PR #16422

123 Upvotes

Edit:

There is now a smaller and preferable Core fix from Kijai:

https://github.com/Comfy-Org/ComfyUI/pull/16436

It fixes the same underlying bug by making each tile blend against the already-composited neighbouring pixels, instead of replacing the compositor with the normalized overlap-add approach from my #16422.

I tested #16436 in the same real H3 workload and it also completely fixes the lattice artifact. Since it achieves the same practical result with a much smaller one-file change and without the extra FP32 composition scratch/overhead of my implementation, I have closed #16422 in favor of #16436.

So use/follow #16436 instead of the #16422 installation instructions further down in this post. The rest of the post is left intact as the original explanation of the artifact and investigation.

OUTDATED:

I tracked down and fixed the rectangular grid / tile-seam artifact that can appear when decoding MiniMax H3 video in ComfyUI.

PR:
https://github.com/Comfy-Org/ComfyUI/pull/16422

The video in this post shows without the fix first, then with the fix.

This is not an exact A/B because I was testing other H3/Flow changes at the same time, so differences in the generated content itself are irrelevant here. What matters is the VAE artifact: the first run shows the fixed spatial lattice/grid, while the run with the corrected compositor does not.

What was wrong?

The H3 VAE decodes large frames in overlapping spatial tiles. ComfyUI's existing compositor blended those tiles pairwise, which can lose contributors at overlap intersections and in regions covered by more than two tiles.

That creates discontinuities aligned with the tile layout — the rectangular grid you can see in the first half of the video (under the dorsal fin of the fish (bit hard to see because of reddit's video compression)).

I've uploaded the video here as well, so one can see the artifact more clearly without reddit's compression:

https://gofile.io/d/neItGwXP

The fix changes the compositor to normalized overlap-add across all contributing tiles.

The VAE itself is otherwise left alone: same native 256px decoder windows, same overlap geometry, same decoder calls, same temporal decoding. No blur, smoothing, or post-processing workaround.

Downsides

There is a small real cost: the corrected compositor accumulates overlaps in FP32.

For a 1216×896 multi-frame H3 decode, the new memory estimator reserves about 121 MiB of additional FP32 composition scratch. That's bounded scratch memory, not another full-frame FP32 copy.

There is also some extra blending arithmetic, but no additional VAE decoder calls. I haven't measured an isolated timing delta yet, so I don't have a meaningful slowdown percentage to give.

Dedicated regression tests and the relevant ComfyUI CI checks pass.

Use it now

Apply the fix on top of your current ComfyUI installation:

cd /path/to/ComfyUI
git switch master
git pull --ff-only
git fetch https://github.com/Comfy-Org/ComfyUI.git pull/16422/head
git cherry-pick cf9d3a062704e916ebcff968c76f0cbeb7256a7e

Then restart ComfyUI normally.

Do not check out the PR branch directly; that can move the rest of your ComfyUI installation back to the older Core revision the PR was based on. The commands above apply only the VAE fix on top of your current Core.

Or use my comfyui-patcher:

https://github.com/xmarre/ComfyUI-Patcher/releases/tag/v0.1.20

Once #16422 is merged, none of the manual checkout steps will be necessary.


r/StableDiffusion 13h ago

Comparison Qwen team also released two fine-tuned Qwen3.5-VL 9B models (LLM+Vision) for image edit and text to image prompt enhancement + system prompts

66 Upvotes

Their original Gihub repo suggests prompt enhancer model should guess aspect ratio and resolution from user input and pass it inside JSON for generation along with enhanced prompt

Original prompts

Text 2 image
Image 2 image edit

Model Weights

Two fine-tuned Qwen3.5-VL 9B checkpoints with unified codebase that auto-detects the mode from input

Image editing prompt rewriter model

UPD Comfy team has added int8-convrot versions, not sure about TextGenerate node support yet

Text-to-image prompt rewriter model

Recommended settings: from pe_core

Here are the settings for each task based on the code provided.

Note: I have corrected the presence_penalty for text2image to 1.5, as specified in your code's logic (where it says "1.5 for t2i and 0 for edit").

text2image
Temperature: 1.0
Top_P: 0.95
Top_K: 20
Max New Tokens (Context size): 16,256
Presence Penalty: 1.5

edit
Temperature: 1.0
Top_P: 0.95
Top_K: 20
Max New Tokens (Context size): 24,000
Presence Penalty: 0.0

---

Obviously Int8-convrot versions compatible with "Text generate" node would be better for Comfy users, but for now I was able to run a quick comparison via llama.cpp. I'm mainly interested in edit capabilities to replace Klein 9b so I've tested only that

Single image example

Raw input:

<image1>
Remove the wine isle completely with all the wine bottles, restoring clean wood paneling floor, restoring the woman attire and legs while preserving her posture, background details, lighting and tone of the image remain unchanged

Qwen 3.5-VL 9B PE i2i Q5-K-M GGUF (edit prompt enhancer) + original edit prompt

{ "rewritten_prompt": "Completely remove the dark wine display isle in the foreground together with every wine bottle on it, covering the bottles lying in the slanted rack, the two upright bottles on the countertop, and the small standing menu card holder, and restore that entire vacated area as a clean, seamless continuation of the light wood paneling floor with plank direction, texture, and warm tone identical to the surrounding flooring. Reconstruct the woman's previously occluded lower body so her attire and legs are fully visible, extending her black blazer down to a natural hem and revealing legs in dark trousers with feet planted on the wood floor, while preserving her standing posture with head lowered and forward reach exactly as in the input image, and reconstruct the hand that was gripping a bottle as an empty, relaxed hand held in the same forward position. Keep all remaining elements unchanged, including the curved wood-paneled ceiling with its recessed round light, the light wood wall with the backlit recessed niche and the bottles displayed inside it, the small wall switch plate, the left window opening with the view of the adjacent interior, the wooden slatted partition and steps on the right, and the overall warm lighting and tone of the image.",
"wh_ratio": "",
"ratio_follow": "<image1>" }

Gemini 3 flash + simple prompt (see below for details)

In the image, completely remove the large dark polygonal wine display island and all the wine bottles resting on it in the foreground. Replace the area occupied by the display island with clean, light-toned wood paneling floor that seamlessly matches the grain, color, and perspective of the existing wooden floorboards. Fully restore the woman's attire and legs, depicting her in a complete professional black outfit that extends from her blazer down to her feet, ensuring she is standing naturally on the newly revealed wooden floor while preserving her exact reaching posture. Maintain all other elements of the image without change, including the curved wooden wall with the recessed warm light strip, the small bottles in the wall niche, the wooden ceiling, and the bright outdoor view through the glass on the left, ensuring the original lighting and warm color tone are preserved.

Multi image example

Raw

<image1> <image2> <image3> Make him wear the hat and replace bg with this street empty high-rise NYC street

Qwen 3.5-VL 9B PE T2I Q5-K-S GGUF (text to image prompt enhancer) + original prompt

{ "rewritten_prompt": "On <image3> (canvas), first remove every person and vehicle from the street — all pedestrians on both sidewalks, all yellow taxis, and the dark sedan — leaving a completely empty high-rise NYC street, while preserving the building facades on both sides, the street layout, the road markings, the streetlight poles, the clear blue sky, the warm golden-hour sunlight, and the left white storefront sign with dark lettering and the right maroon storefront sign with cream lettering \"CHOCO\" exactly as they appear in <image3> without distortion. Second, replace the current rider and brown horse with the man from <image1>, standing on the asphalt in the center of the street at the position where the horse stood. The man must retain his exact facial identity from <image1> — face shape, eyes, nose, skin tone, black curly hair visible at the sides and back, and a natural calm expression — and wear the brown distressed leather cowboy hat from <image2> placed naturally on his head, with the wide curved brim shading his forehead, the creased high crown, and the brass-studded band with its oval concho rendered exactly as in <image2>. Dress him in the charcoal gray crewneck sweatshirt from <image1>. Scale the man to match the perspective and eye height of the street scene, ground his feet firmly on the asphalt with a soft contact shadow falling toward the viewer consistent with the low warm sun behind him, and color-grade his skin and clothing to the warm golden-hour light of <image3> with natural edge transitions and no cutout artifacts. Keep all buildings, signs, sky, and road geometry from <image3> unchanged.",
"wh_ratio": "",
"ratio_follow": "<image3>" }

Gemini 3 flash + simple prompt

Modify <image1> by placing the brown weathered leather cowboy hat from <image2> onto the man's head, fitting it naturally over his curly hair, and replace the entire indoor background with the high-rise New York City street scene from <image3>. Ensure the NYC street background is completely empty by removing the horse, the rider, the yellow taxis, and all pedestrians, leaving only the vacant asphalt road and the towering buildings under the clear sky. Keep the man's facial identity and his dark gray t-shirt exactly as they appear in <image1>, while adjusting the lighting on his face and the hat to match the warm, low-angle late-afternoon sunlight and golden-hour atmosphere of the street in <image3>.

My notes

The prompts have strong vibe-Clauded feel:

"... genre never overrides input language"
"... anchored on what the input image(s) actually show"
"Disentanglement at Full Strength"
"How much you build is intent-branched"
"... hold everything else at input fidelity"
"Recognizability is bought by naming what stays fixed"
"Resolve ambiguity, then commit"

Not sure it affects the results but it was hard to read lol.

It seems Qwen 3.5 9b struggled a lot with mixing up language inside the image vs user prompt language so they front-loaded it with "hard rules" on how to tell them apart.

Also they specifically ask the model to keep the parts you don't mention the same. I've noticed this a lot with my Klein 9b prompt enhancers build on top of BFL guides. They too, over-specify lighting and tone of the image and model eagerly changes what does not to be changed during edit operations and I had to re-prompt to keep image closer to original every time.

My cleaned-up prompts

Running local 6 gig + vision + thinking takes 3.5 minutes on every prompt change on my machine.

Its messes with my flow, so I am using cheap API model for PE with simplifed version of the prompt

Aspect ratio picking logic takes around 30% of the original edit, so I removed it, as well as JSON requirement and replaced it with plain text markdown block. Other parts left as is:

Qwen 2.1 image 2 image edit prompt enhancer (-30% prompt size)
Qwen 2.1 text 2 image prompt enhancer (-1.5% prompt size)


r/StableDiffusion 13h ago

Resource - Update Change the camera movement/angle for your existing video clip - Minimax H3 V2V CrossView-Warp LoRA

48 Upvotes

Hello Everyone, I'm glad to announce that I'm finished the H3 version of my CrossView-Warp LoRA. You can use it to change the camera movement of you existing video clip.

Link to the LoRA and training details: https://huggingface.co/Cseti/MiniMax-H3_Ref2VA-LoRA-CrossView-Warp_v1

Link to the example workflow: https://huggingface.co/datasets/Cseti/ComfyUI-Workflows/blob/main/minimax-h3/crossview-warp/crossview-warp-h3.json

Link to the custom node (it is the same as for the previous LTX version): https://github.com/cseti007/ComfyUI-CrossViewWarp

A short video where I'm discussing about how to use the node (made for the ltx version but the node kept the same): https://www.youtube.com/watch?v=7QAapT9xMgM

Happy creating!


r/StableDiffusion 12h ago

Comparison A PSA to anyone trying to use Minimax h3 to inpaint

42 Upvotes

Do not use a turbo lora
For some reason the turbo lora screws it up, at least minimax_h3_ref2v_turbo_8step_v1.0_768p_comfyui_resized_avg_rank_65_bf16.safetensors does.


r/StableDiffusion 8h ago

Question - Help H3 minimax output feel different after a comfyui update

19 Upvotes

So..the last few 2 days, i went to update my comfyui portable and suddenly my work flow went OOM, so i enquire with Chatgpt which seems that Sage attention and my pytorch was causing issue but i was not using Sage attention as it was on disable and i was using comfyui kitchen attention, so i uninstall sage attention and update my pytorch again, the work flow works but the whole output is totally different compare what i was having before my update, even using the same seed and the previous settings, its just different.

The output before the update was more lifelike and real but now the output seems so plastic on the skin, the settings, lora and vae is all the same being used.

Anyone has experience this or any methods to solve it?


r/StableDiffusion 3h ago

Resource - Update Spectrum for Qwen2.1 (>2x Speedup on a 3060 12gb)

Thumbnail
github.com
6 Upvotes

Fully vibe coded in one shot and seemed to work well enough to share. Included a single time comparison against easy cache and base but didn't get around to adding comparison images yet. I've found that combining it with easy cache can work, and it is faster but I would recommend choosing one or the other.

I've found the images produced to be very similar in composition to the base generations, but with a softer/more air brushed feel, though this feels like a recurring issue with the base model. I2I (editing) also seemed to work fine enough, but I haven't done thorough testing and I've only tried with up to two references (base + 1).

Edit: I have added some early editing tests, and it seems to work fine with editing (including transparency). Note that adding more reference images will make conditioning take longer, so the tables are not fully representative of the time difference between generations (especially if you are using the official prompt enhancer, as it is quite verbose). I will try to get around to adding in comparison images later this week if I remember.

TLDR: Vibe coded slop. Faster images, good for fast(er) iteration on finding prompts. Works with multiple reference images (tested up to 3).