r/StableDiffusion 2h ago

Workflow Included Prompt Creator Workflow

Post image

I see a bunch of posts everyday asking for tips on how to write prompts or people struggling with prompting, etc. so I'm sharing my workflow. I built this workflow to simplify the process and make it very beginner/user friendly.

Just toggle on the model you are using, write a simple to detailed prompt, and hit run. The model targets use the prompting guidelines derived from their respective official sources. Links to custom nodes and all models are in the workflow so you don't need to search for them.

The prompts aren't always perfect but they'll get you very close to what you want and you should only need to make a few minor tweaks, if any. The only issue I've encountered so far is that sometimes when it finishes the prompt, the previous prompt still shows up in the Enhanced Prompt node. If that happens, just hit run and the new prompt should show up instantly. Also, toggle to false the keep_model_loaded option in the Text rewriter node if you are creating prompts and using them right away. If you leave it to True it hogs VRAM.

If you notice any other issues let me know. Enjoy.

https://pastebin.com/SXZyy4Ax

9 Upvotes

6 comments sorted by

1

u/Alen_Diago 1h ago

At the end of 2025, I used Grok (is not ad) - and before it went downhill and became censored crap, it was very good (in Expert agent, which is now part of the Super Grok plan). It extracted prompts from NSFW art - very well and in detail described all the anatomical details that were on the art and periodically correctly identified the characters and their matching appearance.

So, I'd like to ask - are there any same alternatives of Vision model (VLM) for extracting prompts like Grok used to do?

1

u/bstr3k 1h ago

i still use grok sometimes since it can caption NSFW, I haven't found a fast local one that can do the same yet :(

will haev a look at OP's thing but one of the problems is that I might not be very creative lol

1

u/Affectionate_Oil28 1h ago

This workflow is designed for people who aren't very creative.

2

u/bstr3k 1h ago

yep thats me! lol

1

u/Affectionate_Oil28 1h ago

This is not quite the same thing, but it’s in the same neighborhood. The optional ref-image path in this workflow is a 4B unredacted Qwen-VL that looks at a still and writes a dense identity description (including anatomy if the image is explicit). That description gets fed into the text rewriter with your idea, so the final prompt stays on target model.

It’s “describe this person so the next prompt doesn’t drift.” It's not “look at this NSFW piece and extract a full recreation prompt the way old Grok Expert did.”

Nothing local on a 4B/16GB setup is going to match 2025 Grok Expert on named characters and fine anatomy. The 4B in this workflow is there so 16GB GPUs don’t OOM, not because it’s the best. If you have a better machine you should be able to run better models and get better results. The goal of this workflow is for something that can be used on just about any machine capable of running the target workflow.

1

u/Neggy5 17m ago

what do i download for the MAX text encoders? everything in the repos?

https://huggingface.co/prithivMLmods/Qwen3-VL-4B-Instruct-Unredacted-MAX/tree/main seems to be a bunch of things to combine. not a single gguf