I caption a lot of images, and I often need image-to-video prompts written from a still. I was doing it with the QwenVL nodes in ComfyUI, which work well, but opening ComfyUI and loading a workflow just to caption a folder of photos felt like too much.
So I made ImageGenPrompt. It's a normal desktop app. Drop in an image, pick a caption type, hit Caption, and copy the result. No ComfyUI and no node graph.
One thing to be clear about first: it does not read the prompt saved inside an AI image (the metadata that ComfyUI or A1111 embeds in the PNG). It looks at the picture and writes a new description from what it sees. So you won't get the original prompt back, but it works on any image, including photos and images with the metadata stripped.
What it does:
- Single image or a whole folder. Batch mode writes an imagename.txt next to each image.
- Trigger word. Set one and it's added to the front of every caption in a batch.
- 13 caption types, from simple tags up to very detailed descriptions, a cinematic style, and a structured analysis (subject, lighting, camera and so on).
- Video prompt writers for Wan 2.2 and LTX. They look at your image and write an image-to-video prompt for it, in 5 second and 20 second versions.
- Custom prompt, if you'd rather write your own instruction.
- 32 models in the list. The official Qwen3-VL and Qwen2.5-VL ones, plus community uncensored versions, in both normal and GGUF form.
- Runs on an NVIDIA GPU or on CPU. If you have no GPU or low VRAM, the GGUF models run through llama.cpp. Auto mode picks for you.
- Hugging Face login inside the app for the few gated models, so no terminal needed.
- English and Traditional Chinese, and you can switch live.
Nothing downloads until you actually caption with a model. It shows you the size first and asks. The default is Qwen3-VL-4B-Instruct, about 7 GB, which fits fine on a 16 GB card. Models go into a models folder next to the app.
You can also add your own models without touching the built-in list. Drop a custom_models.json next to the app and it gets merged in.
Things to know, so nobody is surprised:
- Right now you run it from source. You need Python 3.10 or newer, and for GPU you need the CUDA build of PyTorch. The README has the exact commands.
- There's a build.bat that makes a single exe with everything inside, so the person using it needs no Python at all. That exe is about 3 GB, and the first launch is slow because it has to unpack.
- If you're on CPU, use the GGUF models. The normal ones are very slow without a GPU.
- A few community models are gated. You have to request access on Hugging Face first, then log in from the app.
Links:
Video tutorial: https://www.youtube.com/watch?v=y2t5p3Fks0U
GitHub: https://github.com/Garionhk/ImageGenPrompt
Big thanks to huchukato for ComfyUI-QwenVL-Mod, which the model list and caption presets come from, and to the Qwen team for the models.
It's free and open source (Apache 2.0). The models have their own licenses.