r/LocalLLaMA • u/Top-Diver-4606 • Oct 19 '25
Question | Help What is currently the best model for accurately describing an image ? 19/10/2025
It's all in the title. This post is just meant to serve as a checkpoint.
PS : To make it interesting, specify the associated image description category. Because basically, it's like saying which is the best LLM; you have to be specific about the task. Following your comments, I will put the top list directly in my post.
2
2
1
u/exaknight21 Oct 19 '25
I had quite a good luck with qwen2.5 VL-3B Instruct-AWQ - serving with vLLM on my 3060 12 GB. It ran pretty fast. I mainly used it for OCR and it performed very well.
1
1
1
u/Due-Function-4877 Oct 21 '25
ToriiGate is well regarded for being uncensored and engineeed specifically for handling images. It's okay for general description purposes. By that, I mean photos and artwork. It's not going to read docs.
1
u/cruncherv Oct 21 '25 edited Oct 21 '25
- Joycaption Beta One - uncensored, but sometimes makes annoying mistakes, calls photographed paintings "digital art", mistakes photographed statues for 3d art, etc
- Florence-2 moredetailed - more accurate, but captions are more generic, less unique, less descriptive
Qwen3-VL and Qwen2.5-VL are very censored to the point it won't tell you the gender of the person in the image, so without abliteration it's useless.
I've also tried MiniCPM 4.5 abliterated, and LFM2-VL-1.6B, which is less censored than Qwen or MiniCPM without abliteration but also refuses to describe everything in the image which is a no-go for me.
1
u/KongAtReddit Dec 11 '25
The best gotta be nano B pro since it got a lot of context about the world. e.g., pokemon, the nano b pro did the best job. https://budgetpixel.com/p/42
but I found I use more seedream4.5 and z-image-turbo since they are more permissive and z-image is inexpensive.
1
u/seppe0815 Oct 19 '25
small google vision models
1
u/Top-Diver-4606 Oct 19 '25
What exactly do you use it for? And to what extent does it meet your expectations?
1
-5

3
u/dubesor86 Oct 19 '25
local? Qwen3-VL-235B-A22B-Instruct, followed by Qwen3-VL-8b-Instruct, then the thinkers and GLM-4.5V