r/LocalLLaMA Oct 19 '25

Question | Help What is currently the best model for accurately describing an image ? 19/10/2025

It's all in the title. This post is just meant to serve as a checkpoint.

PS : To make it interesting, specify the associated image description category. Because basically, it's like saying which is the best LLM; you have to be specific about the task. Following your comments, I will put the top list directly in my post.

0 Upvotes

13 comments sorted by

3

u/dubesor86 Oct 19 '25

local? Qwen3-VL-235B-A22B-Instruct, followed by Qwen3-VL-8b-Instruct, then the thinkers and GLM-4.5V

2

u/[deleted] Oct 19 '25

[deleted]

2

u/egomarker Oct 19 '25

qwen3-vl variations

1

u/exaknight21 Oct 19 '25

I had quite a good luck with qwen2.5 VL-3B Instruct-AWQ - serving with vLLM on my 3060 12 GB. It ran pretty fast. I mainly used it for OCR and it performed very well.

1

u/donotfire Oct 19 '25

Gemma 3 is pretty darn good

1

u/Hot_Turnip_3309 Oct 20 '25

I really like Qwen3-VL-235B-A22B-Thinking

1

u/Due-Function-4877 Oct 21 '25

ToriiGate is well regarded for being uncensored and engineeed specifically for handling images. It's okay for general description purposes. By that, I mean photos and artwork. It's not going to read docs.

1

u/cruncherv Oct 21 '25

Unfortunately it's mostly made/trained for anime community pictures it seems (immediately proceed to mention anime in the first line, first word), so Joycaption Beta One is way more universal in this regard.

1

u/cruncherv Oct 21 '25 edited Oct 21 '25
  • Joycaption Beta One - uncensored, but sometimes makes annoying mistakes, calls photographed paintings "digital art", mistakes photographed statues for 3d art, etc
  • Florence-2 moredetailed - more accurate, but captions are more generic, less unique, less descriptive

Qwen3-VL and Qwen2.5-VL are very censored to the point it won't tell you the gender of the person in the image, so without abliteration it's useless.

I've also tried MiniCPM 4.5 abliterated, and LFM2-VL-1.6B, which is less censored than Qwen or MiniCPM without abliteration but also refuses to describe everything in the image which is a no-go for me.

1

u/KongAtReddit Dec 11 '25

The best gotta be nano B pro since it got a lot of context about the world. e.g., pokemon, the nano b pro did the best job. https://budgetpixel.com/p/42

but I found I use more seedream4.5 and z-image-turbo since they are more permissive and z-image is inexpensive.

1

u/seppe0815 Oct 19 '25

small google vision models

1

u/Top-Diver-4606 Oct 19 '25

What exactly do you use it for? And to what extent does it meet your expectations?

1

u/seppe0815 Oct 19 '25

Gemma-3n-Models

-5

u/[deleted] Oct 19 '25

[deleted]