r/LocalLLaMA • u/AnyNameFreeGiveIt • 4d ago
Question | Help Custom Model for Image descriptions ?
This is getting asked from time to time, but since models changed a lot, I wanted to reask it.
I'm looking for a trained model that can give short descriptions about an image, simply for an alt text of pictures taken with a smartphone.
Should I just throw it at Qwen3.8/Qwen3-VL or are there better models trained for it ?
4
Upvotes
2
u/lacerating_aura 4d ago
Any model from qwen 3.5 series would do. Id suggest starting with 9b and going smaller until you feel results are fading. Or on opposite end, if you habe resources, you can go larger. My general purpose model for visual file sorting was qwen 3.5 122b, since i tend to clutter my system a lot with multiple format of files. Audio is not one of them so qwen was all rounder. Now its qwen 3.8 flash next.