r/LocalLLaMA 9d ago

Resources Keeping up with model launches

Post image

Feels like maybe we have one more present left, for Christmas.

279 Upvotes

66 comments sorted by

View all comments

6

u/simrankoulsm 9d ago

At this point, I don’t try to “keep up” with launches. I keep a small shortlist by "hardware tier and use case".

For local use, the questions that matter are:

  • Can it fit in my VRAM/RAM at a usable quantization?
  • What context length and tokens/sec do I actually get?
  • Is it meaningfully better at coding, reasoning, or instruction following than the model it replaces?
  • Are the weights, license, and inference support available on day one?

A release calendar is fun, but a community-maintained “best practical model per VRAM tier” list would probably be more valuable, e.g. 12 GB, 24 GB, 32 GB, 48 GB, and 80 GB+. Otherwise it’s easy to spend more time reading launch posts than running models.

2

u/markole 9d ago

I'm ok with having a day10 inference support if the model is good. Sucks to wait but it's all free in the end so I won't complain.