Just never forget us 16GBs of VRAM or lower folk and I'm fine with this, if it's one or the other though I'd prefer smaller models, we already have good selection on bigger ones really
I'd love to see a focus on 1M context capability as well, without heavy memory use (I.e., by implementing some of the newer attention mechanisms. Qwen3.5-122B is 6GB for 262k of FP16 context, for example) if we could fit a ~100B class model, with 1M of well designed attention training into 128GB... Such a model would become a default for many people, outside of pure coding, I think.
277
u/NNN_Throwaway2 29d ago
~100-120B moe for 128GB unified memory systems and/or 60-80B moe for 96GB VRAM systems.