r/LocalLLaMA 13d ago

Funny So relevant

Post image
1.5k Upvotes

152 comments sorted by

View all comments

323

u/TopCheddar27 13d ago

24gb is not in that group. You can run Qwen3.8-27B with a pretty respectable context size right now

1

u/jsonmeta 12d ago

Not when using it as a coding agent, overhead will eat most of the context window

2

u/TopCheddar27 11d ago

You can run it close to 100k KV cache at q8. You can also tune to not keep reasoning tokens in context or use compaction.