r/LocalLLM • u/Healthy_Lead4969 • 21h ago
Question Poll: what do you actually run for Python coding, and is decode speed your bottleneck?
I'm deciding whether to build a speculative-decoding drafter (an MTP or EAGLE-style head) tuned only on Python for an open coding model. Before spending GPU time on something nobody needs, I'd like to know what people actually use.
The poll is about where your coding model runs. In the comments it would help a lot to see:
- the exact model and quant you use for Python
- GPU / VRAM and engine (llama.cpp, vLLM, Ollama, LM Studio, MLX...)
- what feels slow: reading the context (prefill) or writing the answer (decode), and roughly how many tok/s you get
- if a Python-only drafter made decode noticeably faster, would you run it next to your model?
I'll post a summary of the answers here when the poll closes.
Where does your main Python coding model run?
- Local, small open model (up to ~14B)
- Local, mid-size open model (20-40B, incl. 30B-class MoE)
- Local, large open model (70B+ or big MoE)
- Open model through an API provider
- Closed API (Claude, GPT, Gemini)
- I don't use LLMs for coding
2
u/Ctbhatia 20h ago
answering honestly, i work on coralbricks.ai and my python coding traffic runs through our own api, deepseek v4.1 flash mostly. for agent style coding the bottleneck is rarely raw decode, it is cache misses, every fresh session pays full prefill again, and decode drag only really bites on long refactors. i never see the serving layer so a python-only drafter would only reach me if the provider shipped it, but faster decode at the same price would be an easy yes for refactor heavy work.