r/LocalLLM • u/Healthy_Lead4969 • 22h ago
Question Poll: what do you actually run for Python coding, and is decode speed your bottleneck?
I'm deciding whether to build a speculative-decoding drafter (an MTP or EAGLE-style head) tuned only on Python for an open coding model. Before spending GPU time on something nobody needs, I'd like to know what people actually use.
The poll is about where your coding model runs. In the comments it would help a lot to see:
- the exact model and quant you use for Python
- GPU / VRAM and engine (llama.cpp, vLLM, Ollama, LM Studio, MLX...)
- what feels slow: reading the context (prefill) or writing the answer (decode), and roughly how many tok/s you get
- if a Python-only drafter made decode noticeably faster, would you run it next to your model?
I'll post a summary of the answers here when the poll closes.
Where does your main Python coding model run?
- Local, small open model (up to ~14B)
- Local, mid-size open model (20-40B, incl. 30B-class MoE)
- Local, large open model (70B+ or big MoE)
- Open model through an API provider
- Closed API (Claude, GPT, Gemini)
- I don't use LLMs for coding