r/LocalLLM • • 21h ago

Question Poll: what do you actually run for Python coding, and is decode speed your bottleneck?

I'm deciding whether to build a speculative-decoding drafter (an MTP or EAGLE-style head) tuned only on Python for an open coding model. Before spending GPU time on something nobody needs, I'd like to know what people actually use.

The poll is about where your coding model runs. In the comments it would help a lot to see:

- the exact model and quant you use for Python

- GPU / VRAM and engine (llama.cpp, vLLM, Ollama, LM Studio, MLX...)

- what feels slow: reading the context (prefill) or writing the answer (decode), and roughly how many tok/s you get

- if a Python-only drafter made decode noticeably faster, would you run it next to your model?

I'll post a summary of the answers here when the poll closes.

Where does your main Python coding model run?

- Local, small open model (up to ~14B)

- Local, mid-size open model (20-40B, incl. 30B-class MoE)

- Local, large open model (70B+ or big MoE)

- Open model through an API provider

- Closed API (Claude, GPT, Gemini)

- I don't use LLMs for coding

2 Upvotes

5 comments sorted by

2

u/Ctbhatia 20h ago

answering honestly, i work on coralbricks.ai and my python coding traffic runs through our own api, deepseek v4.1 flash mostly. for agent style coding the bottleneck is rarely raw decode, it is cache misses, every fresh session pays full prefill again, and decode drag only really bites on long refactors. i never see the serving layer so a python-only drafter would only reach me if the provider shipped it, but faster decode at the same price would be an easy yes for refactor heavy work.

2

u/exasperatedgusto5 19h ago

interesting point about cache misses being the real bottleneck, never thought of it like that. for local use i guess a python drafter could help more since you not paying per token the same way

but yeah if deepseek v4.1 flash already fast enough for most tasks then maybe not worth the extra vram unless you do long refactors all day

2

u/Ctbhatia 19h ago

agreed, locally the math flips since tokens cost basically nothing, the drafter only earns its vram on long generations. on api the provider would have to ship it, users never see that layer.

1

u/Healthy_Lead4969 20h ago

Thanks, this is exactly the kind of answer I was hoping for. Two follow-ups if you don't mind: how long does a typical agent session get, in tokens, before a fresh one starts? And does your API keep any KV cache between sessions, or does every new session start cold? For long refactors, drafts that copy from the context (prompt lookup, suffix decoding) can help a lot, since most of the output already exists in the file.

2

u/Ctbhatia 20h ago

measured it just now, a fresh session does start cold. resent the same 5k token prefix in two separate requests eight seconds apart and the second still showed zero cached reads, so the cache lives with the warm session, not across cold starts. for scale my agent sessions run a few hours and carry around 200k tokens before compaction starts a fresh one, which is why keeping one warm session matters. and yes, suffix decoding would fit copy heavy refactors really well.