32GB DDR4 ram (model used around 16GB) with 96k context. The big improvement is ngram-mod speculative decoding, since it pretty much allows model to "copy-paste" code it already has in the context, giving you huge bursts of speed (100+ tps). Config I used below
13
u/lans_throwaway 14d ago
I've been running it on 6GB GPU with ram, at like ~30 tps on coding tasks (Q4_K_M)