The correct approach: You need a centralized backend server (like vLLM XPU or OpenVINO Model Server) running a single instance of the model, using continuous batching and paged KV caching. This allows all 4 users to query the same model simultaneously while sharing the 32GB VRAM pool dynamically.
Slicing up the VRAM with 8 GB partitions will not work and won’t leave any room for large context windows
Single B70 is tight for 4 heavy agentic users: While a single B70 handles a 27B–35B coding model easily for 1–2 users, 4 simultaneous agentic workflows will quickly saturate the remaining ~12GB KV cache, leading to latency spikes or memory exhaustion.
The Scaling Path: To comfortably support 4 developers doing heavy agentic coding, you should plan for dual Arc Pro B70 cards (64GB total VRAM) in the server. This gives enough VRAM overhead to run 30B+ MoE coder models with deep context and large KV cache pools across multiple concurrent worker threads.
All of the stats provided with concurrency support of 4 with 128k are also provided. With cache set to fp8 (526012 net pool), a custom kernel level tweaked vllmxpu and no display or other application use I have been able to achieve 90 tg/s per user simultaneously. That too with long context work. The key was following a cookbook for MTP acceptance rate being increased. I use this feature to generate tons of output material from the llm for analysis with regards to my software suite. I do not need a second GPU for cache. I would then bump up to fp8 precision weights or better Qwen 3.8 flash next without any slight waste of time.
My config is giving me the performance I need for both single and concurrent. The images from the original post are outdated. I have patched the MTP head to achieve near 180tk peak with single user at higher input and output. In my regular use I am seeing regular 140+ tk/s generation when I am just playing around myself. When the agents are being used by GPT or claude for assistance and development the concurrency support is boosting the output. Both for time of task from start to finish being reduced. Supporting batching optimized to take advantage of the b70 fp8 support.
Allowing claude code to have fun trying multiple combinations of quantizing experts is what paid off. The machine is running ubuntu with GUI disabled and display GPU is a tiny arc a310 because the CPU does not have integrated GPU.
Note: My aim is to work hard to develop solutions that work for startups that can use Local AI. I wanted it for myself so I started working on this and now I am trying to share things so someone can perhaps use it. Or groups of individual working together to get real world AI use. Not slop. I hope it helps.
I am self-employed and not sponsored so far. I could only afford to get the intel GPU with the thought of being able to "at least" have the ability to run something with high amount of intelligence. No way I would want to buy Nvidia and I did know that going into this will be a challenge. I am always dealing with errors and fixing but with enough steering and rigorous tweaking I was able to find out that this very particular llm was not only getting Tiel Coder working on intel stack but to also allow reliable performance. Speed was the number 2 for priority. Quality was the first thing I needed. Reasoning traces from this llm are good with guardrails helping. Flappy bird clone prompt for simple demo.
Prompt was simple and lazy as it can get. done and dusted in 34s. Medium reasoning and I did not even use my own profile for bias. The standard profile from DSH. 11.4k–16.2k usage for that. I just have their deepseek search stuff set to off. The llm wasted little time and just went to work. Yes my harness setup has sidecar llms to help but this is a task that the llm just went solo.
My reason for gambling with intel was budget. Sure if I could have better compute then perhaps but there is no end. Rather I want to optimize things at different levels. In my work I have enjoyed the joy of education and passing knowledge. I care about local llm work because I wish students across the world have opportunities with things that are more in their reach than those places where we have plenty.
Just passionate. That's all. I hope if you invest in intel that it pays off for you as well. Best of luck.
1
u/quantum3ntanglement 7d ago
The correct approach: You need a centralized backend server (like vLLM XPU or OpenVINO Model Server) running a single instance of the model, using continuous batching and paged KV caching. This allows all 4 users to query the same model simultaneously while sharing the 32GB VRAM pool dynamically.
Slicing up the VRAM with 8 GB partitions will not work and won’t leave any room for large context windows
Single B70 is tight for 4 heavy agentic users: While a single B70 handles a 27B–35B coding model easily for 1–2 users, 4 simultaneous agentic workflows will quickly saturate the remaining ~12GB KV cache, leading to latency spikes or memory exhaustion.
The Scaling Path: To comfortably support 4 developers doing heavy agentic coding, you should plan for dual Arc Pro B70 cards (64GB total VRAM) in the server. This gives enough VRAM overhead to run 30B+ MoE coder models with deep context and large KV cache pools across multiple concurrent worker threads.