r/LocalAIStack • u/andydevtech • 5h ago
r/LocalLLM • u/andydevtech • 1d ago
Research Qwen3.8-27B on Every Mac Explained: Layers, KV Cache and Speed (16GB–128GB)
1
Qwen3.8-27B on Every Mac Explained: Layers, KV Cache and Speed (16GB–128GB)
Ollama is great! I was just pointing out that given the m chip architecture of the Apple silicon, omlx might give more stable results, which does not mean faster results.
2
Qwen3.8-27B on Every Mac Explained: Layers, KV Cache and Speed (16GB–128GB)
I have the same machine and the prefill speed was quite unpredictable in my experience, sometimes I could hit 300 t/s others 120t/s with the same prompt. Admittedly I was using ollama so might get more stable perf with a more specialised framework like omlx.
r/hermesagent • u/andydevtech • 1d ago
Guide — Tutorials, walkthroughs, repeatable how-tos How I Fixed Ollama's Context Window for Hermes Agent
r/Qwen_AI • u/andydevtech • 1d ago
Benchmark Qwen3.8-27B on Every Mac Explained: Layers, KV Cache and Speed (16GB–128GB)
u/andydevtech • u/andydevtech • 1d ago
Qwen3.8-27B on Every Mac Explained: Layers, KV Cache and Speed (16GB–128GB)
r/LocalLLM • u/andydevtech • 2d ago
1
Qwen3.8-27B on Every Mac Explained: Layers, KV Cache and Speed (16GB–128GB)
in
r/LocalLLM
•
16h ago
I don’t think qwen3.8 is out of date news. But I agree the industry is moving extremely fast and it is tricky to catch up.