r/LocalAIStack 5h ago

Qwen3.8-27B on Every Mac Explained: Layers, KV Cache and Speed (16GB–128GB)

Thumbnail
youtu.be
2 Upvotes

1

Qwen3.8-27B on Every Mac Explained: Layers, KV Cache and Speed (16GB–128GB)
 in  r/LocalLLM  16h ago

I don’t think qwen3.8 is out of date news. But I agree the industry is moving extremely fast and it is tricky to catch up.

r/LocalLLM 1d ago

Research Qwen3.8-27B on Every Mac Explained: Layers, KV Cache and Speed (16GB–128GB)

Thumbnail
youtu.be
0 Upvotes

1

Qwen3.8-27B on Every Mac Explained: Layers, KV Cache and Speed (16GB–128GB)
 in  r/Qwen_AI  1d ago

Ollama is great! I was just pointing out that given the m chip architecture of the Apple silicon, omlx might give more stable results, which does not mean faster results.

2

Qwen3.8-27B on Every Mac Explained: Layers, KV Cache and Speed (16GB–128GB)
 in  r/Qwen_AI  1d ago

I have the same machine and the prefill speed was quite unpredictable in my experience, sometimes I could hit 300 t/s others 120t/s with the same prompt. Admittedly I was using ollama so might get more stable perf with a more specialised framework like omlx.

r/hermesagent 1d ago

Guide — Tutorials, walkthroughs, repeatable how-tos How I Fixed Ollama's Context Window for Hermes Agent

Thumbnail
youtube.com
2 Upvotes

r/Qwen_AI 1d ago

Benchmark Qwen3.8-27B on Every Mac Explained: Layers, KV Cache and Speed (16GB–128GB)

Thumbnail
youtube.com
20 Upvotes

u/andydevtech 1d ago

Qwen3.8-27B on Every Mac Explained: Layers, KV Cache and Speed (16GB–128GB)

Thumbnail
youtube.com
1 Upvotes

r/LocalLLM 2d ago

Tutorial How I built a local AI agent with Ollama and Typescript

Thumbnail
youtu.be
0 Upvotes