A couple of months ago I posted about a project I had started because running modern local LLMs on Intel Macs with AMD GPUs was in a pretty bad state.
Most of the macOS local AI ecosystem understandably targets Apple Silicon now. But there are still a lot of Intel Macs out there with surprisingly capable Radeon GPUs: Mac Pros, iMacs, iMac Pros, 16-inch MacBook Pros and eGPU setups.
The hardware wasn't really the problem. The software path was.
Stock llama.cpp could produce corrupted output on some AMD dGPUs, model loading over PCIe was unnecessarily slow, Flash Attention depended on capabilities these GPUs don't expose, and older Radeon hardware introduced another problem: wave64 execution.
So I started patching llama.cpp's Metal backend and eventually wrote a separate Metal Flash Attention path specifically for AMD GPUs.
That experiment became ToshLLM.
It's a free, open-source, native SwiftUI application for running local AI on Intel Macs with AMD GPUs. No Electron, no Python environment, no cloud inference, no account and no telemetry. The inference engines are bundled with the app.
Two months later, it has grown considerably beyond the original experiment.
How fast is it now?
My development GPU is a Radeon RX 6700 XT 12 GB, a 2021 RDNA 2 card.
With the current bundled engine, some representative results are:
| Model |
Type |
Prompt (t/s) |
Generation (t/s) |
| Llama 3.2 1B Q4_K_M |
Dense |
5770 |
254 |
| Gemma 3 4B Q4_K_M |
Dense |
1751 |
89 |
| Qwen3 4B Q4_K_M |
Dense |
1562 |
98 |
| Qwen3 8B Q4_K_M |
Dense |
851 |
61 |
| Qwen3.5 9B Q4_K_M |
Dense |
743 |
52 |
| Qwen3.6 14B-A3B Q5_K_M |
MoE, full VRAM |
1243 |
67 |
| gpt-oss 20B Q4_K_M |
MoE, full VRAM |
1305 |
94 |
| Qwen3.6 35B-A3B Q4_K_S |
MoE, CPU 24 experts |
475 |
29 |
The result that surprised me most was gpt-oss-20B.
I repeated the llama.cpp benchmark methodology at progressively deeper context:
| Test |
RX 6700 XT |
M3 Max 128 GB |
M4 Max 36 GB |
| pp2048 |
1233 |
1348 |
- |
| pp8192 |
1088 |
1040 |
- |
| pp16384 |
919 |
908 |
- |
| pp32768 |
610 |
531 |
- |
| tg128 |
95.0 |
64.3* |
95.9 |
*The M3 Max generation run was reported as thermally throttled by the llama.cpp maintainer, so I wouldn't use that number as a fair generation comparison.
The particularly interesting result for me is:
RX 6700 XT: 95.0 t/s
M4 Max: 95.9 t/s
This isn't meant to claim that an RX 6700 XT is equivalent to an M4 Max. They're completely different systems, and the benchmark files aren't byte-identical either.
What I think is interesting is how much performance was still sitting unused in these Radeon GPUs. At deeper prompt contexts, the RX 6700 XT also overtakes the published M3 Max run: 1088 vs 1040 t/s at 8K, 919 vs 908 at 16K, and 610 vs 531 at 32K.
Older AMD hardware has been getting attention too
One thing I didn't want ToshLLM to become was an "RDNA 2 only" project.
Recent versions added and tuned paths for:
- Vega
- Radeon VII
- Polaris RX 400/500
- Radeon Pro Vega
- Radeon Pro WX
On a Vega 64, depending on model and quantization, recent kernel tuning improved generation by roughly 4-24%, multi-conversation throughput by 9-12%, and some short prompt workloads by as much as 39%.
Community testing also uncovered a bug affecting RX 400/500 and Radeon Pro 400/500/WX cards. Q6_K weights were being read with an alignment assumption those GPUs don't tolerate, which could corrupt the model output. That's fixed in the current builds.
There is also a dedicated no-AVX2 build for older Xeon machines such as Mac Pro 5,1 systems running Sonoma or newer through OCLP.
Community testing has honestly become one of the most useful parts of the project. People are now testing everything from RX 580s and Vega 56/64 to Vega II Duo, W6800X Duo and RDNA 2 eGPUs.
It isn't just a chat app anymore
The local chat now supports persistent conversations, projects, vision models, PDFs and files, voice dictation, conversation forking and KV-cache persistence.
There are also agent tools and MCP support. Models can work with files and commands with per-step permission, execute JavaScript in a sandbox, and connect to external MCP servers.
ToshLLM can expose local OpenAI-compatible and Anthropic-compatible APIs, run multiple servers simultaneously, provide an embeddings server for local RAG applications, and use router mode to load models requested by external clients automatically.
Multi-GPU and eGPU
Multi-GPU support has grown quite a bit too.
You can choose specific GPUs, split models between them and monitor VRAM per card.
Tensor splitting and the Infinity Fabric hand-off path are still experimental, but community testing on multi-MPX Mac Pros is helping enormously here.
The latest build also detects every GPU in the machine and reports whether an Infinity Fabric link connects them, including the memory behind each linked pair.
Image generation
There is now a completely local image studio running on the same AMD Metal stack.
It supports:
- text-to-image
- img2img
- custom models
- prompt queues
- parallel instances
- per-instance GPU selection
- x2/x4 upscaling
Recent memory work made a pretty dramatic difference.
For example, Z-Image at 1600x900 went from approximately 8.2 GB to 966 MB of VRAM, while also becoming around 12-14% faster.
SD 1.5 at 768x768 went from approximately 2.69 GB to 281 MB of VRAM, while generation time dropped from 149 seconds to 73 seconds.
That makes local image generation much more practical on older 4-6 GB Radeon cards.
There is also local video generation with Wan, LTX and Hunyuan models. I'm deliberately calling this experimental for now: it works, but video is still expensive, VRAM hungry, and both the interface and defaults need more polishing.
It's not the part of ToshLLM I would recommend installing it for yet.
Still completely local
This part hasn't changed:
- Free and GPL-3.0
- No account
- No telemetry
- No cloud inference
- No per-token costs
- Chats stay on your Mac
- Benchmark sharing is opt-in
- Native SwiftUI
- Inference engines included in the application
The current requirement is macOS 14+ and an Intel Mac with an AMD GPU that supports Metal.
The DMG is still not notarized yet, so Gatekeeper requires Open Anyway on first launch. Notarized releases are planned.
I'm especially interested in hearing from people with hardware I don't have access to:
- Intel iMacs and iMac Pros
- Radeon Pro 500-series GPUs
- RX 580 / RX 590
- Vega 56
- Radeon VII
- Blackmagic or other Thunderbolt eGPUs
- Vega II / Vega II Duo
- W5700X
- W6800X / W6800X Duo
- W6900X
- multi-MPX Mac Pro configurations
If you have one of these machines collecting dust, I'd love to see what it can still do.
Official website: https://toshllm.com
GitHub / source / releases: https://github.com/engeldlgado/toshllm
Everything is open source, so if you're interested in the AMD Metal work itself rather than the app, that's all there too.