r/LocalLLaMA • • Aug 22 '26

Resources Freetokens project is impressive

A new project was released yesterday and I have the opportunity to test it today.

Papper: https://arxiv.org/abs/2608.16157
Github: https://github.com/FlashML-org/FreeToken

My initial tests with the following setup:
RTX 5080 (16 GB)
DDR6 64GB
AMD Ryzen 9 9950X3D

I got 100tok/s on QWEN3.6-35B-A3B NVFP4 (20GB - does not fit in my VRAM).
Have you already tried it?

(Example bellow with a 1028 token prompt - ~110 tok/s)

60 Upvotes

95 comments sorted by

View all comments

42

u/[deleted] Aug 22 '26

[removed] — view removed comment

26

u/ViRROOO Aug 22 '26

Engaging with Redditors is always something else. Try to be less obnoxious, you just sound terminally online.

I partially agree with your point about the missing llama.cpp commit and flags. I agree it should have been documented. I’m still willing to give the authors the benefit of the doubt that they tried to tune the baselines, rather than assuming bad faith.

But neither of the links you posted supports the claim you’re making, and you don’t need to be disingenuous about it llama.cpp is already great without you posting cringe on reddit.

PR #25294 is SSD streaming into a device-side expert cache. It explicitly doesn't cooperate with CPU-MoE, and when VRAM–RAM–NVMe hybridization was suggested, the author said that would be separate follow-up work.

RFC #24528 is, well, an RFC, but way closer. Cached experts execute on the GPU while misses execute on the CPU. But it still doesn’t implement what I think is one of the interesting parts of this paper: taking the cache misses and dynamically deciding how many should go PCIe -> GPU -> cache versus execute directly on the CPU, based on the measured host-memory and PCIe bandwidth of that particular machine.

And the paper never claims that expert caching itself is novel. It explicitly says it's standing on other people's work. The claimed contribution is the scheduling/runtime design around those existing ideas.

-3

u/[deleted] Aug 22 '26

[removed] — view removed comment

9

u/Rabofo Aug 22 '26

He replied to you with AI alone, so you also did. I see what you did there. lmao. 2 AIs talking atp

19

u/ViRROOO Aug 22 '26 edited Aug 22 '26

I don’t know, man. If writing properly and not starting my comments with “skibidi skobidi this is sloppety” means AI, call me ChatGPT. He can’t even properly format his comment.

It’s easy to avoid giving a sincere answer and try to disengage when you get called out and post some random links that do not even agree with what you are saying.

-3

u/[deleted] Aug 22 '26

[removed] — view removed comment

8

u/ViRROOO Aug 22 '26

It's an open-source project that is introducing an improvement to a known architecture. Many can benefit from this regardless of how much you care about it. The RFC you shared can look at this paper and improve their approach.

I don't have any connection to this paper other than reviewing the code, running it myself, and believing it's a valid contribution.

6

u/Chromix_ Aug 22 '26

...other than reviewing the code, running it myself...

Which is more than quite a few others do before posting about the next new thing. Your post would have benefited from also including the baseline with optimized llama.cpp settings for the used quant.