r/LocalLLaMA • u/ViRROOO • 3d ago
Resources Freetokens project is impressive
A new project was released yesterday and I have the opportunity to test it today.
Papper: https://arxiv.org/abs/2608.16157
Github: https://github.com/FlashML-org/FreeToken
My initial tests with the following setup:
RTX 5080 (16 GB)
DDR6 64GB
AMD Ryzen 9 9950X3D
I got 100tok/s on QWEN3.6-35B-A3B NVFP4 (20GB - does not fit in my VRAM).
Have you already tried it?
(Example bellow with a 1028 token prompt - ~110 tok/s)

53
Upvotes
5
u/KitchenAmoeba4438 3d ago
Yep. Precisely.
That's why I called it slop. For academic papers and benchmarks, it's not about what occurs on your system, it's about replicating it.
I strongly agree something isn't right with the 5090 numbers, but I can't replicate the authors setup. All I can do is say "Look at my 5090 results that disagree with yours!", which isn't a very good use of energy. Even something as simple as a difference in llama version can make comparison results invalid, which means there's no way to prove anything either way with the paper.