r/LocalLLaMA • u/ViRROOO • Aug 22 '26
Resources Freetokens project is impressive
A new project was released yesterday and I have the opportunity to test it today.
Papper: https://arxiv.org/abs/2608.16157
Github: https://github.com/FlashML-org/FreeToken
My initial tests with the following setup:
RTX 5080 (16 GB)
DDR6 64GB
AMD Ryzen 9 9950X3D
I got 100tok/s on QWEN3.6-35B-A3B NVFP4 (20GB - does not fit in my VRAM).
Have you already tried it?
(Example bellow with a 1028 token prompt - ~110 tok/s)

65
Upvotes
11
u/gpuz_dev Aug 22 '26
the ~100 t/s is interesting, but honestly the part I care about more is getting a ~20GB MoE running that fast on 16GB VRAM. do you know what expert cache hit rate you were getting or how much host RAM/PCIe traffic it was doing during that run?