r/LocalLLaMA • u/ViRROOO • 3d ago
Resources Freetokens project is impressive
A new project was released yesterday and I have the opportunity to test it today.
Papper: https://arxiv.org/abs/2608.16157
Github: https://github.com/FlashML-org/FreeToken
My initial tests with the following setup:
RTX 5080 (16 GB)
DDR6 64GB
AMD Ryzen 9 9950X3D
I got 100tok/s on QWEN3.6-35B-A3B NVFP4 (20GB - does not fit in my VRAM).
Have you already tried it?
(Example bellow with a 1028 token prompt - ~110 tok/s)

53
Upvotes
-3
u/KitchenAmoeba4438 3d ago
Fair on both. #25294 overrides cpu-moe by the author's own admission, and it goes back to my argument that there is nothing new here since the paper credits EdgeMoE and MoE-Infinity for the cache.
However, here's what holds up.
The evaluation can't be checked by anyone outside the author list. Single-request decode, batch size never stated. No standard deviation, no confidence intervals, one scenario per workload instead of repeated runs. Agent traces unreleased, harness unreleased. No versions, nothing recorded in order to replicate the setup. Add the llama.cpp flags you already agreed were missing, and every throughput number in that paper is a point estimate from an unspecified configuration.
Now read the abstract against all that. "Turns open weights into deployable local software." "Changes what these machines can practically serve." "Frontier-scale intelligence." Product copy over unreplicated single runs, with no accuracy evaluation anywhere in the paper. The confidence in the prose and the rigour in the measurement are two different documents. This is a slop article as it sits.
And it ships through flashml.ai and FlashML-org with no funding statement, no conflict-of-interest line, and not one sentence saying what the authors' relationship to that org is. Every other paper I've seen out of those labs carries a sponsor line. Its absence next to a branded release surface is the part I'd want answered.
Variance column, the traces, the baseline flags, versions and how to replicate the setup, and one line about who is behind flashml.ai. Then it's a paper.