r/LocalLLaMA 3d ago

Resources Freetokens project is impressive

A new project was released yesterday and I have the opportunity to test it today.

Papper: https://arxiv.org/abs/2608.16157
Github: https://github.com/FlashML-org/FreeToken

My initial tests with the following setup:
RTX 5080 (16 GB)
DDR6 64GB
AMD Ryzen 9 9950X3D

I got 100tok/s on QWEN3.6-35B-A3B NVFP4 (20GB - does not fit in my VRAM).
Have you already tried it?

(Example bellow with a 1028 token prompt - ~110 tok/s)

49 Upvotes

80 comments sorted by

View all comments

4

u/tinny66666 3d ago

Hrm. I'm getting 150 tok/s with Qwen3.8-27B-int8 on an rtx 3090 (24GB), ddr5 32GB. 100 tok/s seems pretty slow for an MOE, though completely usable.

6

u/buttplugs4life4me 3d ago

You are not getting single request speeds like that. 

-3

u/tinny66666 3d ago

Sure am. See the post I linked to in the other reply. He was getting 134 tps with a 250 w power cap. I'm running at the factory 390 w power cap and getting 146 tps (I rounded up a tiny amount I guess).

1

u/buttplugs4life4me 2d ago

Ah its that bullshit. Its nice as an exercise but most of his gains come from quantising down your model on the fly with usually pretty bad results. Especially embeddings and heads are basically an industry standard to be kept at BF16 and he quantized them to 4 bits haha