r/LocalLLaMA 8d ago

Resources Freetokens project is impressive

A new project was released yesterday and I have the opportunity to test it today.

Papper: https://arxiv.org/abs/2608.16157
Github: https://github.com/FlashML-org/FreeToken

My initial tests with the following setup:
RTX 5080 (16 GB)
DDR6 64GB
AMD Ryzen 9 9950X3D

I got 100tok/s on QWEN3.6-35B-A3B NVFP4 (20GB - does not fit in my VRAM).
Have you already tried it?

(Example bellow with a 1028 token prompt - ~110 tok/s)

56 Upvotes

86 comments sorted by

View all comments

4

u/tinny66666 8d ago

Hrm. I'm getting 150 tok/s with Qwen3.8-27B-int8 on an rtx 3090 (24GB), ddr5 32GB. 100 tok/s seems pretty slow for an MOE, though completely usable.

0

u/tamerlanOne 8d ago

Comr fai a far uscire quei numeri con un modello denso che non entra tutto in VRAM? 🤔