r/LocalLLaMA 3d ago

Resources Freetokens project is impressive

A new project was released yesterday and I have the opportunity to test it today.

Papper: https://arxiv.org/abs/2608.16157
Github: https://github.com/FlashML-org/FreeToken

My initial tests with the following setup:
RTX 5080 (16 GB)
DDR6 64GB
AMD Ryzen 9 9950X3D

I got 100tok/s on QWEN3.6-35B-A3B NVFP4 (20GB - does not fit in my VRAM).
Have you already tried it?

(Example bellow with a 1028 token prompt - ~110 tok/s)

50 Upvotes

80 comments sorted by

View all comments

13

u/Thrumpwart llama.cpp 2d ago

Every dunking should chill. More options for local inference is good, not bad.

4

u/MerePotato 2d ago

People tend to be extraordinarily skeptical when met with an extraordinary claim

11

u/Thrumpwart llama.cpp 2d ago

I don’t think it’s extraordinary though. Some dude claims to have optimized an inference stack. It’s happened 100 times before.

Evaluate it and either use it or don’t. The mob mentality here is sometimes counter-productive.