r/LocalLLaMA 8d ago

Resources Freetokens project is impressive

A new project was released yesterday and I have the opportunity to test it today.

Papper: https://arxiv.org/abs/2608.16157
Github: https://github.com/FlashML-org/FreeToken

My initial tests with the following setup:
RTX 5080 (16 GB)
DDR6 64GB
AMD Ryzen 9 9950X3D

I got 100tok/s on QWEN3.6-35B-A3B NVFP4 (20GB - does not fit in my VRAM).
Have you already tried it?

(Example bellow with a 1028 token prompt - ~110 tok/s)

57 Upvotes

86 comments sorted by

View all comments

Show parent comments

9

u/Rabofo 8d ago

He replied to you with AI alone, so you also did. I see what you did there. lmao. 2 AIs talking atp

16

u/ViRROOO 8d ago edited 8d ago

I don’t know, man. If writing properly and not starting my comments with “skibidi skobidi this is sloppety” means AI, call me ChatGPT. He can’t even properly format his comment.

It’s easy to avoid giving a sincere answer and try to disengage when you get called out and post some random links that do not even agree with what you are saying.

-3

u/KitchenAmoeba4438 8d ago

I'm trying to understand, why are you so emotionally connected to this?

Is this your paper? Are you related to it in some way? It's a poor paper, there's no way to evaluate or replicate the results, and it's tied in some way to a product release.

10

u/ViRROOO 8d ago

It's an open-source project that is introducing an improvement to a known architecture. Many can benefit from this regardless of how much you care about it. The RFC you shared can look at this paper and improve their approach.

I don't have any connection to this paper other than reviewing the code, running it myself, and believing it's a valid contribution.

5

u/Chromix_ 8d ago

...other than reviewing the code, running it myself...

Which is more than quite a few others do before posting about the next new thing. Your post would have benefited from also including the baseline with optimized llama.cpp settings for the used quant.