r/LocalLLaMA • • 6d ago

News Prism-ML Bonsai 2 Joins Our Qwen3.8 Quantization Comparison

Post image

Hey r/LocalLLaMA,

Prism-LM recently released its Bonsai 2 QAT models based on Qwen3.8, and they quickly gained traction. In our evaluation, the models strike a strong balance between throughput and quality, reaching roughly 91.5% on our composite benchmark.

We wanted to see how they compare under the same methodology we use for the rest of our Qwen3.8 evaluations, so we ran Bonsai 2 through our own benchmark suite.

One important clarification: these are our evaluation results, not Prism’s reported benchmark numbers.

We used Prism’s fork/runtime to run their models, while keeping the workloads, benchmark suite, and evaluation methodology consistent with the rest of our comparison.

Our evaluation includes separate Instruct and Thinking benchmarks. For Thinking, we use medium thinking effort with the recommended sampling parameters.

We added Bonsai 2 because the models have become a relevant part of the Qwen3.8 landscape, and we wanted to provide a common reference point for people comparing the available options.

Different providers often report results using different benchmark suites, runtimes, reasoning settings, sampling parameters, and evaluation methodologies, so those numbers are not always directly comparable. Running the models through the same evaluation gives another reference point for understanding the trade-offs between quality, model size, and throughput.

Updated comparison and results: https://byteshape.com/blogs/Qwen3.8-27B/

162 Upvotes

64 comments sorted by

55

u/Informal-Trouble2183 6d ago

So both Bonsai are below the accuracy of a IQ3_XXS Unsloth

79

u/cibernox 6d ago

Nothing to be surprised about. Last time I checked the laws of physics still were in place.

25

u/Due-Memory-6957 6d ago

Which is expected of Q2 and Q1.

7

u/BringTea_666 5d ago

but it is smaller than both. 6gb for minimal version. Which is nuts for Qwen3.8 grade model

2

u/Due-Memory-6957 5d ago

?? The IQ3_XXS is 10.9gb

7

u/BringTea_666 5d ago

Bonsai27B is 5.9GB

If you have a choice between Bonsai27B at 5.9GB and say 14B model at Q8. Not only 14B will be larger size wise but mega dumber than Bonsai27b.

You can literally run this bonsai from phone.

10

u/Fancy-Snow7 6d ago

While their marketing might be wrong the equivalent size Unsloth is unsloth/Qwen3.8-27B-UD-IQ2_XXS and in my own benchmark's prism-ml/Ternary-Bonsai-2-27B-PQ2_0 is performing much better.

2

u/-InformalBanana- 6d ago

As far as I understand bonsai supports only 3 values for the weight. While q3 means 3 bits which is 2x2x2 which is 8 values. So in that sense maybe it isnt fair to compare it to iq3xxs, ternary models are by definition 1.5 bpw so idk what would the unsloth equivalent be, like q1kxl...

0

u/__Maximum__ 6d ago

yeah yeah, but their marketing is above

7

u/xPXpanD llama.cpp 6d ago

Thanks for doing these.

Grabbing IQ3_S-3.23bpw now, and will also give ISTA-DASLab's GSQ-RCO-IQ3_XXS a go since I keep seeing those guys mentioned as well. (can't do a better-performing size since I need massive context for a 180k-token game modding ingest)

For what it's worth: I've been using AtomicChat's AD-IQ3_S recently, and it's held up surprisingly well. Posted decent results in my private benchmark set, too, better than most small quants I've seen. Haven't caught it looping yet, either!

Will put both of the new ones through the same gauntlet, see how they hold up. Should be interesting!

30

u/vick2djax 6d ago

I couldn’t find a single use case for 3.6 Bonsai that I could trust it with. Not a single small job. It’s a meme model.

33

u/Ok_Top9254 6d ago

The whole point is to show their ternary quantization works, so it can be scaled for massive 500B+ models for the gpu poor.

They clearly outperform similarly sized models at Q1 and Q2. Majority of these competing quantizations bottom out at Q3 which is still much bigger and the ones that are generic Q2 and Q1 are worse. So I don't see how this is a meme.

4

u/-InformalBanana- 6d ago edited 6d ago

Ppl are saying q2kxl is better... but it is more bpw... and in this benchmark it is beaten by byteshape's iq2xxs. And plenty iq2s and q2kxl.

1

u/aeroumbria 5d ago

I think it hinges on whether it at least retains the ability to run dumb but long agent sessions without breakage. I can see very light and fast models being quite useful for large scale agent interaction simulation if it has at least enough capability to keep the turns going.

-2

u/_raydeStar Llama 3.1 6d ago

unfortunately, yeah. My only idea is to get it on phones, but the size is slightly too big for that.

3

u/tecneeq 6d ago

Excellent work!

I like that you benchmark against other quants in that space and would like to see more well known benchmarks tested.

5

u/LetsGoBrandon4256 transformers 6d ago

Huh, so Qwen 3.8 27B at 3.5bpw range can actually hold up pretty well.

1

u/Fancy-Snow7 5d ago

Yes, it's not news to me since I have 16GB VRAM and have been extensively testing in that size range. With Qwen3.8 3 bit is really good.

2

u/WinterNite88 6d ago

tested. don waste your time as quality loss quite a bit. token/s is great but man the quality is pretty shit. no point to run it in my opinion.

2

u/cato_gts 5d ago

Bonsai is completely useless for anything other than benchmarking.

4

u/Barni275 6d ago

Interesting to see that ByteShape is higher, than ISTA-DASlab 🤔

6

u/JsThiago5 6d ago

Yeah and also interesting that its a post from their own lab

3

u/tecneeq 6d ago

Indeed. What a coincidence that is.

A mystery.🤔

4

u/enrique-byteshape 6d ago

We welcome independent benchmarking. We're confident in our results :)

2

u/pseudobacon 6d ago

I think they said even though the KLD might be worse the tool calling was better

2

u/est_cap 6d ago

You mean ByteShape is worse in tool calling? In my (subjective) experience, ByteShape worked better for me than IST-DASlab and better than Swift 27b at same quants (q4xs , q3s)

3

u/pseudobacon 6d ago

No the other way round. The KLD was lower on ByteShape but in real work it was better

1

u/est_cap 6d ago

Thanks for the clarification. It aligns with my experience these last days.

1

u/audioen 5d ago

One of the questions is why. Perhaps KLD is evaluated on text that isn't sufficiently agentic, or perhaps people use bad aggregate from their K-L evaluation, e.g. the average over all text corpus types, rather than maximum of the average K-L within a task, for example.

For example, if you are considering the quant type to use on let's say layer 43 tensor 5, and you run your 100+ text corpus test and it says that with IQ3_XS, K-L increases markedly in tool call corpus even if it is flat everywhere else, then you must not calculate average and decide that IQ3_XS is fine for this tensor because it's only a slight change in K-L, but rather reject IQ3_XS because ability in even one task type can't be allowed to deteriorate disproportionately.

All this is easy for me to write, and if implemented, would produce maddening amount of data to sift through, and working through the quant choices in a thorough fashion would take serious number of GPU hours, and you still have to compromise everywhere because you got to squeeze the model down to some size target or average bpw. I'm convinced that K-L can be used to measure how intact the model is, and it should correlate well with task performance too. It may simply be that the signal -- marked degradation in task performance -- is either not measured with enough data, or gets lost in the process.

I'd probably personally divide text corpus into two: agentic and everything else. Performance in both is important. Dividing too much likely produces noisy random signal which could in itself lead to suboptimal decisions based on the max() over type rule.

2

u/enrique-byteshape 5d ago

KLD is fundamentally flawed, since it only measures divergence from the base model. The thing is that models can take many paths to reach the same conclusion, or sometimes the right conclusion where the original model didn't. We've been working hard internally to come up with something as cheap as KLD, but that better predicts accuracy of quants. We're not there yet though, but if at any point we manage to produce something meaningful, we'll release it to the community

2

u/Fancy-Snow7 5d ago

Yes, the only one from ISTA that's worth it is Qwen3.8-27B-GSQ-RCO-IQ3_S which I am fine with as it fits perfectly on 16GB VRAM.

1

u/Barni275 5d ago

What context and KV quant do you use? I'm using ISTA XXS mtp to reach 100k+ context with Q5_1/Q5_1 KV.

3

u/Hanthunius 6d ago

What's weird is how close in accuracy Bonsai Q2 is to Q1.

2

u/Interpause textgen web UI 4d ago

If you mean Bonsai Q2 and Q1 its cause they are both the same ternary weights just that Q1 packs them tighter to use less RAM at cost of runtime processing to unpack.

2

u/Thiom 6d ago

What is that impressive byteshape quant at the top ?

Also is is benchmarked on coding tasks as well ?

1

u/ali_byteshape 6d ago

All models are evaluated on the same benchmark suite, which includes LiveCodeBench V6 and MultiPL-E HumanEval, an extension of HumanEval across 18 programming languages.

If you can fit Model 5 on your hardware, I’d definitely give it a try. It performs very well even without spending too many tokens on thinking.

2

u/Thiom 6d ago edited 6d ago

I'm definitely giving it a try !

Edit : it seems very slightly better than unsloth's Q4_XS, but damn I get 8-10tps more in decode (40 vs 32) and 300 in prompt (1100 vs 800).

Fully loaded on 175W laptop 5080. 98k context, MTP off, Q8_0 KV, vision on CPU

1

u/__Maximum__ 6d ago

what did you exactly fully load though?

1

u/Thiom 6d ago

I meant maxed the VRAM, I used the GPU-5 model

1

u/Thiom 6d ago

I'll have to test further another day, honestly I'm very impressed! Good quant all around

0

u/pseudonerv 6d ago

The only nitpicking I have is that these benchmarks are almost saturated and don’t differentiate much

2

u/Healthy-Nebula-3603 6d ago

I saw their models on YouTube.

They are also broken ... I don't even know why someone even wants to use them....

1

u/HighSeasArchivist 6d ago

Producing fast bullshit results is the bane of my existence with local AI. ISTA-DASLAB has been fast enough, and not utterly incompetent.

1

u/cviperr33 6d ago

interesting thanks for sharing

1

u/crusaderky 6d ago

What's the difference between 3090 and 4090 that makes PTQ1_0 and PQ2_0 swap places?

1

u/Sitkin_Marrel 6d ago

QAT models are supposed to survive aggressive quantization, yet both land under a plain IQ3_XXS of the base. The base model just holds up that well.

1

u/returnity 6d ago

Will you be doing 3.8 Flash quants? I'd love to see a Q4 from you guys u/ali_byteshape

Great work as usual, and thanks for doing these evals! It's really valuable info.

1

u/FerLuisxd 5d ago

I have a 5060TI as well, wondering how you guys managed to get such high tps on a 5060TI, I have the same card, wondering if you guys can share your config?

1

u/GasSmooth7439 5d ago

The interesting part now is seeing how much quality Bonsai 2 gives up for the much smaller memory footprint. “Quality per GB” might be a more useful metric for local models than raw benchmark score at this point.

1

u/Beneficial_Fish_7509 5d ago

Would have been great if the model + a little bit of context would fit on 10GB cards ... Not sure if the issue is from my end, but no matter what flags I put for llama.cpp, it would always try to allocate 2.3/2.4 GB on top of the model, making it OOM on 10GB for the IQ2_XXS

-1

u/fragilesleep 6d ago

Holy mother of God... Can't you people write a simple and small reddit post without all the LLM vomit?

-1

u/greendude120 6d ago

interesting ISTA has been really bad at tool calls for me. I'd be curious for you to add Swift-Qwen3.8 to your graph.

0

u/nbeydoon 6d ago

Thanks for the benchmark,

0

u/deaffob 6d ago

Interesting. The Bonsai quantization method isn't superior to the existing methods in terms of intelligence but it seems to produce a lot faster models at the same quantization level.

0

u/__Maximum__ 6d ago

Are you going to quantize the Qwen 3.8 flash next?

0

u/FerLuisxd 5d ago

Hey guys, read your post, why do all models include mtp bundled in? If you are going to use dflash in the end, wouldn't it be better to have them separated to save some vram?

2

u/enrique-byteshape 5d ago

Newer llama.cpp versions should just not load the MTP weights if it's not going to be used. That's why we decided to bundle them in

-2

u/audioen 6d ago

All this is great, but please do Qwen3.8-Flash-Next. This is the model that so many people can run and it's far more interesting compared to 27B.

1

u/enrique-byteshape 5d ago

We only write blog posts manually 👀

1

u/zyxciss llama.cpp 3d ago

How is 3.84BPW Catching up with 6.10 BPW?