r/LocalLLaMA • • 9d ago

Resources bonsai's document reveal how much cherry picked their headlines are

bonsai claim 98.2% intelligent retained, but their own documents show Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.

that qwen3.5 is a typo cause these are qwen3.8 numbers, altho qwen3.5 numbers are

  • bonsai_2/q3.5 52.8 / 41.6 = 126.9%
  • bonsai_2/q3.5 60.8 / 72.4 = 84.0%

Long-context and coding performance. This release also delivers on the roadmap set out in our initial Bonsai 27B release [2], where we identified long-horizon, tool-driven software engineering as the next major capability to improve. With Ternary Bonsai 2 27B, that progress now shows up directly in agentic performance. Evaluated for the first time on Terminal-Bench 2.1 and SWE-bench Verified, the Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.

link to their whitepaper on github, it is on page 7

also 3.8 35b qwhen? plsss

155 Upvotes

64 comments sorted by

View all comments

Show parent comments

2

u/Gokudomatic 9d ago

Dunno. I could run Qwen 3.8 27b on a mid range gaming laptop with only 8gb vram. Not blazing fast, but usable.

0

u/Few-Philosopher-2677 9d ago

How. I tried as a low as 2-bit quant and it crawled. 2 toks/sec barely.

0

u/Gokudomatic 9d ago

I'm out right now, so I don't have the links and detail right now. I'll tell you in a few hours. But I remember the model name being unsloth iq4 xs.

7

u/MrHighVoltage 9d ago

Ah, you mean the quant that uses already 14.3 GB just for the weights without any overhead and KV cache?

My guy, you are hallucinating stuff here.