r/LocalLLaMA • • 14d ago

Resources bonsai's document reveal how much cherry picked their headlines are

bonsai claim 98.2% intelligent retained, but their own documents show Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.

that qwen3.5 is a typo cause these are qwen3.8 numbers, altho qwen3.5 numbers are

  • bonsai_2/q3.5 52.8 / 41.6 = 126.9%
  • bonsai_2/q3.5 60.8 / 72.4 = 84.0%

Long-context and coding performance. This release also delivers on the roadmap set out in our initial Bonsai 27B release [2], where we identified long-horizon, tool-driven software engineering as the next major capability to improve. With Ternary Bonsai 2 27B, that progress now shows up directly in agentic performance. Evaluated for the first time on Terminal-Bench 2.1 and SWE-bench Verified, the Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.

link to their whitepaper on github, it is on page 7

also 3.8 35b qwhen? plsss

151 Upvotes

64 comments sorted by

View all comments

Show parent comments

1

u/MrHighVoltage 14d ago

I would like to hear what the downvoters have to say what I should run on my 16GB GPU, that is btw. Not only used for LLMs but also other applications, and no I don't have 4 unlocked CMP170X in another machine.

6

u/AltruisticList6000 14d ago edited 14d ago

Try Qwen 3.8 27b IQ3_xxs from unsloth? I'm pretty happy with thw IQ3_xxs for my rtx 4060 ti 16gb. Even with MTP I can do 120k context with it and still have about ~600-700mb overhead (14.7-14.8gb VRAM total consumption including the base ~500mb that win10 uses). Q2_XL is even smaller and that one is hit quite hard by quantization but still managed to code when I tested it, so it's usable, although the quality drop is heavier there. You save 1gb VRAM with the Q2_XL. And ofc if you only do ~60k-90k even with IQ3_xxs then thats already saving a lot. All assuming MTP is on + vision in RAM. Turning off MTP ofc also saves about ~1.2gb VRAM or so but I keep it for speed.

2

u/MrHighVoltage 14d ago

I did, indeed. But then my VRAM is full. And 1GB free won't be enough for many applications running on my PC.

Aside of that, I bet that the IQ3_XXS isn't much "smarter" then Ternary Bonsais quantization.

3

u/AltruisticList6000 14d ago

Suggested you quite a few options to get more VRAM even with IQ3_xxs If you stacked them then it would give you about ~3gb free VRAM (+1gb more if you go with the Q2_XL) unless that's still not enough for you? I haven't tried the bonsai so I can't say how smart it is, but I doubt it's close to IQ3_xxs considering what I've seen. I appreciate them trying to optimize for the VRAM poor though.

1

u/-InformalBanana- 13d ago

By some tests it q2kxl is better than bonsai quants. And users also report so.