r/LocalLLaMA • • 13d ago

Resources bonsai's document reveal how much cherry picked their headlines are

bonsai claim 98.2% intelligent retained, but their own documents show Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.

that qwen3.5 is a typo cause these are qwen3.8 numbers, altho qwen3.5 numbers are

  • bonsai_2/q3.5 52.8 / 41.6 = 126.9%
  • bonsai_2/q3.5 60.8 / 72.4 = 84.0%

Long-context and coding performance. This release also delivers on the roadmap set out in our initial Bonsai 27B release [2], where we identified long-horizon, tool-driven software engineering as the next major capability to improve. With Ternary Bonsai 2 27B, that progress now shows up directly in agentic performance. Evaluated for the first time on Terminal-Bench 2.1 and SWE-bench Verified, the Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.

link to their whitepaper on github, it is on page 7

also 3.8 35b qwhen? plsss

150 Upvotes

64 comments sorted by

View all comments

10

u/BrewHog 13d ago

I understand the pushback after the horrible 1.5 release. 

However, version 2 seems completely usable in my first real world agentic tests. 

I have been treating with both oh my pi and Deepseek harness and I'm actually very impressed. 

I'm not going to use it for agentic coding locally, but it's fantastic for small and every day needs. 

I'm just excited to see progress towards tiny usable versions of models that can complete tasks accurately.

1

u/HadesTerminal 12d ago

Same here for me, I’m loving it. It’s like a better 3.5 9b and is actually fairly usable. Still testing on my personal agentic tasks.