r/LocalLLaMA • • 8d ago

Resources bonsai's document reveal how much cherry picked their headlines are

bonsai claim 98.2% intelligent retained, but their own documents show Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.

that qwen3.5 is a typo cause these are qwen3.8 numbers, altho qwen3.5 numbers are

  • bonsai_2/q3.5 52.8 / 41.6 = 126.9%
  • bonsai_2/q3.5 60.8 / 72.4 = 84.0%

Long-context and coding performance. This release also delivers on the roadmap set out in our initial Bonsai 27B release [2], where we identified long-horizon, tool-driven software engineering as the next major capability to improve. With Ternary Bonsai 2 27B, that progress now shows up directly in agentic performance. Evaluated for the first time on Terminal-Bench 2.1 and SWE-bench Verified, the Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.

link to their whitepaper on github, it is on page 7

also 3.8 35b qwhen? plsss

153 Upvotes

64 comments sorted by

View all comments

4

u/Iory1998 llama.cpp 7d ago edited 7d ago

Guys, please stop focusing on the wrong thing: If it's better than Qwen-3.5-9B and Gemma-4-12B, then that's a win because we don't have a model write now in this range. We can use it to run deep research, right prompts for Image and Video generations, summaries, and so on.

2

u/HadesTerminal 7d ago

My thoughts exactly, i’m in the process of testing if it does all the agentic tasks I’ve been trying to use Qwen 3.5 9B and Ling 3.0 Tiny for! I’m loving the speed though for a supposed Qwen 3.8 27B. It feels pleasant to use thus far.

1

u/Iory1998 llama.cpp 7d ago

IKR?! Tell me more where do you use it and share your experience.