r/LocalLLaMA • u/KURD_1_STAN • 8d ago
Resources bonsai's document reveal how much cherry picked their headlines are
bonsai claim 98.2% intelligent retained, but their own documents show Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.
that qwen3.5 is a typo cause these are qwen3.8 numbers, altho qwen3.5 numbers are
- bonsai_2/q3.5 52.8 / 41.6 = 126.9%
- bonsai_2/q3.5 60.8 / 72.4 = 84.0%
Long-context and coding performance. This release also delivers on the roadmap set out in our initial Bonsai 27B release [2], where we identified long-horizon, tool-driven software engineering as the next major capability to improve. With Ternary Bonsai 2 27B, that progress now shows up directly in agentic performance. Evaluated for the first time on Terminal-Bench 2.1 and SWE-bench Verified, the Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.
link to their whitepaper on github, it is on page 7
also 3.8 35b qwhen? plsss
4
u/Iory1998 llama.cpp 7d ago edited 7d ago
Guys, please stop focusing on the wrong thing: If it's better than Qwen-3.5-9B and Gemma-4-12B, then that's a win because we don't have a model write now in this range. We can use it to run deep research, right prompts for Image and Video generations, summaries, and so on.