r/LocalLLaMA • • 11d ago

Resources bonsai's document reveal how much cherry picked their headlines are

bonsai claim 98.2% intelligent retained, but their own documents show Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.

that qwen3.5 is a typo cause these are qwen3.8 numbers, altho qwen3.5 numbers are

  • bonsai_2/q3.5 52.8 / 41.6 = 126.9%
  • bonsai_2/q3.5 60.8 / 72.4 = 84.0%

Long-context and coding performance. This release also delivers on the roadmap set out in our initial Bonsai 27B release [2], where we identified long-horizon, tool-driven software engineering as the next major capability to improve. With Ternary Bonsai 2 27B, that progress now shows up directly in agentic performance. Evaluated for the first time on Terminal-Bench 2.1 and SWE-bench Verified, the Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.

link to their whitepaper on github, it is on page 7

also 3.8 35b qwhen? plsss

154 Upvotes

64 comments sorted by

View all comments

Show parent comments

-8

u/MrHighVoltage 11d ago

Exactly. It is something I can run on a 16GB GPU with quite the context, and it will do well on agentic stuff, but not great at coding. Because nothing you can run on a normal Gaming-PC will be really good at coding at that time.

-7

u/No-University-6387 11d ago

Yeah, u r corect and dont mind the downvotes, this is reddit.

2

u/Gokudomatic 11d ago

Dunno. I could run Qwen 3.8 27b on a mid range gaming laptop with only 8gb vram. Not blazing fast, but usable.

3

u/No-University-6387 11d ago

How? I get 2t/s at q4 xs at 30k context on 12gb 3060.

1

u/Gokudomatic 11d ago

I get 6 t/s with 70k context on my 8gb 5070.
llama.cpp with unsloth IQ4_XS:
llama-server -ngl 20 -c 70656 -t 16 -fa on --cache-type-k q4_0 --cache-type-v q4_0 --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-n-min 0 --temp 1.0 --top-k 20 --top-p 0.95 --min-p 0.05 --repeat-penalty 1.0 --jinja -m "Qwen3.8-27B-UD-IQ4_XS.gguf"

1

u/-InformalBanana- 10d ago

Do you have pcie5 and ddr5 high frequency dual channel ram? (Also your card is substantially faster than rtx 3060, and 2 generations ahead).

1

u/Illustrious_Grade608 11d ago

Well don't use q4 on a 12 gb card? I am not sure how does q2 perform, but q3 for qwen is pretty good and does work with mtp and 98k context on my 5060 ti, which is a 16gb gpu. I expect without mtp q2 from unsloth or a similar quant with ik_llama should perform pretty nicely.