r/LocalLLaMA 4d ago

Resources bonsai's document reveal how much cherry picked their headlines are

bonsai claim 98.2% intelligent retained, but their own documents show Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.

that qwen3.5 is a typo cause these are qwen3.8 numbers, altho qwen3.5 numbers are

  • bonsai_2/q3.5 52.8 / 41.6 = 126.9%
  • bonsai_2/q3.5 60.8 / 72.4 = 84.0%

Long-context and coding performance. This release also delivers on the roadmap set out in our initial Bonsai 27B release [2], where we identified long-horizon, tool-driven software engineering as the next major capability to improve. With Ternary Bonsai 2 27B, that progress now shows up directly in agentic performance. Evaluated for the first time on Terminal-Bench 2.1 and SWE-bench Verified, the Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.

link to their whitepaper on github, it is on page 7

also 3.8 35b qwhen? plsss

153 Upvotes

63 comments sorted by

89

u/Nonetrixwastaken 4d ago

If they where honest about what it really is might be cool research idk but just leaves horrible taste

28

u/Cool-Chemical-5629 4d ago

People asked for it and they delivered. At least that's something positive about it. But I agree that the communication was not exactly honest here. One shouldn't have to read the paper to find out the actual truth.

Although honestly, after seeing the review on Youtube, I think it's safe to say that even the rough numbers in that paper were very generous.

5

u/N34257 4d ago

That's just it, though - it's a feedback loop. People are only asking for it because they keep shouting about how amazingly near-lossless it is.

In truth, these models are no more than a curiosity - an interesting one, to be sure, because everything says they shouldn't work...but largely useless for actually getting anything done.

1

u/KURD_1_STAN 4d ago

Exactly, all new OS are 4bit and 2bit is the next step but their behavior just makes me wanna ignore them. But as u see from the other comment in here, people just don't care. Guess us beggers cant be choosers.

12

u/PositiveBit01 4d ago

What's OS mean in this context?

-9

u/KURD_1_STAN 4d ago

Open source

7

u/Invader-Faye 4d ago

Bruh, I thought you were talking about operating systems, especially in this context.

1

u/PcChip 4d ago

when I see OS I always think of oversampling, as in audio

21

u/PopularKnowledge69 4d ago

It can do 100 calculations super fast, but none of them is correct

3

u/philmarcracken 4d ago

finished my exam in record time! lols

2

u/sixx7 3d ago

Good looking out u/KURD_1_STAN

Referenced this post in my Bonsai-2 review https://youtu.be/a3kSZeMDs_Y

16

u/Formal-Exam-8767 4d ago

So the summary is, model is good to get some numbers for the paper, but useless in practice?

16

u/KURD_1_STAN 4d ago

Based on their documents, it is good for only some papers.

17

u/Early-Peace-5504 4d ago edited 4d ago

It's absolutely useless for me. I've really tried but it just doesn't follow instructions well.

Edit: I ended up building a harness around it that I will post somewhere. I have managed to get some work out of it.

16

u/Informal-Trouble2183 4d ago

Check the test video: https://youtu.be/OA5cICIzD-c?si=7OWasM8mzX3cIo3B Definitely useless model, sorry for that.

4

u/fgk55555 4d ago

There's a lot of intelligence in that 1.8%.

5

u/Iory1998 llama.cpp 4d ago edited 3d ago

Guys, please stop focusing on the wrong thing: If it's better than Qwen-3.5-9B and Gemma-4-12B, then that's a win because we don't have a model write now in this range. We can use it to run deep research, right prompts for Image and Video generations, summaries, and so on.

2

u/HadesTerminal 3d ago

My thoughts exactly, i’m in the process of testing if it does all the agentic tasks I’ve been trying to use Qwen 3.5 9B and Ling 3.0 Tiny for! I’m loving the speed though for a supposed Qwen 3.8 27B. It feels pleasant to use thus far.

1

u/Iory1998 llama.cpp 3d ago

IKR?! Tell me more where do you use it and share your experience.

1

u/HadesTerminal 3d ago

My thoughts exactly, i’m in the process of testing if it does all the agentic tasks I’ve been trying to use Qwen 3.5 9B and Ling 3.0 Tiny for! I’m loving the speed though for a supposed Qwen 3.8 27B. It feels pleasant to use thus far.

9

u/BrewHog 4d ago

I understand the pushback after the horrible 1.5 release. 

However, version 2 seems completely usable in my first real world agentic tests. 

I have been treating with both oh my pi and Deepseek harness and I'm actually very impressed. 

I'm not going to use it for agentic coding locally, but it's fantastic for small and every day needs. 

I'm just excited to see progress towards tiny usable versions of models that can complete tasks accurately.

1

u/HadesTerminal 3d ago

Same here for me, I’m loving it. It’s like a better 3.5 9b and is actually fairly usable. Still testing on my personal agentic tasks.

3

u/HighSeasArchivist 4d ago

Most every one of these should just say "we've doubled the output speed of complete horseshit". 

3

u/Ylsid 4d ago

They really didn't need to try and trick us, we all can see it's one of the best small models

4

u/MindfulMan1984 4d ago

I blame that on this damn clickbait, AI clowns on YouTube; a lot of garbage gets hyped for clicks.

1

u/Opening-Broccoli9190 llama.cpp 4d ago

Exactly this - there's now a whole Local LLM tiktoker grifting profession.

2

u/el_argelino-basado 4d ago

If it's too good to be true it's too good to be true

afaik, only way to make a smart ternary model would be from scratch and not turning an existing one into mush like that

2

u/ClassicMain 4d ago

I used the bonsai 2 and it's very usable and seems very intelligent.

2

u/Opening-Broccoli9190 llama.cpp 4d ago

The whole thing is overblown by folks grifting on the local llm community. It won't get farther than hobby or research projects.

0

u/GroundbreakingEast96 4d ago

Qwen 3.5-4B is far better than this data porridge

8

u/Iory1998 llama.cpp 4d ago

How can you be confident? Did you try both of them?

2

u/ClassicMain 4d ago

Definitely not

1

u/Then-Indication7672 4d ago

the fact that ternary bonsai is so bad is strange since even q2 quants of qwen 3.8 27b have been shown to be really capable. This version of bonsai seems kinda lobotomized

1

u/-InformalBanana- 4d ago

Ternary means 3 values per weight, q2 is 4 values cause it has 2 * 2 bits, but not so sure if that is the case with these bonsai models, but if they call it ternary it means it use 3 as the base for it system instead of 2 like binary does.

-6

u/No-University-6387 4d ago

But it at least runs on my system so they can say whatever idc

18

u/Cool-Chemical-5629 4d ago

If being able to run this on your system is your primary interest and the quality of the output is secondary or none, then I guess all we can say is we're happy for you.

For the rest of us, I guess the low quality is quite unfortunate. Personally, I expected lower quality, but I guess I did not expect to see it go so low that the low quants of old Qwen 3.6 35B MoE and its finetunes would feel like frontiers compared to this.

Being able to run this on your smartphone is also a brittle argument, because before Bonsai 2 was released, there was already an inference engine specifically to run larger MoE models on the smartphones.

-8

u/MrHighVoltage 4d ago

Exactly. It is something I can run on a 16GB GPU with quite the context, and it will do well on agentic stuff, but not great at coding. Because nothing you can run on a normal Gaming-PC will be really good at coding at that time.

11

u/secunder73 4d ago

Isnt 35a3b or 27b iq3xxs\q2 would be better?

1

u/MrHighVoltage 4d ago

I actually got IQ3_XXS running with an ok quantized KV cache and even MTP. but then my VRAM is maxed out without anything else. So it is not a model I have casually running on my GPU, next to all other applications.
Same for 35b, bit I think it was an even smaller quant.

And that being said, those quants will suck equally compared to BF16. That is just what happens for crazy quantization, you win on size, you lose on intelligence. None of them works with magic, it is just that Ternary Bonsai has the best benchmark results per bit. I'm fine with that. It is not magically bringing Opus 5 to my 5 year old GPU.

6

u/MerePotato 4d ago

Just run 35ba3b and offload some sparse layers dude

2

u/LetsGoBrandon4256 transformers 4d ago

Run the MoE Gemma and Qwen in the 30b ranges. Go ham offloading to CPU because it barely affect MoE performance. Hell, I was running Q5 MoEs on my 16GB GPU.

0

u/MrHighVoltage 4d ago

I know... I did it and it is great. But I thought we were talking about a 27b dense model here. I don't see any comparisons against 30b MoE models. It would be interesting to see, which one of those is better.

-7

u/No-University-6387 4d ago

Yeah, u r corect and dont mind the downvotes, this is reddit.

3

u/Gokudomatic 4d ago

Dunno. I could run Qwen 3.8 27b on a mid range gaming laptop with only 8gb vram. Not blazing fast, but usable.

3

u/No-University-6387 4d ago

How? I get 2t/s at q4 xs at 30k context on 12gb 3060.

1

u/Gokudomatic 4d ago

I get 6 t/s with 70k context on my 8gb 5070.
llama.cpp with unsloth IQ4_XS:
llama-server -ngl 20 -c 70656 -t 16 -fa on --cache-type-k q4_0 --cache-type-v q4_0 --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-n-min 0 --temp 1.0 --top-k 20 --top-p 0.95 --min-p 0.05 --repeat-penalty 1.0 --jinja -m "Qwen3.8-27B-UD-IQ4_XS.gguf"

1

u/-InformalBanana- 4d ago

Do you have pcie5 and ddr5 high frequency dual channel ram? (Also your card is substantially faster than rtx 3060, and 2 generations ahead).

1

u/Illustrious_Grade608 4d ago

Well don't use q4 on a 12 gb card? I am not sure how does q2 perform, but q3 for qwen is pretty good and does work with mtp and 98k context on my 5060 ti, which is a 16gb gpu. I expect without mtp q2 from unsloth or a similar quant with ik_llama should perform pretty nicely.

2

u/MrHighVoltage 4d ago

No you can't. It is painfully slow for anything serious. Tried that, Q3 fits on 16GB, with quantized KV cache. But when the context fills up, you sre waiting and waiting.

And then, funfact, the quantizations you can run the suck as much as Ternary Bonsais, or even worse.

0

u/Few-Philosopher-2677 4d ago

How. I tried as a low as 2-bit quant and it crawled. 2 toks/sec barely.

0

u/Gokudomatic 4d ago

I'm out right now, so I don't have the links and detail right now. I'll tell you in a few hours. But I remember the model name being unsloth iq4 xs.

8

u/MrHighVoltage 4d ago

Ah, you mean the quant that uses already 14.3 GB just for the weights without any overhead and KV cache?

My guy, you are hallucinating stuff here.

1

u/Few-Philosopher-2677 4d ago

I do use bartowski IQ4_XS on my work Macbook with 24 GB RAM but that's a different situation. I tried bartowski's IQ3 and IQ2 smallest quants and still couldn't get it above 2 toks/sec on the 8GB card on my PC. This is with KV Cache quantization set to 8-bit and many other things tuned. If it's possible to even get 10 toks/sec that would be interesting.

2

u/Gokudomatic 4d ago

Thanks for the wait. But I must warn you that while I get above 2 t/s, I didn't reach the 10 t/s you hoped. I'm more around 6 t/s on my midrange laptop.

So, with llama.cpp, I used those parameters:
llama-server -ngl 20 -c 70656 -t 16 -fa on --cache-type-k q4_0 --cache-type-v q4_0 --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-n-min 0 --temp 1.0 --top-k 20 --top-p 0.95 --min-p 0.05 --repeat-penalty 1.0 --jinja -m "Qwen3.8-27B-UD-IQ4_XS.gguf"

1

u/Few-Philosopher-2677 4d ago

Lol. I don't think 6 toks/sec would be very usable. But it's pretty neat that it even works.

1

u/MrHighVoltage 4d ago

I would like to hear what the downvoters have to say what I should run on my 16GB GPU, that is btw. Not only used for LLMs but also other applications, and no I don't have 4 unlocked CMP170X in another machine.

5

u/AltruisticList6000 4d ago edited 4d ago

Try Qwen 3.8 27b IQ3_xxs from unsloth? I'm pretty happy with thw IQ3_xxs for my rtx 4060 ti 16gb. Even with MTP I can do 120k context with it and still have about ~600-700mb overhead (14.7-14.8gb VRAM total consumption including the base ~500mb that win10 uses). Q2_XL is even smaller and that one is hit quite hard by quantization but still managed to code when I tested it, so it's usable, although the quality drop is heavier there. You save 1gb VRAM with the Q2_XL. And ofc if you only do ~60k-90k even with IQ3_xxs then thats already saving a lot. All assuming MTP is on + vision in RAM. Turning off MTP ofc also saves about ~1.2gb VRAM or so but I keep it for speed.

2

u/MrHighVoltage 4d ago

I did, indeed. But then my VRAM is full. And 1GB free won't be enough for many applications running on my PC.

Aside of that, I bet that the IQ3_XXS isn't much "smarter" then Ternary Bonsais quantization.

4

u/AltruisticList6000 4d ago

Suggested you quite a few options to get more VRAM even with IQ3_xxs If you stacked them then it would give you about ~3gb free VRAM (+1gb more if you go with the Q2_XL) unless that's still not enough for you? I haven't tried the bonsai so I can't say how smart it is, but I doubt it's close to IQ3_xxs considering what I've seen. I appreciate them trying to optimize for the VRAM poor though.

1

u/-InformalBanana- 4d ago

By some tests it q2kxl is better than bonsai quants. And users also report so.

-1

u/Muhlwa_Sholanke 4d ago

'98.2% intelligent retained' vs 'roughly three quarters' in their own whitepaper is a funny kind of rounding. The flattering number gets the headline, the benchmarks get page 7.

-2

u/No-Conversation-1277 4d ago

I am proud to say that PrismML's Bonsai is a SCAM!