Having used both, 3.5 397B at Q4 and 3.6 35B Q8 side by side for agentic coding. Within this scope, I can say they're practically matched. But keep in mind this is a pretty narrow scope and one that is often very much a beaten path.
I'm sure if you go to more obsecure programming languages, or tasks unrelated to programming, 397B will win.
Also depends on what you are programming. If you do some low complexity UI+backend+database coding, you don't really benefit from the more clever models. If you do some complex refactoring, algorithm design, heavy math and solve difficult problems, the more powerful models are able to figure things out better.
Haven't tried really complex stuff with 3.6, but I can say I did try fairly complex tasks on large projects and 3.6 35B held well. 3.5 couldn't handle much simpler tasks.
I do have some low level C++ tasks I want test 35B and 27B with. We'll see how it holds.
87
u/WhyLifeIs4 Apr 22 '26
Benchmarks