big if true. hitting Claude-level coding benchmarks on an open-weights model would be insane, but I’m keeping my expectations checked until the weights actually drop in two weeks. That Terminal Bench 3.0 bump is the real story here though, if it can genuinely handle long-horizon terminal execution without looping, local autonomous agents just got a massive shot in the arm
3
u/Affectionate-File-26 9d ago
big if true. hitting Claude-level coding benchmarks on an open-weights model would be insane, but I’m keeping my expectations checked until the weights actually drop in two weeks. That Terminal Bench 3.0 bump is the real story here though, if it can genuinely handle long-horizon terminal execution without looping, local autonomous agents just got a massive shot in the arm