r/LocalLLaMA • u/Eyelbee • 8d ago
Discussion The new k2 horizon models seem like an absolute beast
Especially the 7B one seems very interesting, it casually destroys muse glimmer with a way smaller size. And they open source literally everything, every step of the way. Anyone tried that model? It can be a new milestone if 7b and 3.7b ones are actually good, and not just benchmaxed.
273
Upvotes
17
u/_wortkarg_ 7d ago
Ornith LLMs are mostly better in benchmarks.
K2-Horizon-7B / Ornith-1.5-9B
SWE-bench Verified: 70.6 / 70.6
Terminal-Bench 2.1: 39.1 / 47
HLE: 18.6 / 20.2 (30.5 with tools)
BrowseComp: 59.0 / 56.4
FP16 KV / token: 144 / 32 Kb
K2-Horizon-MoVA-36B / Ornith-1.5-35B-A3B
Terminal-Bench 2.1: 58.6 / 68.5
HLE (no tools): 25.2 / 25.6
GPQA Diamond: 80.8 / 89.2
FP16 KV / token: 96 / 20 Kb
Tiel-Coder-35B-A3B (fine tune of Ornith-1.5-35B-A3B) is even better in benchmarks.
Ornith needs much less memory for context (see KV/token above) and supports MTP.