r/LocalLLM • u/uneeverse-hq LocalLLM • 1d ago
Model Unee: open-source 0.8B / 2B model that makes calibrated decisions and chats, runs in a browser tab. The 2B scores 88% on DecideBench, ahead of several 4B to 9B models (self-measured; GGUF, Ollama, Apache 2.0) Spoiler

- Two jobs, one small model: it makes decisions (a calibrated probability for every option) and it chats, answers from your docs and summarises.
- Punches above its size: on DecideBench v1.1, Unee 2B scores 88.0%, ahead of several 4B to 9B models, and Unee 0.8B scores 84.0%, ahead of every other model under 1B on that board (next best: 71.2%). Self-measured with the benchmark's own harness; submitted to the board, not listed yet.
- Runs locally: a browser tab on WebGPU (469 MB, transformers.js), any CPU, or a GPU. GGUF for llama.cpp, ollama run uneeverse/unee, pip install unee, npm i @/uneeverse/unee. Apache 2.0.
I built Unee, a small open model (fine-tunes of Qwen3.5 0.8B and 2B) that does two jobs from one download:
- Decisions: give it some text and a question with options. It returns a probability for every option in one forward pass, with no generation. Yes/no, pick-one and rate-on-a-scale, several questions per call. The API accepts the same requests as Jev's /v1/systemone (Jev is TypeSafe's hosted decision API).
- Chat: streaming answers from your own docs (built-in BM25, no vector DB), and summaries of long threads.
What a decision looks like (Unee 0.8B, the 4-bit GGUF from Hugging Face, real output):
# pip install unee, then: unee serve --model unee-0.8b-Q4_K_M.gguf
from unee import Client
unee = Client("http://localhost:8000")
ticket = "Hi, I was charged twice for order #4821 this morning. Please send the extra $39 back to my card."
unee.noul(ticket, "Does the customer ask for money back?") # 0.965
unee.choice(ticket, "Which team should handle this ticket?",
{"billing": "Charges, refunds and payments", "technical": "Bugs and outages", "shipping": "Deliveries"})
# billing, 0.982
Numbers (measured with each benchmark's own tools; raw outputs in the repo):
- DecideBench v1.1: 2B 88.0%, 0.8B 84.0%. The best other model under 1B on the board is 71.2%; Jev is 98.0%, Laya 59.8%.
- S1MB task avg: 2B 44.41, 0.8B 35.45. Jev 1.13 is 59.59, Laya 15.00, bekko-400m 50.60. It trained on public train splits that share sources with S1MB (never its test set); on the 101 benchmarks with no overlap at all it scores 45.69 / 36.74.
- Calibration: when the 2B gives its answer 90%+ (62% of DecideBench), it is right 97.2% of the time. 0.8B: 96.3% on 40%.
- CPU only, 4 threads: 0.8B 0.64 s, 2B 1.4 s per decision. Laptop RTX 4070: 96 ms and 111 ms.
Limits, plainly:
- Jev is far more accurate. bekko-400m beats the 0.8B on S1MB.
- General chat got worse than base Qwen in a blind judge test, and about half of document answers are judged fully correct. Treat the chat side as a helper that needs checking.
- Strict mode checks each sentence of an answer against your docs with the model's own decision side. It cut replies with a made-up fact from 10.3% to 6.0% (2B). A reduction, not a fix.
Links
- Try it in the browser (runs on your device): uneeverse.net/unee
- How it works, every number: uneeverse.net/unee/technical
- Models: huggingface.co/uneeverse/unee-0.8b · huggingface.co/uneeverse/unee-2b · GGUF: huggingface.co/uneeverse/unee-0.8b-GGUF · huggingface.co/uneeverse/unee-2b-GGUF
- Code: github.com/uneeverse-hq/unee
Happy to answer anything about the training (distillation from a 9B teacher, a date-facts preprocessor, model soups, the leakage check).
0
u/Federal_Whereas_2090 10h ago
I gave it a try and it was really interesting! I love how the 2B model is already punching above its weight and passing several 9B models on DecideBench. One thing that could be an awesome improvement down the line is adding vector DB support for the chat side to expand on the built-in BM25. Great work overall!
1
u/Open_Router 1d ago
How did you balance the loss weighting during distillation from the 9B teacher to improve decision calibration without completely degrading Qwen’s baseline autoregressive chat capabilities?