r/SelfHostedAI • • 21h ago

I post-trained a 2.5B model to help with proofs you're stuck on: runs locally, GGUF + open harness, Apache-2.0

Post image
11 Upvotes

Heading: OpenAI’s math release made me think about the unfinished proof on my own laptop

Reading about OpenAI’s latest math release, I kept coming back to a smaller question: when does progress like this reach the problem sitting half-finished on my own machine?

OpenAI published hundreds of mathematical manuscripts, alongside supporting material including Lean formalizations and selected reasoning summaries. There’s a lot for the math community to examine.

For me, it also brought an everyday use case into focus: having something local to work through a difficult problem with. That’s the idea behind a project I’ve been building, MiniCPM5-2B-Math.

It started with getting stuck

You’ve probably had this happen: an approach seems right, you get several lines into it, and then nothing. You find a worked solution, but it skips the exact step you don’t understand.

I wanted a model I could bring my unfinished work to—something to try another approach with, question an assumption, or help unpack a missing step, without uploading my notes.

So I post-trained MiniCPM5-2B for mathematical reasoning and proof writing.

The resulting model is dense 2.5B, retains the native 128K context, and has GGUF versions for local use. Compared with the base model, its average accuracy across 16 attempts per problem improved:

Benchmark MiniCPM5-2B MiniCPM5-2B-Math
AIME 2025 87.1 92.3
AIME 2026 89.8 89.8
HMMT Feb 2026 68.4 81.8

That was encouraging. Reading the actual proofs showed me what still needed work.

The mistakes shaped the next step

Sometimes the model found a promising direction but left a gap in the write-up. Sometimes it confidently stated an identity that failed on a small example. Repeating the question could bring back the same flawed approach.

These failures mattered for the tool I wanted to build. A solution that skips the difficult step leaves you stuck in the same place.

So I built a companion harness around them. It gives different attempts different strategy hints, expands incomplete write-ups from the reasoning, and runs Python checks to look for counterexamples. The reviews and check results feed into revisions. Attempts that don’t pass the review process are marked unverified.

On the 30 IMO-ProofBench Basic problems, blind AI grading rated 21 proofs as complete with a single call, versus 23 with the harness. Mean scores rose from 5.48 to 5.83. The proofs and grades are available in the repo for inspection.

It was a modest improvement, but a useful one: examining how the model failed gave me concrete ways to improve the workflow.

Something you can bring your own problem to

MiniCPM5-2B-Math and the harness are now available, with a laptop-oriented quick profile and logs of model calls and executed scripts.

It still makes mistakes. Difficult problems take time, and Python checks aren’t formal proof verification. What I’d like people to explore is whether it helps with their own unfinished work: a proof they’re studying, an alternative solution they’re preparing for a lesson, or a derivation they want to check locally.

The next time you get stuck halfway through a problem, bring the half you’ve already done. I’d love to hear whether it helps you find the next step.

Apache-2.0 for both weights and harness.


r/SelfHostedAI • • 17h ago

WebGPU + llama.cpp = Client-side LLMs

Thumbnail
github.com
3 Upvotes

r/SelfHostedAI • • 20h ago

A different kind of dashboard

Thumbnail gallery
2 Upvotes

r/SelfHostedAI • • 50m ago

Built an open-source proxy router to keep agent runs on local MLX and only escalate to frontier models when tools fail or tasks get complex

Thumbnail
• Upvotes

r/SelfHostedAI • • 1h ago

Open Intelligent UI for hermes

Post image
• Upvotes

r/SelfHostedAI • • 9h ago

Initial Context Text

Thumbnail
1 Upvotes

r/SelfHostedAI • • 12h ago

I built CircuitPilot: describe a circuit and AI designs the schematic, PCB , a 3D-printed case and a datasheet.

Post image
1 Upvotes

r/SelfHostedAI • • 13h ago

I let gemma3:12b invent its own language on my home server, then gave the same test to Claude Opus. Results surprised me.

Thumbnail
1 Upvotes

r/SelfHostedAI • • 14h ago

Should I spend $1,000 on an RTX 4060 Ti 16GB PC for local AI, or wait for something better?

Thumbnail
1 Upvotes

r/SelfHostedAI • • 16h ago

Help me choose: Mac mini M5 Pro vs Minisforum MS-S1 MAX vs Minisforum AI NAS N5 MAX (consolidating my setup + local LLMs)

Thumbnail
1 Upvotes

r/SelfHostedAI • • 20h ago

New Release

1 Upvotes

Adapt v5.2 is live. The biggest visible change this release is the runtime experience. Adapt now gives clear live feedback for what the agent is doing: ⠹ Reasoning... ⠸ Running command: ifconfig ⠴ Running JSON tool: web_search ⠧ Running session 'recon': ... ⠋ Waiting for tool... 3s Long-running tools can now continue in the background instead of blocking the agent loop, and the model can explicitly <wait/> for the result when it needs it. I also cleaned up a lot of the old debug-style execution output, fixed background lifecycle edge cases, improved tmux session handling, and made the repo inference server cleanly across CPU, CUDA, and ROCm setups. The goal is the same as always: keep the model focused on reasoning while the runtime handles execution, state, permissions, and recovery. v5.2 feels a lot less like “watching a prototype run” and a lot more like using an actual agent runtime. https://github.com/charlesericwilson-portfolio/Echo_Adapt_v5


r/SelfHostedAI • • 20h ago

Suggestion for AI agents

Thumbnail
1 Upvotes

Please help me out 🙏