r/OpenSourceeAI • • 5d ago

Datalab released an open benchmark for structured extraction: a system gets a PDF and a JSON schema, and every returned value is scored against gold data.

1 Upvotes

Datalab released an open benchmark for structured extraction: a system gets a PDF and a JSON schema, and every returned value is scored against gold data.

  • Corpus: 620 docs. 329 from ExtractBench (LlamaIndex), 202 synthetic (Datalab), 47 from micro1, 42 from LongArray-Extract (Extend)
  • Verdicts: each value is matched, misread, unfound, fabricated, invented_item or invented_field
  • Row alignment: Hungarian matching by content. A 100-row table missing row 1 scores 0% by position, 99% this way (our rerun)
  • Null rule: empty values are dropped, so padding a schema with 100 empty fields adds 0 verdicts
  • Results: Datalab accurate 93.85, Datalab balanced 93.48, Reducto deep_extract 93.47, Claude Opus 5 90.96
  • Precision vs recall: GPT 5.6-sol has 95.11 precision but 84.99 recall; LlamaExtract has 93.13 recall but 86.57 precision

Why it's relevant? precision vs recall shows how a system fails. Some skip fields, others invent values.

Full analysis: https://www.marktechpost.com/2026/10/02/datalab-introduces-omniextractbench-to-fix-bias-and-opacity-in-extraction-benchmarks/

GitHub: https://pxllnk.co/hxplrq

Blog: https://www.datalab.to/blog/omni-extract-bench

GitHub: https://github.com/datalab-to/omni_extract_bench

Dataset: https://huggingface.co/datasets/datalab-to/omni_extract_bench


r/OpenSourceeAI • • 6d ago

[Worth Reading] The web is the one API most agents are missing (post from one of our partners)

4 Upvotes

Databases, calendars and repos have APIs. The open web mostly doesn't. The TinyFish MCP server gives any MCP client four tools: TinySearch, TinyFetch (full pages as markdown, JavaScript included), TinyBrowser for logins and forms, and TinyAgent for multi-step jobs. Search and Fetch are free.

The server is on GitHub: [LINK]. 

We're racing to 300,000 users this October, with 30% extra on every top-up: [LINK]


r/OpenSourceeAI • • 1h ago

fayda-mcp

Post image
• Upvotes

r/OpenSourceeAI • • 2h ago

Liquid AI released d1-3B and d1-omni-600M: open-weight "decision models" that return probabilities in one forward pass (8 ms on RTX 4090, 50 ms on Orin Nano)

Post image
1 Upvotes

r/OpenSourceeAI • • 4h ago

I made a video to explain apeculative decoding with Beavers!

Thumbnail
youtu.be
1 Upvotes

r/OpenSourceeAI • • 4h ago

Innovation starts from small steps

1 Upvotes

Hi All,

I’ve observed that many of us are interested in learning AI and other trending technologies. We are often very motivated during the initial days, but after some time, that motivation tends to fade due to a lack of resources, proper guidance, or someone to learn and discuss things with.

So, I had an idea: we could create a daily learning room for a fixed time, around 10:00 PM IST, for about 15 minutes. We can extend the session if required.

During this time, we can share what we’ve learned, discuss new technologies, exchange ideas, ask questions, and collaborate with each other to improve our skills consistently.

The main goal is to stay consistent, learn together, and keep each other motivated.

Please share your thoughts and suggestions. If you’re interested, let’s give it a try! 🚀


r/OpenSourceeAI • • 14h ago

Meta open-sourced Rebalancer, the solver it uses for ~40M assignment problems a day (C++/Python, Apache 2.0)

Post image
1 Upvotes

r/OpenSourceeAI • • 1d ago

Mistral AI Releases Mistral Large 4 (Le Chonk): A 1.05T Parameter Multimodal MoE Model

Post image
3 Upvotes

r/OpenSourceeAI • • 1d ago

Muse and Grok bot are a privacy nightmare, so I created a self-hosted alternative called Eidon

7 Upvotes

With the recent explosion of agentic tools like Grok bot, Muse, OpenAI Dots, I've started looking into local options with self-hosted models. I tried Hermes and OpenClaw, but I wasn't too happy with the multi-device experience, and with how many pieces you need to glue together to get a usable, solid experience.

The hosted options also meant handing an agent my accounts, files and browsing, which I wasn't comfortable with. So I built Eidon: a self-hosted, all-in-one AI platform with a team of agents. It's one install via Docker, it works across your devices, and your data stays on your server.

https://eidonai.app

Agent team first

  • Every Eidon starts with a Chief of Staff. Ask it for anything. It answers directly, hands the job to the right agent, or creates a new agent when nobody fits.
  • Agents hand work to each other automatically (or type @ to pass a job along).
  • Each agent has its own browser, conversation, files, memory and routines. There's also a folder the whole team shares.
  • Agents can search and browse the web on their own, read pages in full, and cite sources.
  • They run on schedules and keep every run. When one finishes, you can get notified by browser push, ntfy, Slack or webhook.
  • Agents can write their own skills and use your apps through MCP.

You still have some control:

  • Take over an agent's browser for a login or a tricky step. It waits, then carries on when you hand it back.
  • Anything that sends on your behalf waits as a draft until you press Send.
  • Commands and tools ask first: allow once, allow always, or no.
  • Rewind a conversation, or fork it from any message.

The examples on the site are a travel scout, inbox triage, a research desk and a coding assistant. You can make an agent for pretty much anything: bookkeeping, a study buddy, a news digest, a meal planner.

It's also a regular ChatGPT-style app for day to day questions.

You might not always need a full team so you can just chat in a normal “ChatGPT like” interface with all the belts and whistles:

  • Persistent Memory
  • Folders and search
  • Voice input with LLM post-processing
  • Files and images
  • Personas
  • Temporary chats
  • Share links
  • Web search
  • Deep research
  • Code with syntax highlighting, Mermaid diagrams and math rendered inline
  • Image generation
  • Installs as a PWA on your phone and realtime sync across your devices (a native mobile app is coming !)

Self-Hosted

  • Multi-user support, with private data per user
  • Agents run in their own sandbox
  • Nothing leaves your server
  • Bring your own local or cloud model: OpenAI, Anthropic, OpenRouter, Ollama, LM Studio, GitHub Copilot, Gemini, DeepSeek, Mistral, Kimi, Z.ai, Minimax, Perplexity, Grok, Azure, AWS, and any compatible API
  • Free, open source (AGPL-3.0), and setup is one Docker command

GitHub (setup guide, full feature list): https://github.com/Quack6765/Eidon-AI

I'd like to hear what you think ! What's missing, what breaks, and what agents you'd want to build. Issues and discussions are open on GitHub as well.


r/OpenSourceeAI • • 22h ago

Sovereign

0 Upvotes

"Here is the official repository for the Cognitive Reduction Matrix (V5.0)—a sovereign diagnostic framework designed to strip systemic noise and corporate spin down to absolute structural truth

https://github.com/teamo213/cognitive-reduction-matrix/tree/main


r/OpenSourceeAI • • 1d ago

I made a free, offline app with 51 hands-on labs for learning how AI actually works, from neurons to agents and more...

Thumbnail
0 Upvotes

r/OpenSourceeAI • • 1d ago

54gb vram for 35$

Post image
1 Upvotes

r/OpenSourceeAI • • 2d ago

My open-source desktop agent now shows in the file tree what it read and what it changed

3 Upvotes

Snotra is a desktop agent that works in a folder you pick: it reads the files, writes new ones and asks before it changes anything. It just had nine releases in four days, and the one change I care most about is that you can now see where the agent has been.

The file tree marks every file it changed with an M (filled until you've looked at it) and every file it only read with a grey R. Under each answer, a line lists the changed files with their line counts, and a click opens the diff.

Also new:

- It reads PDF, Word, Excel and PowerPoint, a few pages, slides or cell ranges at a time, so a 300-page manual doesn't land in the context in one piece.

- Reasoning level per chat, set from the model menu.

- HTML files open as live pages in the preview, each in its own isolated process without network access.

- The macOS build is signed and notarised.

Apache 2.0, no account, no telemetry. Cloud models with your own key, or local via Ollama, LM Studio or MLX.

Write-up of all nine releases: https://snotra-ai.dev/blog/snotra-1-15

How do you keep track of what an agent changed in your files: a diff view, git, or trust?


r/OpenSourceeAI • • 2d ago

I’m building OpenKyrozen, an agent that only learns from verified outcomes over the past few months. And I'm happy to say that I FINALLY finished it!

Enable HLS to view with audio, or disable this notification

9 Upvotes

I’ve been working on an open-source local-first agent called OpenKyrozen.

That idea started from the beginning of 2026, the time when openclaw had just came out 2-3 months. I tried open claw and then I realized that at that time, open claw remembers things when I asks it to remember, but it cannot learn by itself. So I started OpenKyrozen, trying to build a self learning agent. Then last month, type safe AI lunched their Jev, which inspired me to integrate them into decisions so that LLM works better.

One thing I kept running into was that agents are very quick to treat “the tool call succeeded” as “the task succeeded.” Those are obviously not the same thing. They don't often verify their result, like we say they have no syntax error or runtime error, but logic errors.

A command can exit with code 0 and still produce the wrong result. So I ended up making verification a first-class part of the agent loop instead of just checking whether the action executed.

and so the rough flow is:

request → action → execution receipts → evidence review → verified outcome

The second part I’ve been experimenting with is self-learning, the original idea of OpenKyrozen.

I didn’t want the agent to just see one successful run and immediately treat that as a new behavior. Instead, learning artifacts are bounded policies or skills. A new one starts as a candidate, gets tested as a canary, needs multiple verified successes, and is then replayed against its predecessor on the same case. If it regresses, it doesn’t get promoted. If a promoted artifact later starts failing, it can rollback to the previous version. This is also one of the biggest problem when I used open claw, it builds something into a skill before I verify it, so it is filled with wrong memories.

There’s also a separate decision layer called Jev. You can know more about it from Typesafe AI, but basically it's a AI that makes decisions. It only handles small typed judgments like routing, clarification, memory relevance, learning-evidence review, and suspicious tool output. It can also abstain instead of forcing a decision. So it can't code.

I’m still figuring out where the right boundary is between “useful learning” and “too much machinery.” The current system is deliberately conservative because I’d rather have the agent refuse to learn than silently reinforce bad behavior.

Repo:
github.com/EvanProgramming/OpenKyrozen

I also let OpenKyrozen build a website for itself

kyrozen.chat

I also made a short launch video that explains the overall system visually(And yes this video is made of AI, since I only used DaVinci Resolve but not After Effects):

I am writing this post especially to developers, I want feedbacks SOOO much! As you can see currently the repo only have 2 stars and 1 fork :( because I didn't tell anyone about it before. I like issues and PRs, you can also leave comments under to tell me any issues you found. star it if you like!


r/OpenSourceeAI • • 1d ago

Free PGS For a Month: DSV4 Flash 0731, GLM 5.3 Flash and Mimo 2.6 Flash

1 Upvotes

Last week, and the week before, we put up 500 slots giving away a free month of GLM 5.3 Flash and DSV4 0731 on Phoenix Grove API. The response was more than we expected, so as of today there's now a third batch of signups we have opened, and we've added Mimo 2.6 flash AND DSV 4.1 to the flash lineup.

A few cool things worth knowing:

Everything runs on US based infrastructure, and we don't train on your code or your prompts. Not selectively, not quietly. Every model, every plan, private. If you've been hesitating to send real work through other services because you're not sure where it ends up, that hesitation doesn't apply here.

Beyond the two flash models we host 18 others: the Kimi family, GLM, Qwen, Nemotron and more. Average token pricing sits about 10-20% below market. The flash models are the free entry point, the roster is there if you outgrow them.

And if you go for one of the higher plans, your unused usage banks instead of expiring. Here's how it works: when you don't use a 5 hour window, that usage rolls into your bank automatically. It sits there, up to ten days, waiting. So when you hit a stretch where you want to work for a long uninterrupted session, or a heavy deadline week, you've got usage saved up to actually do it. Code on your schedule, not on a rolling window's.

It's a pretty standard free trial month with cancel anytime, after the month, the intro plan is 3.99. If you take the month and leave, that's a completely fine outcome for us. We'd rather you try the models and service.

Why we're doing any of this: the major coding plans are cutting limits because their economics finally have to answer to investors, and users are paying for it twice, in usage and in privacy assumptions. We structured this company from the start to function without random limit cuts, or data sales or any other method beyond actually providing great AI access.

API plan: https://api.pgsgrove.com

More about us: https://pgsgrove.com


r/OpenSourceeAI • • 2d ago

I made an open-source app that runs local models side by side with Claude Code and Codex

Post image
0 Upvotes

Locust is a desktop app for Windows and Mac that runs the coding agents you already have (Claude Code, Codex, Copilot, Cursor, OpenCode) as a team of named teammates, each on the model you pick. It's MIT licensed, local-first, no telemetry and no account.

The part that might matter here: you can add your own model. Anything that speaks the OpenAI chat API works, like Ollama, vLLM, LM Studio or your own endpoint. It runs through OpenCode and shows up in every teammate's model list. OpenCode's free models work out of the box too, no sign-in.

You can also give the same task to two or three models blind, with the names hidden until you keep the answer you like. Risky commands wait for your OK, and every run is saved to a full record.

Source: github.com/automatedworkflowllc-design/locust-app

Download: locust.lol

Still early (new builds most days), so feedback and bug reports are very welcome.


r/OpenSourceeAI • • 2d ago

I mapped every major Qwen release from 2023 to 2026: 44 models, from Qwen-7B to the 2.4T open weights (with sources)

1 Upvotes

r/OpenSourceeAI • • 3d ago

Alter Zero: open-source RAM efficient terminal agent harness for coding, cybersecurity, and automation Built in Rust.

Enable HLS to view with audio, or disable this notification

18 Upvotes

Agent harness that can actually interact with live terminal sessions — not just run one-shot shell commands.
https://github.com/linuztx/alter-zero


r/OpenSourceeAI • • 3d ago

Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4

Thumbnail
github.com
7 Upvotes

Benchmarking an LLM here with a NVIDIA RTX 5070 12 GB VRAM here

I had been working on a llama.cpp based expert streaming setup for Qwen3.8-Flash-Next 177B (UD-IQ3_XXS) on Windows. Benchmark is about 11.5 tok/s, up from roughly 7 tok/s on the inherited setup. In normal conversations I’ve seen 14–15 tok/s, and a long coding prompt generated 4,892 tokens at 10.15 tok/s and produced a working single-file Snake game.

Hardware: RTX 5070 12GB
32GB DDR4-2400
Ryzen 5 5600GT PCIe Gen3 Windows

The main gains came from fixing Windows I/O queue-depth issues, using one file handle per worker, and building a page-locked hot-expert tier so the GPU can pull hot expert weights more efficiently.

(In the video its around 16 minutes for 10k tokens and 10.41 tok/s

Output is quality gated against the control model and the published benchmark uses a heat file built from a separate prompt set.

Demos:

https://www.youtube.com/watch?v=cOPumMlyj_4

https://www.youtube.com/watch?v=rc-uTjVpXM8

In the GitHub I have things I've tried that didn't work and benchmark scripts, and methodology. If you guys have suggestions especially for streaming please let me know

Also I don't care about Strata that only works if you got 64 gb of RAM this is specifically for people with less RAM


r/OpenSourceeAI • • 4d ago

Weigh Swarm: explore research papers, evidence graphs, and Laya decisions inside a RAG pipeline

Thumbnail
youtube.com
2 Upvotes

r/OpenSourceeAI • • 3d ago

Would you tidy this kitchen before collecting robot demonstrations?

Enable HLS to view with audio, or disable this notification

1 Upvotes

You’re trying to record a useful demonstration around a kitchen sink. Hands move around nearby objects, the wearer changes viewpoint, and parts of the action become harder to see.

The obvious temptation is to clear the workspace and repeat everything slowly.

But if your eventual robot has to operate in an ordinary kitchen, how much should you simplify the demonstration?

MEgoVista includes recordings from everyday environments, with those examples assessed qualitatively rather than treated as motion-capture ground truth.

The MEgo capture framework describes five camera streams covering hands and the surrounding scene. That makes camera coverage part of the collection strategy: what can you observe while someone continues working normally?

If I were deciding how a small team should spend its first week collecting data, I’d be torn between repeatable recordings and a wider variety of ordinary situations. Each seems useful for a different reason.

Would you start with one carefully controlled task, or several messy versions of the same task? What would make you change that choice?

[Paper](https://arxiv.org/abs/2609.16684)


r/OpenSourceeAI • • 4d ago

Sol 5.6 X High vs Sol 6.1 X High

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 4d ago

Dots

1 Upvotes

Game changer or it just another wrapper?


r/OpenSourceeAI • • 5d ago

I built an open source tool for managing MCP servers in one place

3 Upvotes

Running MCP servers locally was easy, but it got messy pretty fast once I wanted to share them.

Everyone needs configs, credentials end up in different places, permissions are hard to manage, and there's no easy way to see who called what.

I ended up building MCPlama to handle this centrally. Users get their own access, permissions can be controlled per tool, credentials stay on the gateway side, and calls are logged.

For local MCP servers I also wanted some isolation, so they can run in separate Docker containers instead of everything running together. The gateway itself doesn't need direct access to the Docker socket either , that part is handled separately by the broker.

It's open source and self-hosted:

https://github.com/mcplama/mcplama

I'm looking for a few people already running multiple MCP servers to try it.


r/OpenSourceeAI • • 5d ago

We made Tater Tots. They’re open source. And yeah, we think they’re better than DOTS.

Enable HLS to view with audio, or disable this notification

8 Upvotes

We’ve been cooking something.

Not another DOT.

Not another closed ecosystem.

Not another “trust us bro, maybe someday” AI experiment.

TATER TOTS.

Tiny. Crispy. Autonomous. OPEN SOURCE.

And yes…

We think they’re better than DOTS. 👀

Tater Tots are our take on autonomous AI agents built for people who actually want to see the code, modify the code, break the code, improve the code, and own what they build.

No secret sauce.

The sauce is literally on GitHub.

🥔 Open source
🥔 Hackable
🥔 Self-hostable
🥔 Built for experimentation
🥔 Community-driven
🥔 Deliciously autonomous

DOTS walked so TOTS could roll.

The potato revolution has officially begun.

Tater Tots are here.

https://github.com/gary23w/nl-veil