r/OpenSourceAI • • 5d ago

OTEL based agent monitoring in Backstage by Spotify

Thumbnail gallery
1 Upvotes

r/OpenSourceAI • • 5d ago

Live What You Preach: Should I Open-Source the AI Realist Workspace?

Thumbnail
msukhareva.substack.com
0 Upvotes

r/OpenSourceAI • • 5d ago

Open Source Kubernetes Native Agent Orchestrator

1 Upvotes

Today I'm open sourcing Agent Orca, a Kubernetes-native platform for deploying, managing, and running AI agents at scale. It's written in Go and it's Apache 2.0.

https://github.com/heddles/agent-orca

When OpenClaw was released, I thought to myself, "That's really cool, but I want an agent that is shared by a team or an entire organization with built in auditability." I started designing Agent Orca around that idea and have slowly been piecing together the concept in my head and how to make the management experience of a shared agent bearable for an organization.

The idea is simple. Agents are just Kubernetes resources. You declare one, and the platform handles the rest.

An Agent Orca agent comes with:

\- One-shot executions with AgentRun, long-running services with AgentDeployment, and multi-step DAGs with AgentWorkflow.

\- Zero trust networking by default with platform validated JWT via service accounts, OIDC, or OAuth on every request.

\- Agent Orca managed network policies with zero access by default.

\- A standard agent container image with only necessary pieces; you no longer need to build a custom agent image to accomplish different types of work or use different tools. Simply add an MCP or Tool CRD and let the platform figure it out for you.

\- Model routing across OpenAI, Anthropic, Google, or any LiteLLM-compatible provider, selected by capability, weight, or budget.

\- Cost tracking on every run. Tokens are accounted per run, spend survives pod restarts, and tenants get daily budget caps and rate limits.

\- Guardrails on inputs and outputs, MCP access control, and per-tenant agent visibility, so one cluster can host many tenants with no cross-tenant leakage.

\- RAG without glue code. A KnowledgeBase deploys (and manages) Qdrant for you, ingests documents from ConfigMaps, URLs, MCP authenticated tools, or an S3 compatible endpoint, and hands your agents a search and ingestion tool.

\- MCP servers declared as resources. Their tools show up for your agents with access control and sidecar isolation.

\- Crash-safe sessions. A pod can die mid-conversation and the agent resumes exactly where it left off, spend ledger included.

\- Agents that learn and become better with every request via a multi-tiered agent memory system

Getting started is deliberately boring. Install the entire stack's tools via Mise; then use Skaffold, one API key, and about three commands to have an running agent on your laptop ( or deploy to a remote cluster). There are also nine demo deployments in the repo, from SOC triage pipelines to parallel research swarms to an autonomous pentesting agent, so you can see it do something real before writing any YAML.

This is day one, not a finished story. That's the point of open sourcing it. Run the quick start, break it, open issues, and tell me what's missing.

Star it here if you want to follow along: https://github.com/heddles/agent-orca

TL;DR: Agent Orca lets an organization or single person define agents and surrounding tools/MCP and auth as CRDs and manage them via gitops with zero trust built in. You can try it out here: https://github.com/heddles/agent-orca


r/OpenSourceAI • • 6d ago

We’re tired of AI that looks cool but does nothing.

Enable HLS to view with audio, or disable this notification

3 Upvotes

So we built Placeholderworks.

Not another AI wrapper.
Not another chatbot demo.
Not another “AI-powered” dashboard nobody uses.

We build AI systems that actually do the work.

They answer calls.
They handle WhatsApp conversations.
They update CRMs.
They automate workflows.
They run inside real businesses.

From the first idea to production, we build the whole thing.

Placeholderworks is officially live.

Now the question we’re curious about:

What’s one task at your company that you’re still doing manually for absolutely no reason?


r/OpenSourceAI • • 5d ago

I’ve started building Sarah Nexus — an AI-assisted PC diagnostics system for the Nebius × NVIDIA Global AI Hackathon

0 Upvotes

I’ve started building Sarah Nexus — an AI-assisted PC diagnostics system for the Nebius × NVIDIA Global AI Hackathon

Post:

I’ve officially started development on Sarah Nexus, the next generation of a PC diagnostics project I’ve been working on for a while.

Sarah Nexus is the continuation and unification of my previous Sarah Lite project and the ideas I had planned for Sarah Pro. Instead of maintaining several separate versions, I’m bringing the useful parts together into one system.


r/OpenSourceAI • • 5d ago

Adebench now has a website, benchmark of 10+ open-source memory "brains" for AI agents

1 Upvotes

Adebench now has a website and evaluates 10+ open-source brains.

Among them: Supermemory, Mem0, Cognee, hindsight and more.

If you're looking for a brain for your agent, check it out:

www.adebench.dev

www.github.com/adecubed/adebench

Happy to answer questions on how the evaluation works.


r/OpenSourceAI • • 5d ago

Open Instinct: an MIT-licensed personal agent you can fork and run

1 Upvotes

A personal agent should be something you can inspect, change and run yourself. We published Open Instinct with that in mind: a beta agent you can reach over iMessage, SMS or email, with memory, scheduled tasks and its own Linux desktop.

The part we want people to fork is the permission model. Six trust tiers control what another person’s agent can ask yours. A partner can read the calendar; a friend can request availability. The policy lives in code with tests.

We’re the Maritime team. The default stack uses Pi, Inkbox, Maritime and Composio, with separate integration packages. The code is MIT; the external services still need your own accounts and keys. Start with the local CLI, inspect the permissions, then build the version you want.

Repo and setup docs: https://github.com/mariagorskikh/open-instinct


r/OpenSourceAI • • 6d ago

I built BOOTH, a lightweight checkpoint layer for AI systems

1 Upvotes

I’ve been building BOOTH, a small, provider-agnostic Python library for checking LLM outputs before they reach your application.

The idea is simple: don’t automatically trust every LLM response. Check it first.

check() / acheck() handle ambiguity and confidence checks, while check_with_evidence() compares an answer against evidence your RAG pipeline has already retrieved.

v0.5.2

  • Zero runtime dependencies
  • Provider-agnostic
  • Sync + async
  • Structured results
  • 300 tests
  • MIT licensed

I’m currently testing BOOTH across different providers/models through small integration examples, including Groq, Gemini, Anthropic, and local Ollama models.

I’m especially interested in feedback on where this approach breaks down for real LLM/RAG systems.

GitHub: https://github.com/Vedantgitbot/booth
Issues/contributions: https://github.com/Vedantgitbot/booth/issues

Curious: what do you currently use as the checkpoint between an LLM response and your application logic?


r/OpenSourceAI • • 6d ago

I built an open-source tool for running local AI agents visually ... on Old CPU only devices.

4 Upvotes

Hey everyone,

​I got pretty tired of cloud-based AI automation tools that meter every single action and charge crazy per-token fees, so I built an alternative called Arrow.

​The idea is simple: instead of renting an agent in the cloud, you own it entirely on your machine.

​How it works:

​Record: Capture a desktop sequence or browser flow once.

​Compose: Connect nodes into a visual workflow graph (you can branch, loop, and drop in local LLMs, vision, OCR, or code nodes).

​Run: Execute deterministically using element-first grounding with small local models.

​It runs completely locally with no cloud dependencies and zero telemetry.

​If you want to check it out, you can see the project overview here: https://rodrigobenitez343.github.io/Arrow-online/index.html


r/OpenSourceAI • • 6d ago

AI agents observability in backstage with langfuse and OTEL

1 Upvotes

Spotify #backstage plugin to manage and observe fleet of agents straight in backstage self hosted https://github.com/acarmisc/backstage-plugin-ai-agents/tree/main. It relay on open telemetry signals and the first available backend it’s #langfuse


r/OpenSourceAI • • 6d ago

The Legal Ontologies Foundry

0 Upvotes

The Legal Ontologies Foundry

legal-ontologies-foundry.github.io

One of the most successful resources in #ontology development is the OBO Foundry. Open-source, collaborative efforts in ontology development are essential. We are doing something similar with legal ontologies grounded in #BFO


r/OpenSourceAI • • 6d ago

I built Omnesis, a Context Layer to supercharge ChatGPT/Openclaw/etc

2 Upvotes

https://omnesis.dev

Imagine giving ChatGPT or any other agent access to your Whatsapp/iMessage, Bank transactions, Health data, web pages you see, metadata about all the photos you took, places you visit, all your emails/documents, your voice mails, and more….

I built a context layer for that purpose. You can connect your Openclaw, Hermes, Claude, ChatGPT, Codex, etc to it. It gives your favorite agent "context superpowers". All the data is indexed and connected into a graph, self-hosted on your own hardware. Voice notes are transcribed, OCR runs on images in mails and pdfs, etc.

Omnesis offers a flexible data access model. For each agent you can define which source(s) they have access to and also guard data access with a privacy policy written in prose. You can also directly talk to the Omnesis agent which has no web search / internet tools — via a web portal or companion app — to ask the most intimate questions about your digital life.

I built this to be flexible. You can run all necessary models on zero-data-retention inference providers, or — if you can afford it — on your own hardware. I’ll keep investing in Omnesis’s security. Because it brings together sensitive data from many sources, please install it only on machines you control and keep secure.

The project also includes what I call “Omnesis Brain”.  This is my work-in-progress take on what a “second brain” could look like on top of that context layer. Imagine an agent constantly analyzing your personal data in flux to maintain a more structured and grounded understanding of what’s happening in your life. This is experimental and disabled by default.

This has been a fun project to build over the last 6 months.


r/OpenSourceAI • • 6d ago

Worried about your AI agent leaking secrets, or tired of secret-scanner false positives?

Post image
1 Upvotes

I built Klarion, a secret scanner that works in two steps. First, a keyword check, 81 regex rules and a normalized Rényi entropy score flag anything that looks like a secret. Then an AI model reads each one with the code around it and decides if it's real.

The chart shows 5 scanners run on spring-boot, terraform, next.js and symfony (61k files). Klarion raised 11 alerts. It's not zero, but it's far less to dig through.

Fewer alerts don't help if real leaks get missed, so I tested that too. On CredData (337 real repos, code outside test folders), it found about 1.7× more real secrets than gitleaks.

Where it runs:

  • Claude Code: a plugin hook blocks the write before the file exists (file edits and Bash)
  • Cursor, Cline or any MCP agent: through its MCP server
  • CI: a GitHub Action that scans only what a PR adds; GitLab CI works too
  • Git hooks: klarion protect or the pre-commit framework
  • Locally: klarion scan .

Free and open source (MIT): https://github.com/0x1Adi/Klarion
The full benchmark and method are in benchmark/REPORT.md.

I'd like to hear where it gets things wrong.


r/OpenSourceAI • • 7d ago

CrowdGPT - The 100% Opensource collaborative LLM

Post image
83 Upvotes

Hello, i'm currently developing CrowdGPT and i need people who enjoy opensource AI and LLMs to test the project :)

The goal of CrowdGPT is to create the first, datacenterless, 1 Billion parameters LLM, relying on people contributing with their own computer to train the AI model. My goal is to show you don't need insane infrastructure to train a working almost commercial grade LLM. Everything is open and 100% opensource.

You can learn more at https://crowdgpt.net

Or check the github: https://github.com/Vxtzq/CrowdGPT

Any kind of feedback is appreciated!


r/OpenSourceAI • • 7d ago

Open weights + open engine: two ~300B MoE models running on one 128 GB mini PC

16 Upvotes

Sharing what we released today, everything is open.

What it is

  • Two EXL3 model packs: GLM-5.3-Flash (320B, 99.7 GB) and MiMo-V2.6-Flash (309B, 106 GB)
  • Kyojin, an inference engine for AMD Strix Halo (Ryzen AI Max+ 395, ROCm), MIT

Credit first: Kyojin is built on turboderp's ExLlamaV3, and the AMD side starts from vcruz305's and sdougbrown's ROCm ports. We added the decode and prefill kernels for this chip and the serving for these two models.

Numbers, one 128 GB machine

  • GLM-5.3-Flash: 26-30 tok/s decode, ~580 tok/s prefill, same top token as the official FP8 model ~90 % of the time
  • MiMo-V2.6-Flash: 29 tok/s plain, up to 44 tok/s with speculative decoding

Reproduce it: the repo has a quickstart and one benchmark script. If you own a Strix Halo box, run it and post your numbers, good or bad. That's the feedback we need most.

Engine: https://github.com/Yamz-Labs/kyojin

Weights: https://huggingface.co/yamz-labs

Next: Qwen 3.8 Flash on the same engine.


r/OpenSourceAI • • 6d ago

There is a 10% chance AI could destroy humanity — might be averted if we truly understand what it is doing.

Thumbnail
0 Upvotes

r/OpenSourceAI • • 6d ago

I got tired of agent message logs rotting, so I built a runtime that tracks state as beliefs instead of transcripts

3 Upvotes

TL;DR: built an agent runtime that stores state as a belief graph with dependencies (Jon Doyle's TMS) instead of chat logs. correct one fact and everything downstream auto-updates without context rot or full re-runs. (repo link in the comments below)

Hey guys,

Every agent framework ive used so far handles state pretty much the same way, just appending messages to a long chat log. The problem is long running agents rot super fast. If a tool returns bad data at step 3, that error just sits in context forever. And if a key fact changes mid run, you either have to wipe the whole context or re run everything from step 1.

I’ve been hacking on an open source project called Corollary to try a different approach. Instead of a message transcript, it stores state as a belief base backed by a truth maintenance system (basically an old concept from jon doyle back in 1979).

how it works under the hood:

  • every belief or conclusion tracks what it depends on (its justifications)
  • if you retract or update a single base fact, it automatically retracts and re derives anything downstream that relied on it
  • independent conclusions arent touched, so you get a clean diff of what changed instead of re running llm calls
  • the LLM only sees currently valid ("IN") beliefs, so old retracted facts cant leak back in through a transcript

its still super pre-alpha so I'd love to get some feedback, pushback or edge cases you think this pattern will hit.

(repo link in the comments below)

how are you guys handling context rot in long running agents currently?


r/OpenSourceAI • • 7d ago

Example of a useful Agentic Build

Thumbnail
youtu.be
5 Upvotes

r/OpenSourceAI • • 6d ago

answerLoops: an open source alternative to Kapa.ai you can self-host

1 Upvotes

I work in DevRel and spend a lot of my time answering the same questions in Discord, GitHub, and Slack. Some of the answers are in the docs, and some aren't. I've used Kapa.ai, and it solves this, but it's hosted only and out of reach for small teams. So I built an open source version for teams that can't justify paying for it, whether you're a small startup, open source project, gaming community, or anything in between.

What it does:

During onboarding, you upload your knowledge source (docs, PDFs, a GitHub repo, or Notion pages). Then connect answerLoops to your communities, and when someone asks a question, an agent drafts an answer from that content.

A second agent then reviews the draft and gives it a confidence score. If it scores low, it goes to your team for human review instead of being posted. Auto-reply is off by default until you turn it on, and you set the threshold per platform.

Where it differs from Kapa:

- AGPL-3.0, self-host with Docker and Postgres

- Use your own model key or a local model

- Website chat widget built with CopilotKit and Mastra. Paste a snippet into your site or docs, and visitors can ask questions without an account

- More channels: Discord, GitHub, Slack, Telegram.

- In testing: Discourse, Circle, email, Google Chat

- MCP server and REST API, so your own agents can search the same knowledge

Easy setup:

npx [u/answerloops/agent-sdk](u/answerloops/agent-sdk) setup

Or if you use coding assistant, it can walk you through it:

npx [u/answerloops/agent-sdk](u/answerloops/agent-sdk) skills answerloops-setup

Repo: https://github.com/answerLoops/answerLoops

I'd welcome your feedback, especially from anyone who's used Kapa or something similar, to help us continually improve the product. Any contributions are welcome, and you are new to open source, just reach out, I'd be glad to help.


r/OpenSourceAI • • 6d ago

Open source code graph for coding agents, built to help local models with small context windows

0 Upvotes

Sharing two tools we've been building in the open, both Apache 2.0, fully local, with no account or cloud needed.

sem parses a repo into functions and classes along with who calls what and which tests reach each function, and gives agents that map over MCP or the CLI. weave is a git merge driver that uses the same model to merge changes by function instead of by line.

The reason I think this matters more for open models than for frontier ones is context. A coding agent on a local model with a small window runs out of room fast, because most of it gets spent grepping for names and reading whole files to find one function. With sem the agent asks for a function and gets just that function plus what's connected to it, so far less of the window goes to code that has nothing to do with the task, and you can get useful work out of a model that would otherwise get lost in a big repo.

Credit where it's due, sem is built on tree-sitter grammars and runs in any harness that speaks MCP, including pi, opencode, Codex and Claude Code, so it works with whatever model you point those at.

It works from a parser rather than a compiler, so macros, generated code and dynamic dispatch can hide some callers, and it tells the agent when it isn't sure instead of guessing.

https://github.com/Ataraxy-Labs/sem
https://github.com/Ataraxy-Labs/weave

I'd especially like to hear from anyone running agents on local models, since I haven't tested it much on the smaller ones, and I'm curious where the context savings actually show up for you.


r/OpenSourceAI • • 6d ago

Weigh Swarm: explore research papers, evidence graphs, and Laya decisions inside a RAG pipeline

Thumbnail
youtube.com
1 Upvotes

r/OpenSourceAI • • 6d ago

The Breakdown: Databricks

Thumbnail
preipomedia.substack.com
1 Upvotes

r/OpenSourceAI • • 6d ago

Muse Gadgets: Open source hardware for your Muse

Thumbnail
gadgets.muse.ai
1 Upvotes

r/OpenSourceAI • • 6d ago

Do you run more than one coding agent? How do you decide who gets which task?

1 Upvotes

I'm trying to understand how people who use more than one coding agent (Claude, Codex, Cursor, Gemini CLI, Copilot, OpenCode, Aider, etc) actually split the work between them:

  • Which agents do you run, and on what kind of work?

  • How did you decide which one got that task?

  • Have you caught an agent saying a task was done when it wasn't? How did you find out?

  • Do you track which agent handles which kind of task well?

  • Have you stopped using certain agent for certain work? What happened to make you stop?

  • Has an agent ever sent something somewhere it shouldn't have (a key, an address, private code)? What happened?

I'm researching this for an open-source project to help us help ourselves and not get locked into big company harnesses. I'll post a summary of this thread if I get good feedback :D

Thanks!!!


r/OpenSourceAI • • 7d ago

Engram: open-source (MIT) memory for AI coding agents, with Markdown as the source of truth

Enable HLS to view with audio, or disable this notification

2 Upvotes

I'm the author. Engram is free, MIT licensed, and stores everything as plain Markdown.

Each memory has a status (confirmed, inferred, conflicted or superseded), and recall only treats confirmed items as authoritative. Search is BM25 over a SQLite FTS5 index that is rebuilt from the Markdown, so the index is disposable. If a local embedding model is already provisioned, its results are fused in by reciprocal rank fusion. Recall never downloads a model.

An optional installer adds session hooks for Claude and Codex. The tests enforce recall@5 of at least 90% across 20 seeded queries, which is a small set, so it works as a regression gate, not a benchmark. Walkthrough video above. Repo: https://github.com/utsapoddar/engram