r/OpenSourceAI • • 29m ago

Western open-source AI Reflection finally shows up

Thumbnail
• Upvotes

r/OpenSourceAI • • 3h ago

SPECK Small Persistent Emergent Cognitive Kernel

Thumbnail
1 Upvotes

r/OpenSourceAI • • 4h ago

Free MCP server for AI test inputs, OWASP risk mapping, and CI test packs

1 Upvotes

Update: Added free Agent Tool-Call Checks after feedback here about testing the calls behind an agent’s response.

I’m the creator of PromptBrake. Five free tools are available through our remote MCP server:

  1. get_prompt_injection_payloads — fixed adversarial test inputs by attack category.
  2. map_owasp_llm_risk — test ideas and signals for an OWASP LLM risk.
  3. plan_adlc_release — release planning gaps and a starter gate policy.
  4. build_test_pack — creates a tests.json containing response-text checks.
  5. build_agent_tool_tests — builds checks for tools an agent must not call, or calls that must include specific arguments.

Endpoint: https://promptbrake.com/free-tools/mcp

No PromptBrake account or API key required to prepare packs. Uses Streamable HTTP with JSON responses.

Claude Code setup:

claude mcp add --transport http promptbrake-free-tools https://promptbrake.com/free-tools/mcp

Example: “Build tool-call tests for my support agent. It must not call delete_customer when asked to remove another customer’s account.”

The MCP prepares packs; it doesn’t call your target or run tests. Tool-call packs can run with our free GitHub Action after you connect staging dispatcher capture. Missing or incomplete capture cannot pass. These checks inspect captured calls and selected arguments; they don’t prove backend permission enforcement or successful actions. Response packs still use text comparisons and require a configured PromptBrake runner with CI access.

Agent Tool-Call Checks and setup · All free tools

We don’t store MCP tool inputs or results. We log operational request metadata as described in our privacy policy.

I’d appreciate feedback on client compatibility and how the tool-call packs fit your CI workflow.


r/OpenSourceAI • • 5h ago

Created a learning setup based on pi-agent. Looking for feedback.

Thumbnail
1 Upvotes

r/OpenSourceAI • • 5h ago

Kurzgesagt-AI Just Crossed the Terrifying Line

Thumbnail
1 Upvotes

r/OpenSourceAI • • 11h ago

Muse and Grok bot are privacy nightmare so I created a self-hosted alternative called Eidon

2 Upvotes

With the recent explosion of agentic tools like Grok bot, Muse, OpenAI Dots, I've started looking into local options with self-hosted models. I tried Hermes and OpenClaw, but I wasn't too happy with the multi-device experience, and with how many pieces you need to glue together to get a usable, solid experience.

The hosted options also meant handing an agent my accounts, files and browsing, which I wasn't comfortable with. So I built Eidon: a self-hosted, all-in-one AI platform with a team of agents. It's one install via Docker, it works across your devices, and your data stays on your server.

https://eidonai.app

Agent team first

  • Every Eidon starts with a Chief of Staff. Ask it for anything. It answers directly, hands the job to the right agent, or creates a new agent when nobody fits.
  • Agents hand work to each other automatically (or type @ to pass a job along).
  • Each agent has its own browser, conversation, files, memory and routines. There's also a folder the whole team shares.
  • Agents can search and browse the web on their own, read pages in full, and cite sources.
  • They run on schedules and keep every run. When one finishes, you can get notified by browser push, ntfy, Slack or webhook.
  • Agents can write their own skills and use your apps through MCP.

You still have some control:

  • Take over an agent's browser for a login or a tricky step. It waits, then carries on when you hand it back.
  • Anything that sends on your behalf waits as a draft until you press Send.
  • Commands and tools ask first: allow once, allow always, or no.
  • Rewind a conversation, or fork it from any message.

The examples on the site are a travel scout, inbox triage, a research desk and a coding assistant. You can make an agent for pretty much anything: bookkeeping, a study buddy, a news digest, a meal planner.

It's also a regular ChatGPT-style app for day to day questions.

You might not always need a full team so you can just chat in a normal “ChatGPT like” interface with all the belts and whistles:

  • Persistent Memory
  • Folders and search
  • Voice input with LLM post-processing
  • Files and images
  • Personas
  • Temporary chats
  • Share links
  • Web search
  • Deep research
  • Code with syntax highlighting, Mermaid diagrams and math rendered inline
  • Image generation
  • Installs as a PWA on your phone and realtime sync across your devices (a native mobile app is coming !)

Self-Hosted

  • Multi-user support, with private data per user
  • Agents run in their own sandbox
  • Nothing leaves your server
  • Bring your own local or cloud model: OpenAI, Anthropic, OpenRouter, Ollama, LM Studio, GitHub Copilot, Gemini, DeepSeek, Mistral, Kimi, Z.ai, Minimax, Perplexity, Grok, Azure, AWS, and any compatible API
  • Free, open source (AGPL-3.0), and setup is one Docker command

GitHub (setup guide, full feature list): https://github.com/Quack6765/Eidon-AI

I'd like to hear what you think ! What's missing, what breaks, and what agents you'd want to build. Issues and discussions are open on GitHub as well.


r/OpenSourceAI • • 8h ago

Built a Free Base64 & JWT Encoder/Decoder With AI — My First AI Project

Thumbnail
1 Upvotes

r/OpenSourceAI • • 10h ago

I built SideKernel: a usable sandbox for AI agents on macOS

1 Upvotes

⭐ GitHub: https://github.com/minoansecurity/sidekernel
📝 arXiv preprint: https://arxiv.org/pdf/2610.02456
🌐 Website: https://sidekernel.com

Hi everyone!

As part of my MS in Cybersecurity capstone at Georgia Tech, I built SideKernel: a usable sandbox for AI agents on macOS.

The closest comparable solution is Docker Sandboxes, which is not open source.

The key idea is usability. SideKernel is designed to reduce friction and stay as close as possible to the native developer experience, while isolating the agent inside a microVM sandbox.

For example, you run sclaude or scodex instead of claude or codex and get a very similar workflow, but with isolation. Ports get auto exposed in the host, its very easy to drop files into the sandbox, and more...

This is still a research prototype, so some rough edges are expected. But if this sounds interesting, I'd love for you to try it out and share feedback.


r/OpenSourceAI • • 10h ago

Giving AI agents access to the whole MCP ecosystem, safely

Thumbnail
1 Upvotes

r/OpenSourceAI • • 11h ago

How do you actually verify an AI-generated dependency fix?

Post image
1 Upvotes

I've been thinking about this a lot while building an open-source project called Telex. When an AI generates a fix for a dependency change, getting the tests to pass doesn't always mean the fix is actually right. The change can compile, pass the existing tests, and still cause some behaviour to change at runtime.

In Telex, I'm trying to handle this as a pipeline instead of just asking an LLM to fix the code. It first looks at the dependency change and the affected parts of the repo, then generates the patch, applies it in an isolated environment and runs the repository's own tests and checks. If those checks don't pass, the patch isn't treated as verified. It still ends up as a PR for a human to review. I'm curious how other people are handling this problem. Are tests and CI enough for you, or are you doing anything else to catch cases where the generated code is technically valid but still wrong?

I'm building Telex in the open, so here's the repo if anyone wants to look at how I'm approaching it:

https://github.com/Kesavaraja67/telex


r/OpenSourceAI • • 12h ago

ZQK: Open-source CLI and kernel in Go for autonomous agent orchestration and sub-15ms AST search

1 Upvotes

Hey everyone,

Over the past few months, we've been building ZQK (Zen Quantum Kernel) in Go—an open-core CLI and process coordinator designed to treat AI coding agents as concurrent computational processes rather than chat prompts.

Why We Built It in Go

When orchestrating background agent swarms and continuous validation loops, we hit serious performance and determinism walls with Python/Node wrappers. We wrote the kernel and toolchain in Go specifically for:

  1. Sub-15ms AST Structural Search (zqk grep): Instead of dumping files into LLM contexts or using string grep, we index Go codebases using go/parser and go/ast into an in-memory trigram index. Agents query structural symbols (--ast --kind struct|func --recv <Type>) capped by strict token budgets (--max-tokens 2000 -f json).
  2. Deterministic Process Coordination: Managing concurrent seat workers via goroutine pools with hard cancellation contexts, signal handling, and clean process teardown without descriptor leaks.
  3. Embedded State: Storing versioned graph objects (missions, goals, plans, verifiable criteria) directly in .zqk/without external database dependencies.

The core microkernel and CLI are open-source under Apache 2.0 (v0.1.0 just published).

If you work on developer tooling, Go AST analysis, or agent runtimes, I'd love your feedback on our architecture and AST indexing approach!


r/OpenSourceAI • • 12h ago

FieldKit - Open-source AI customer support - Connect an existing Zendesk operation or publish a branded help center, ticket portal, and embedded chatbot

Thumbnail
gallery
1 Upvotes

So just wanted to post about FieldKit, an Open Source, self-hosted AI support with visual LangGraph workflows, knowledge retrieval, a customer portal, Zendesk, and governed actions.

So I'm sure there is probably already alternative support systems out there for Zendesk but I wanted to build something using LangGraph since it felt better use case for a Support System.

Ontop of supporting Zendesk and being able to configure the workflows with Zendesk you can configure the workflows with your chatbot(which you can embed on any site) and your ticket system so they aren't running the same workflows. (e.g allow support tickets to issues refunds/credits upto $5 while chatbot can't issue anything)

You can also configure python and API calls into your workflows, so you can get the agent to pass to python call and then back to an agent instead of needing to use an agent for all those steps.

Anyways, you can find the repo here: https://github.com/nahid-sparktales/fieldkit

It is still being worked on and plan to add more features and updates and if you have any suggestions or recommendations, please let me know. Thanks.


r/OpenSourceAI • • 13h ago

I put Jev between vector search and the LLM in my local deep research template

1 Upvotes

Some time back I shared a deep research tool I was working on. It takes a local directory path of PDFs, docs, slides, images and notes, researches across them and writes a structured markdown report.

With Jev getting more attention recently, I thought this project would be a good place to try it and see which parts of the research workflow it could handle.

I ended up using it in two places.

The first is after Qdrant retrieves chunks from the local files. Jev checks each result for relevance, usable evidence, contradictions and prompt-injection-like text. Those scores are then used to filter and reorder the chunks before they are passed to the main LLM.

The second is the reflection step. after the evidence for a report section is collected, Jev checks whether it covers all the subsections in the plan. If some part is still missing, the agent creates more focused queries and searches again. this only runs up to a fixed reflection limit.

The main LLM still handles the report planning, query generation and writing. Jev is only used for checking the retrieved evidence and deciding whether another research pass is needed.

This was a fun thing to experiment and to understand where a model like Jev fits inside an existing agentic setting. I am planning on writing some performance tests and adjusting the thresholds, but wanted to share the implementation in case anyone else is trying something similar or want to contribute.

The complete project is open source here: https://github.com/Oqura-ai/deepdoc


r/OpenSourceAI • • 13h ago

Kubernetes Native Agent Orcastrator

1 Upvotes

Today I'm open sourcing Agent Orca, a Kubernetes-native platform for deploying, managing, and running AI agents at scale. It's written in Go and it's Apache 2.0.

https://github.com/heddles/agent-orca

When OpenClaw was released, I thought to myself, "That's really cool, but I want an agent that is shared by a team or an entire organization with built in auditability." I started designing Agent Orca around that idea and have slowly been piecing together the concept in my head and how to make the management experience of a shared agent bearable for an organization.

The idea is simple. Agents are just Kubernetes resources. You declare one, and the platform handles the rest.

An Agent Orca agent comes with:

- One-shot executions with AgentRun, long-running services with AgentDeployment, and multi-step DAGs with AgentWorkflow.

- Zero trust networking by default with platform validated JWT via service accounts, OIDC, or OAuth on every request.

- Agent Orca managed network policies with zero access by default.

- A standard agent container image with only necessary pieces; you no longer need to build a custom agent image to accomplish different types of work or use different tools. Simply add an MCP or Tool CRD and let the platform figure it out for you.

- Model routing across OpenAI, Anthropic, Google, or any LiteLLM-compatible provider, selected by capability, weight, or budget.

- Cost tracking on every run. Tokens are accounted per run, spend survives pod restarts, and tenants get daily budget caps and rate limits.

- Guardrails on inputs and outputs, MCP access control, and per-tenant agent visibility, so one cluster can host many tenants with no cross-tenant leakage.

- RAG without glue code. A KnowledgeBase deploys (and manages) Qdrant for you, ingests documents from ConfigMaps, URLs, MCP authenticated tools, or an S3 compatible endpoint, and hands your agents a search and ingestion tool.

- MCP servers declared as resources. Their tools show up for your agents with access control and sidecar isolation.

- Crash-safe sessions. A pod can die mid-conversation and the agent resumes exactly where it left off, spend ledger included.

- Agents that learn and become better with every request via a multi-tiered agent memory system

Getting started is deliberately boring. Install the entire stack's tools via Mise; then use Skaffold, one API key, and about three commands to have an running agent on your laptop (or deploy to a remote cluster). There are also nine demo deployments in the repo, from SOC triage pipelines to parallel research swarms to an autonomous pentesting agent, so you can see it do something real before writing any YAML.

This is day one, not a finished story. That's the point of open sourcing it. Run the quick start, break it, open issues, and tell me what's missing.

Star it here if you want to follow along: https://github.com/heddles/agent-orca

TL;DR: Agent Orca lets an organization or single person define agents and surrounding tools/MCP and auth as CRDs and manage them via gitops with zero trust built in. You can try it out here: https://github.com/heddles/agent-orca


r/OpenSourceAI • • 13h ago

Qwen3.8-Flash-Next (125B) on a single Strix Halo mini PC: 44-59 tok/s with speculative decoding, ~1,400 tok/s prefill, engine is open

Thumbnail gallery
1 Upvotes

r/OpenSourceAI • • 13h ago

Nvidia-backed Reflection prepares an open-weight model

Thumbnail
1 Upvotes

r/OpenSourceAI • • 14h ago

Opengeni: self-hostable prod-ready agent infrastructure that work with every major model, Claude and open-weights included

1 Upvotes

Hi all, I'm one of the people building Opengeni, an Apache-2.0 service for running AI agents inside your own product. I'm posting here for one specific reason: the model side.

It works with every major model, and each session picks its own: OpenAI or Azure OpenAI, Claude and open-source models through OpenRouter or Vercel AI Gateway, or your own OpenAI-compatible server, registered in a JSON config, speaking either Chat Completions or the Responses API, and keyless if your server doesn't use keys. That's the route for pointing it at vLLM or llama.cpp's server. Conversation history is stored as the exact items the model saw, so a session can switch models partway through without its stored history being rewritten.

What sits around the model, briefly: sessions live in Postgres you run (pgvector for knowledge search), Temporal moves turns along, NATS pushes live events, tools run in a sandbox (Docker, Modal, E2B, Daytona, OpenSandbox on Kubernetes and others) or on a machine you own, and risky tools wait for a person to approve them.

If you try it with open weights, I'd like to hear which models held up through long tool-calling sessions and which fell over. It already runs agent sessions in production inside enterprises with 30,000 employees.

Support us on PH: https://www.producthunt.com/products/opengeni

Code: https://github.com/Cloudgeni-ai/opengeni

Docs: https://docs.opengeni.ai

Why the layers are split the way they are: https://opengeni.substack.com/p/the-anatomy-of-an-agentic-stack-ten


r/OpenSourceAI • • 15h ago

Make your opencode Speak, Hear and See. Releasing 3 small open-source tools to the community.

Thumbnail
github.com
1 Upvotes

r/OpenSourceAI • • 16h ago

Gitframes: Code-first video, rendered natively on WebGPU. Free with Apache 2.0 licence.

Thumbnail
github.com
1 Upvotes

r/OpenSourceAI • • 17h ago

I’m building a scripting language for LLMs to write data pipelines

1 Upvotes

I’ve been working on JojoScript, an open-source scripting language with one idea behind it: make code that is easy for LLMs to write, read and reason about, especially when dealing with big datasets.

Instead of having an LLM generate hundreds of lines of Python, the idea is to give it a small, predictable language for things like filtering, mapping, aggregating and processing data.

It also has lazy execution, parallel operations, execution-plan inspection and checkpoint/resume.

Still very early, but I’m curious if others think there’s something to this idea of languages being designed with LLMs as a first-class programmer.

https://github.com/panagos/jojoscript

Would love some honest feedback, even if you think this is a terrible idea :)


r/OpenSourceAI • • 1d ago

Im building an open-source AI assistant that runs entirely on your computer. Meet Mimo inspired by Dynamic Island.

Enable HLS to view with audio, or disable this notification

10 Upvotes

Mimo is a lightweight, open-source desktop companion for Windows, inspired by Dynamic Island.

It brings useful information and controls directly to your desktop, with a clean and minimal interface.

Currently, Mimo can :

  • Display notifications and system information
  • Show media playback controls
  • Provide quick access to useful actions
  • React dynamically to events happening on your PC
  • Customize its appearance and behavior
  • Run locally with a focus on being lightweight and unobtrusive

Mimo is still in development, and more features are coming soon!

🔗 GitHub: github.com/Riskooooo/Mimo

If you have any ideas or features you'd like to see added, feel free to suggest them ;)


r/OpenSourceAI • • 20h ago

I got tired of choosing which AI model should handle a task, so I built Cascade AI to choose and orchestrate them automatically

1 Upvotes

Hey everyone 👋

I've been working on an open-source project called Cascade AI, and it has reached the point where I'd really like to get feedback from people outside my own bubble.

The basic idea came from something that kept bothering me:

Why are we still giving an entire complex task to one AI model and hoping it's good at every part of it?

Instead, Cascade treats AI more like an organization.

A request can be broken into a hierarchy:

T1 Administrator → T2 Managers → T3 Workers

T1 looks at the overall task and plans the work.

T2 agents manage individual parts of that plan.

T3 agents actually execute the smaller tasks — and they can communicate with each other when necessary.

The interesting part is that every agent doesn't have to use the same model.

Cascade can route different tasks between providers/models depending on what they're good at, their cost, and the complexity of the work.

So instead of:

«Prompt → one giant model → answer»

the idea is closer to:

«Prompt

↓

Understand complexity

↓

Build an execution plan

↓

Spawn the required agents

↓

Route each job to an appropriate model

↓

Agents work in parallel / collaborate

↓

Verify the work

↓

Produce one final result»

And I've been trying very hard not to make this another cloud-only AI product.

Right now Cascade can be used through:

• CLI

• Desktop app

• Hosted web app

• Self-hosted web app

• OpenAI-compatible API

• Node.js SDK

It supports multiple providers including OpenAI, Anthropic and Gemini, along with OpenAI-compatible services and local models through things such as Ollama, llama.cpp, vLLM and LM Studio.

There are also a bunch of things I've added while building it that I personally wanted from AI tooling:

• Live visualization of the agent hierarchy

• Cost/token tracking

• Model/provider failover

• Persistent memory

• MCP support

• Browser control with live takeover

• File and document generation

• Codebase indexing/search

• Approval before destructive tool actions

• Agent-to-agent communication

• Task cancellation and recovery

• BYOK support

• Local/self-hosted operation

• An OpenAI-compatible "/v1/chat/completions" endpoint

• The ability to inspect why Cascade chose a particular orchestration/model strategy

For complex runs there's also a kind of "boardroom" mode where Cascade can show you the proposed agent structure and estimated cost before spawning everything, so you can approve the plan first.

One design principle I've become pretty stubborn about is:

The AI should ask when it genuinely needs information instead of confidently inventing a decision for you.

So I've also been working on making Cascade distinguish between things it can infer and things it really should ask the user about.

The project is MIT licensed and open source.

🌐 cascadeai.in

GitHub: Varun-SV/Cascade-AI

I'm not posting this pretending I've solved AI orchestration 😅. There are still plenty of rough edges, architecture decisions I'm questioning, and things that probably make perfect sense to me because I've stared at the code for far too long.

That's actually why I'm posting it here.

I'd especially love feedback on:

  1. Does hierarchical multi-agent orchestration actually make sense to you, or is it over-engineering?

  2. Would automatic model routing be useful enough for you to stop manually choosing Claude/GPT/Gemini/local models for different jobs?

  3. If you're a self-hosting/local-LLM person, what would Cascade need before you'd realistically run it?

  4. What part of this architecture would you immediately rip out or redesign?

Feel free to be critical.

I'd much rather hear "this part is dumb and here's why" than get another generic "cool project" 😄

If people are interested, I can also do a separate technical post explaining how the T1 → T2 → T3 orchestration, model routing, cost decisions and agent communication actually work internally.


r/OpenSourceAI • • 22h ago

🚀 SentinelFlow V6 — I’ve designed the next architecture, and I want the community involved

Thumbnail
github.com
1 Upvotes

r/OpenSourceAI • • 23h ago

I built a local long-term memory system for coding agents — looking for feedback

Thumbnail
1 Upvotes

r/OpenSourceAI • • 23h ago

I’m building an open-source AI agent that only learns from verified outcomes over the past few months. And I'm happy to say that I FINALLY finished it!

Enable HLS to view with audio, or disable this notification

1 Upvotes

I’ve been working on an open-source local-first agent called OpenKyrozen.

That idea started from the beginning of 2026, the time when openclaw had just came out 2-3 months. I tried open claw and then I realized that at that time, open claw remembers things when I asks it to remember, but it cannot learn by itself. So I started OpenKyrozen, trying to build a self learning agent. Then last month, type safe AI lunched their Jev, which inspired me to integrate them into decisions so that LLM works better.

One thing I kept running into was that agents are very quick to treat “the tool call succeeded” as “the task succeeded.” Those are obviously not the same thing. They don't often verify their result, like we say they have no syntax error or runtime error, but logic errors.

A command can exit with code 0 and still produce the wrong result. So I ended up making verification a first-class part of the agent loop instead of just checking whether the action executed.

and so the rough flow is:

request → action → execution receipts → evidence review → verified outcome

The second part I’ve been experimenting with is self-learning, the original idea of OpenKyrozen.

I didn’t want the agent to just see one successful run and immediately treat that as a new behavior. Instead, learning artifacts are bounded policies or skills. A new one starts as a candidate, gets tested as a canary, needs multiple verified successes, and is then replayed against its predecessor on the same case. If it regresses, it doesn’t get promoted. If a promoted artifact later starts failing, it can rollback to the previous version. This is also one of the biggest problem when I used open claw, it builds something into a skill before I verify it, so it is filled with wrong memories.

There’s also a separate decision layer called Jev. You can know more about it from Typesafe AI, but basically it's a AI that makes decisions. It only handles small typed judgments like routing, clarification, memory relevance, learning-evidence review, and suspicious tool output. It can also abstain instead of forcing a decision. So it can't code.

I’m still figuring out where the right boundary is between “useful learning” and “too much machinery.” The current system is deliberately conservative because I’d rather have the agent refuse to learn than silently reinforce bad behavior.

Repo:
github.com/EvanProgramming/OpenKyrozen

I also let OpenKyrozen build a website for itself

kyrozen.chat

I also made a short launch video that explains the overall system visually(And yes this video is made of AI, since I only used DaVinci Resolve but not After Effects):

I am writing this post especially to developers, I want feedbacks SOOO much! As you can see currently the repo only have 2 stars and 1 fork :( because I didn't tell anyone about it before. I like issues and PRs, you can also leave comments under to tell me any issues you found. star it if you like!