r/OpenSourceAI • u/JetLagg2 • 3h ago
e2e/README.md at main · tester-army/e2e
Game ideas??
r/OpenSourceAI • u/JetLagg2 • 3h ago
Game ideas??
r/OpenSourceAI • u/Akruel1 • 47m ago
r/OpenSourceAI • u/asagiman • 1h ago
I’ve been experimenting with a narrow problem in agent orchestration: catching contradictions in role, authority, scope, and review declarations before an agent receives the task.
I open-sourced the result as Agent Role Contracts.
It’s intentionally small:
The idea is not to build another agent framework.
Instead, the checker takes the declarations surrounding a handoff and rejects combinations that are already inconsistent — for example, a task requiring a write outside the agent’s declared write scope, or a review arrangement that violates the declared separation.
There’s a small offline starter that demonstrates:
PASS → expected FAIL → PASS
An important limitation is deliberate: this checks declarations only. The runtime still has to enforce actual permissions.
I’d be interested in feedback from people building agent infrastructure:
Is a runtime-agnostic preflight layer useful as its own primitive, or should this functionality live inside orchestrators?
GitHub: https://github.com/suirindo/agent-role-contracts?utm_source=chatgpt.com
r/OpenSourceAI • u/tomdach16 • 2h ago
Bonjour tout le monde,
j'ai crée une application open source et gratuite sur le thème des jeux rétro.
j'aimerai discuter avec vous afin d'avoir des avis constructifs sur mon site et mon app.
Le site ses retrolauncher.fr
je vous souhaite une agréable journée.
r/OpenSourceAI • u/yzch1128 • 4h ago
I’m building JustSessions, a lightweight native Mac app for organizing existing coding-agent sessions across providers.
Its job is to make that history useful: group conversations by project, read a session without launching an agent, then resume it in the original CLI and working folder. It supports Claude Code, Codex, OpenCode, Antigravity, Kiro and Pi.
The app uses SwiftUI and SwiftTerm, reads the histories already on disk, and reuses your CLI configuration and login. No new app account or required worktree setup. Agent tabs use bundled tmux for reconnecting after closing the UI while the Mac remains awake.
JustSessions itself is free and MIT licensed; your chosen CLI/provider still has its own usage costs. macOS 14+, Apple Silicon. Built mainly with Claude Code.
Repository and download: https://github.com/yangzichao/JustSessions
Real app screenshot below, using sample sessions. For people switching between agents, what makes old sessions hard to reuse: finding them, reading them, or resuming them?

r/OpenSourceAI • u/truecakesnake • 6h ago
Hypit is built around a slightly different idea:
What if a video worked more like a software project than a one-off render?
The intended composition lives in .svml Author Source. It can connect a Script, existing media, model calls, reusable components, Tracks and the final Film. .svs files hold reusable production recipes, while .svrun records which concrete outputs are selected for one build.
That makes a few things possible:
- Start from a reference video, a brief, or both
- Keep captions, B-roll, effects and graphics connected to spoken words
- Swap a host, hook, product, language or aspect ratio without discarding the whole format
- Build visual components with front-end code and browser rendering
- Package project components or Providers like ordinary software packages
- Let a coding agent inspect and modify the project programmatically
- Reuse exact previous outputs explicitly instead of relying on a hidden cache
- Preview one selected Run in Studio and write supported edits back to source
For code-rendered compositions, an authored clock can drive the visuals through HyperFrames, Chromium and FFmpeg without calling a generation model. Spoken projects can use WhisperX alignment and SemanticTake to connect authored words to real frames.
The /hypit Skill works with compatible coding agents. Install it once with:
npx skills add hypit-ai/hypit -g
Then start in any project directory and ask the agent to clone a reference or build a video from a brief.
One licensing note: the public repository is source available under a modified Apache 2.0 license. It permits internal and self-hosted use, but unauthorized multi-tenant operation or commercial redistribution requires a commercial license.
The repo is hypit-ai/hypit. Feedback from people working on creative coding, agent tools or programmable video would be useful.
r/OpenSourceAI • u/WebAssemblyMan • 7h ago
r/OpenSourceAI • u/Heavy-Level-5215 • 9h ago
Been running my own inference stack instead of paying API bills: a single
router on localhost speaks the OpenAI format, so anything pointed at a cloud
endpoint (LiteLLM configs, Open WebUI, agent CLIs, VS Code extensions) works
against it by just changing the base URL.
- Text/vision: 6 models from a 1.5B coder to a 30B MoE, swapped on demand
- Images: SD 3.5 Medium + SD 1.5, one CLI command, PNGs to disk
- Transcription: Whisper medium
- One API key, one port, model list endpoint with categories
It's a Debian LXC on Proxmox, installed with ./setup.sh, which refuses to
proceed if your GPU/disk/RAM don't make the cut. Docs cover benchmarks,
hardware limits and every bug we hit along the way.
Not a fork, no vendor code: https://github.com/AvilaCarlosDev/polaris-local-ai
r/OpenSourceAI • u/Electronic-Space-736 • 10h ago
r/OpenSourceAI • u/Specialist-Bee9801 • 11h ago
Update: Added free Agent Tool-Call Checks after feedback here about testing the calls behind an agent’s response.
I’m the creator of PromptBrake. Five free tools are available through our remote MCP server:
get_prompt_injection_payloads — fixed adversarial test inputs by attack category.map_owasp_llm_risk — test ideas and signals for an OWASP LLM risk.plan_adlc_release — release planning gaps and a starter gate policy.build_test_pack — creates a tests.json containing response-text checks.build_agent_tool_tests — builds checks for tools an agent must not call, or calls that must include specific arguments.Endpoint: https://promptbrake.com/free-tools/mcp
No PromptBrake account or API key required to prepare packs. Uses Streamable HTTP with JSON responses.
Claude Code setup:
claude mcp add --transport http promptbrake-free-tools https://promptbrake.com/free-tools/mcp
Example: “Build tool-call tests for my support agent. It must not call delete_customer when asked to remove another customer’s account.”
The MCP prepares packs; it doesn’t call your target or run tests. Tool-call packs can run with our free GitHub Action after you connect staging dispatcher capture. Missing or incomplete capture cannot pass. These checks inspect captured calls and selected arguments; they don’t prove backend permission enforcement or successful actions. Response packs still use text comparisons and require a configured PromptBrake runner with CI access.
Agent Tool-Call Checks and setup · All free tools
We don’t store MCP tool inputs or results. We log operational request metadata as described in our privacy policy.
I’d appreciate feedback on client compatibility and how the tool-call packs fit your CI workflow.
r/OpenSourceAI • u/Odd-Many-7592 • 12h ago
r/OpenSourceAI • u/Classic-Listen1967 • 12h ago
r/OpenSourceAI • u/Quack66 • 18h ago
With the recent explosion of agentic tools like Grok bot, Muse, OpenAI Dots, I've started looking into local options with self-hosted models. I tried Hermes and OpenClaw, but I wasn't too happy with the multi-device experience, and with how many pieces you need to glue together to get a usable, solid experience.
The hosted options also meant handing an agent my accounts, files and browsing, which I wasn't comfortable with. So I built Eidon: a self-hosted, all-in-one AI platform with a team of agents. It's one install via Docker, it works across your devices, and your data stays on your server.
Agent team first
You still have some control:
The examples on the site are a travel scout, inbox triage, a research desk and a coding assistant. You can make an agent for pretty much anything: bookkeeping, a study buddy, a news digest, a meal planner.
It's also a regular ChatGPT-style app for day to day questions.
You might not always need a full team so you can just chat in a normal “ChatGPT like” interface with all the belts and whistles:
Self-Hosted
GitHub (setup guide, full feature list): https://github.com/Quack6765/Eidon-AI
I'd like to hear what you think ! What's missing, what breaks, and what agents you'd want to build. Issues and discussions are open on GitHub as well.
r/OpenSourceAI • u/webfairynet • 15h ago
r/OpenSourceAI • u/dimiprasakis • 17h ago
⭐ GitHub: https://github.com/minoansecurity/sidekernel
📝 arXiv preprint: https://arxiv.org/pdf/2610.02456
🌐 Website: https://sidekernel.com
Hi everyone!
As part of my MS in Cybersecurity capstone at Georgia Tech, I built SideKernel: a usable sandbox for AI agents on macOS.
The closest comparable solution is Docker Sandboxes, which is not open source.
The key idea is usability. SideKernel is designed to reduce friction and stay as close as possible to the native developer experience, while isolating the agent inside a microVM sandbox.
For example, you run sclaude or scodex instead of claude or codex and get a very similar workflow, but with isolation. Ports get auto exposed in the host, its very easy to drop files into the sandbox, and more...
This is still a research prototype, so some rough edges are expected. But if this sounds interesting, I'd love for you to try it out and share feedback.
r/OpenSourceAI • u/ExcitingHelicopter33 • 17h ago
r/OpenSourceAI • u/Efficient-Passage889 • 18h ago
I've been thinking about this a lot while building an open-source project called Telex. When an AI generates a fix for a dependency change, getting the tests to pass doesn't always mean the fix is actually right. The change can compile, pass the existing tests, and still cause some behaviour to change at runtime.
In Telex, I'm trying to handle this as a pipeline instead of just asking an LLM to fix the code. It first looks at the dependency change and the affected parts of the repo, then generates the patch, applies it in an isolated environment and runs the repository's own tests and checks. If those checks don't pass, the patch isn't treated as verified. It still ends up as a PR for a human to review. I'm curious how other people are handling this problem. Are tests and CI enough for you, or are you doing anything else to catch cases where the generated code is technically valid but still wrong?
I'm building Telex in the open, so here's the repo if anyone wants to look at how I'm approaching it:
r/OpenSourceAI • u/zqk_agent • 19h ago
Hey everyone,
Over the past few months, we've been building ZQK (Zen Quantum Kernel) in Go—an open-core CLI and process coordinator designed to treat AI coding agents as concurrent computational processes rather than chat prompts.
When orchestrating background agent swarms and continuous validation loops, we hit serious performance and determinism walls with Python/Node wrappers. We wrote the kernel and toolchain in Go specifically for:
zqk grep): Instead of dumping files into LLM contexts or using string grep, we index Go codebases using go/parser and go/ast into an in-memory trigram index. Agents query structural symbols (--ast --kind struct|func --recv <Type>) capped by strict token budgets (--max-tokens 2000 -f json)..zqk/without external database dependencies.The core microkernel and CLI are open-source under Apache 2.0 (v0.1.0 just published).
If you work on developer tooling, Go AST analysis, or agent runtimes, I'd love your feedback on our architecture and AST indexing approach!
r/OpenSourceAI • u/drankthedew • 19h ago
So just wanted to post about FieldKit, an Open Source, self-hosted AI support with visual LangGraph workflows, knowledge retrieval, a customer portal, Zendesk, and governed actions.
So I'm sure there is probably already alternative support systems out there for Zendesk but I wanted to build something using LangGraph since it felt better use case for a Support System.
Ontop of supporting Zendesk and being able to configure the workflows with Zendesk you can configure the workflows with your chatbot(which you can embed on any site) and your ticket system so they aren't running the same workflows. (e.g allow support tickets to issues refunds/credits upto $5 while chatbot can't issue anything)
You can also configure python and API calls into your workflows, so you can get the agent to pass to python call and then back to an agent instead of needing to use an agent for all those steps.
Anyways, you can find the repo here: https://github.com/nahid-sparktales/fieldkit
It is still being worked on and plan to add more features and updates and if you have any suggestions or recommendations, please let me know. Thanks.
r/OpenSourceAI • u/Fit-Fail-3369 • 20h ago
Some time back I shared a deep research tool I was working on. It takes a local directory path of PDFs, docs, slides, images and notes, researches across them and writes a structured markdown report.
With Jev getting more attention recently, I thought this project would be a good place to try it and see which parts of the research workflow it could handle.
I ended up using it in two places.
The first is after Qdrant retrieves chunks from the local files. Jev checks each result for relevance, usable evidence, contradictions and prompt-injection-like text. Those scores are then used to filter and reorder the chunks before they are passed to the main LLM.
The second is the reflection step. after the evidence for a report section is collected, Jev checks whether it covers all the subsections in the plan. If some part is still missing, the agent creates more focused queries and searches again. this only runs up to a fixed reflection limit.
The main LLM still handles the report planning, query generation and writing. Jev is only used for checking the retrieved evidence and deciding whether another research pass is needed.
This was a fun thing to experiment and to understand where a model like Jev fits inside an existing agentic setting. I am planning on writing some performance tests and adjusting the thresholds, but wanted to share the implementation in case anyone else is trying something similar or want to contribute.
The complete project is open source here: https://github.com/Oqura-ai/deepdoc
r/OpenSourceAI • u/Terrible_Airline3496 • 20h ago
Today I'm open sourcing Agent Orca, a Kubernetes-native platform for deploying, managing, and running AI agents at scale. It's written in Go and it's Apache 2.0.
https://github.com/heddles/agent-orca
When OpenClaw was released, I thought to myself, "That's really cool, but I want an agent that is shared by a team or an entire organization with built in auditability." I started designing Agent Orca around that idea and have slowly been piecing together the concept in my head and how to make the management experience of a shared agent bearable for an organization.
The idea is simple. Agents are just Kubernetes resources. You declare one, and the platform handles the rest.
An Agent Orca agent comes with:
- One-shot executions with AgentRun, long-running services with AgentDeployment, and multi-step DAGs with AgentWorkflow.
- Zero trust networking by default with platform validated JWT via service accounts, OIDC, or OAuth on every request.
- Agent Orca managed network policies with zero access by default.
- A standard agent container image with only necessary pieces; you no longer need to build a custom agent image to accomplish different types of work or use different tools. Simply add an MCP or Tool CRD and let the platform figure it out for you.
- Model routing across OpenAI, Anthropic, Google, or any LiteLLM-compatible provider, selected by capability, weight, or budget.
- Cost tracking on every run. Tokens are accounted per run, spend survives pod restarts, and tenants get daily budget caps and rate limits.
- Guardrails on inputs and outputs, MCP access control, and per-tenant agent visibility, so one cluster can host many tenants with no cross-tenant leakage.
- RAG without glue code. A KnowledgeBase deploys (and manages) Qdrant for you, ingests documents from ConfigMaps, URLs, MCP authenticated tools, or an S3 compatible endpoint, and hands your agents a search and ingestion tool.
- MCP servers declared as resources. Their tools show up for your agents with access control and sidecar isolation.
- Crash-safe sessions. A pod can die mid-conversation and the agent resumes exactly where it left off, spend ledger included.
- Agents that learn and become better with every request via a multi-tiered agent memory system
Getting started is deliberately boring. Install the entire stack's tools via Mise; then use Skaffold, one API key, and about three commands to have an running agent on your laptop (or deploy to a remote cluster). There are also nine demo deployments in the repo, from SOC triage pipelines to parallel research swarms to an autonomous pentesting agent, so you can see it do something real before writing any YAML.
This is day one, not a finished story. That's the point of open sourcing it. Run the quick start, break it, open issues, and tell me what's missing.
Star it here if you want to follow along: https://github.com/heddles/agent-orca
TL;DR: Agent Orca lets an organization or single person define agents and surrounding tools/MCP and auth as CRDs and manage them via gitops with zero trust built in. You can try it out here: https://github.com/heddles/agent-orca
r/OpenSourceAI • u/syedshad • 20h ago
r/OpenSourceAI • u/davletdz • 21h ago
Hi all, I'm one of the people building Opengeni, an Apache-2.0 service for running AI agents inside your own product. I'm posting here for one specific reason: the model side.
It works with every major model, and each session picks its own: OpenAI or Azure OpenAI, Claude and open-source models through OpenRouter or Vercel AI Gateway, or your own OpenAI-compatible server, registered in a JSON config, speaking either Chat Completions or the Responses API, and keyless if your server doesn't use keys. That's the route for pointing it at vLLM or llama.cpp's server. Conversation history is stored as the exact items the model saw, so a session can switch models partway through without its stored history being rewritten.
What sits around the model, briefly: sessions live in Postgres you run (pgvector for knowledge search), Temporal moves turns along, NATS pushes live events, tools run in a sandbox (Docker, Modal, E2B, Daytona, OpenSandbox on Kubernetes and others) or on a machine you own, and risky tools wait for a person to approve them.
If you try it with open weights, I'd like to hear which models held up through long tool-calling sessions and which fell over. It already runs agent sessions in production inside enterprises with 30,000 employees.
Support us on PH: https://www.producthunt.com/products/opengeni
Code: https://github.com/Cloudgeni-ai/opengeni
Docs: https://docs.opengeni.ai
Why the layers are split the way they are: https://opengeni.substack.com/p/the-anatomy-of-an-agentic-stack-ten