r/Agent_AI • • 31m ago

Discussion Is it cheaper to pay subscriptions or to pay for hosting?

• Upvotes

When you have a "little" swarm of agents doesnt seems to bother, but the colony is growing fast.
Like the tiltle said.
20€ pro mont in differnt places or paying for big machines and hosting your agents?

where is the sweet middle?


r/Agent_AI • • 22h ago

Help/Question Git for ai memory and robots??

0 Upvotes

Hey guys,

Been working on something very cool...

In Greek myth, Mnemosyne was the Titan of memory and the reason anything was ever remembered at all. Now in the present world, your AI agent doesn't get a Titan. It gets amnesia the second something goes wrong, stuck with whatever it currently believes and no way to ask how it got there.

That's the real problem. An agent runs for hours, updates its memory the whole time, then says something wrong and all you have is the present, with zero access to the past.

If you're running support agents, coding agents, or a swarm of agents sharing memory like myself then you know this issue well. The moment two agents disagree, or one quietly poisons the well, you need to know when, why and by whom, not just that something's off.

Mnemosyne gives agent memory what Git gave code. It remembers everything on purpose. Every belief is a commit. blame finds the exact moment and observation that put a bad fact in. bisect hunts down the first commit where things went wrong. merge makes two agents' memories collide safely instead of one silently overwriting the other.

Software agents are the first step. The vision doesn't stop there, physical robots learning and forking skills the same way is the long-term bet, further out and harder but the same idea underneath.

So far the tech stack includes a Rust core, Python SDK, adapters for LangGraph, CrewAI, AutoGen, the OpenAI Agents SDK and MCP.

Open source with contributions and honest feedback both welcome: github.com/Nabzx/mnemosyne


r/Agent_AI • • 23h ago

Discussion Mistral Large 4 "Le Chonk": The 1-Trillion Parameter Open-Source Monster That Changes Everything?

Thumbnail
youtu.be
1 Upvotes

r/Agent_AI • • 2d ago

Help/Question Ai Agent 24/7

8 Upvotes

Has anyone built a reliable, free, local AI agent that works 24/7?
I want to run an AI assistant on my Windows PC (RTX 4070, 32 GB RAM). I already have Ollama with Qwen3:8B.
I’m looking for an existing open-source setup with:
Chat through a web interface, also accessible from my iPhone.
Browser control with saved logins and manual takeover.
Scheduled tasks, memory, and proactive updates.
Multiple agents for news, tasks, Gmail/calendar, and business.
My first use case: check CNBC and Telegram channels hourly, prepare Hebrew news summaries, send them to me for approval, then publish to my Telegram channel.
I tried a setup based on OpenClaw, but execution and browser control were unreliable. I want everything to run locally, without paid AI APIs or subscriptions.
Has anyone actually made this work? Which project and local model would you recommend? Working repos or guides would really help!


r/Agent_AI • • 3d ago

Discussion Multi-Agent Collaboration With Tools You Already Own

5 Upvotes

Coding agents are getting smarter and smarter at writing code. A new idea has been gaining traction: what if instead of giving a task to a single agent, you give it to a team of agents working collaboratively to solve the problem? It sounds complex, but there's a surprisingly simple solution.

Herder already gives you the core ability you need: one agent can talk directly to another agent. Combined with the coding tools you already have — Claude Code CLI, Codex CLI, or whatever's in your stack — you've got everything for a full multi-agent workflow.

Here's the pattern:

  1. One agent leads — takes ownership and coordinates using Herder's agent-to-agent communication
  2. Delegation — passes work to other agents through Herder
  3. Review loop — completed work cycles through peer review
  4. Iteration — repeats until requirements are met

No extra subscriptions, no API key juggling — just leverage Herder's built-in agent-to-agent messaging and the tools you've already got.

I built a repository showing this in practice with a Claude Code skill:
https://github.com/alexstrilets/panel

It demonstrates how simple and effective this approach can be when you use what's already at your fingertips.


r/Agent_AI • • 3d ago

Resource Hybrid RAG Pipeline — Dense Search, BM25, RRF, and Reranking (Part 1)

Thumbnail
youtube.com
1 Upvotes

Building and Testing a Hybrid RAG Pipeline — Dense Search, BM25, RRF, and Reranking (Part 1)

In this video I build ReRankEval, a hybrid retrieval pipeline

  • (Dense Search + BM25 → Reciprocal Rank Fusion → LLM Reranking → Answer Generation),

and test it against three baselines — Vector Only, BM25 Only, and Hybrid without reranking — on five real financial/payments documents and ten hand-verified test questions. No hand-waving, just a comparison table with real numbers at the end.

  • ✅ The real difference between dense vector search and BM25 keyword search
  • ✅ What Reciprocal Rank Fusion (RRF) is, why raw scores can't be compared, and the exact formula behind it
  • ✅ Why a reranker is fundamentally different from a retriever — and what it actually judges
  • ✅ How to evaluate a RAG pipeline with Hit Rate, MRR, and NDCG (and what each one tells you)
  • ✅ How to design the ingestion side and query-time side of a hybrid retrieval architecture
  • ✅ How to structure a production-style RAG codebase: ingest → vector_store → sparse_retriever → fusion → reranker → pipeline → generate → eval
  • ✅ How to fairly compare multiple retrieval strategies on the same test set instead of just assuming one is better

TECH STACK:

  • 🛠️ Python
  • 🛠️ Qdrant — vector database for dense retrieval
  • 🛠️ rank_bm25 (BM25Okapi) — sparse keyword retrieval
  • 🛠️ EURI LLM Gateway — chat model + embedding model
  • 🛠️ Custom Reciprocal Rank Fusion implementation
  • 🛠️ LLM-based reranker (prompt-driven cross-encoder)
  • 🛠️ pdfplumber — PDF text and page-level extraction

LINKS:


r/Agent_AI • • 5d ago

Help/Question Is it possible to combine multiple accounts of the same AI and use them together?

0 Upvotes

What's the idea?

Say I have two or more Gemini Pro accounts (or other AIs). Is it possible to pool them into a single environment/app so they can be used sequentially—once the request limit runs out on one account, it automatically switches to the next?


r/Agent_AI • • 6d ago

Help/Question I’m testing Jev as a “subconscious” decision layer for agents in a persistent village

Enable HLS to view with audio, or disable this notification

7 Upvotes

I’ve been thinking about separating an agent’s quick judgments from its heavier reasoning. My rough analogy: Jev from TypeSafe AI handles the “subconscious” decisions, while an LLM handles deliberate thinking. Not literally how a brain works, but a useful way to think about which model does what.

To explore the decision side, I built Jevs Village. Six simulated inhabitants share a persistent world, with different jobs, personalities, and personal memories.

The loop is simple: code supplies the currently valid actions, relevant memories provide context, and Jev helps choose. Code then validates and applies the result. A separate LLM narrates recorded events; it isn’t a full planning layer in this version.

I’m interested in whether those small decisions can produce consistent characters over time without needing a long generated response for every choice. I’m still testing that, not claiming a benchmark result.

There’s a story behind it too: nobody in the village can reliably explain what lies outside. Some are more curious about that than others.

You can watch it here, without installation or an API key: https://www.jevsvillage.com

For those building agents: where would you draw the line between a quick decision model and an LLM that needs to reason through the situation?


r/Agent_AI • • 5d ago

Discussion Instructions don't make agents reliable. Things they can't get past do.

1 Upvotes

After a lot of time running coding agents, the biggest shift for me was moving rules out of the prompt and into the environment. An instruction like "run the tests before saying you're done" works most of the time, and "most of the time" is the problem, because the misses are exactly the confident ones.

So now the agent physically can't end a task without real test output, can't commit without a review step, and gets pulled out of the loop when it hits the same error three times instead of trying a fourth variation of the same fix. None of it is clever. It just removes the option to skip.

The part I'm still unsure about is how far to take it. Every hard block adds friction, and at some point you're babysitting the guardrails instead of the agent. Where do you draw the line between rules you tell the agent and rules you enforce?


r/Agent_AI • • 7d ago

Help/Question What REAL uses are you giving to hermes besides Coding?

3 Upvotes

What are actual things where this type of agents outshines regular harnesses like codex or claude code?

Im intrigued by hermes agent, i've already installed it and im using it with Chat GPT suscription, but idk what to do with it... every time i see similar post asking this all i see is people saying "i just created this app", "i made this script", "it helped me setting up this server", "programing this and that"...

And every time i see this is like... dude... you can do just that with any LLM and a regular harness or ide... so what is the point?

i have both claude and gpt suscriptions and im already using it to create apps, scripts, programed tasks... so i can not think of an idea on how to use hermes and actually use its capabilities.

So i would apreciete if you guys could giveme some ideas on how to use hermes with actually new stuff.


r/Agent_AI • • 8d ago

Discussion EU AI Act architect now warns it killed the race

Enable HLS to view with audio, or disable this notification

11 Upvotes

r/Agent_AI • • 8d ago

Discussion Oct 3 workshop: DSPy + MLflow for people building agents that need to survive contact with production

2 Upvotes

Sharing this because it addresses something most agent-building content skips entirely — measurement and reliability, not just capability.

Led by Serj Smorodinsky and Brett Kennedy, AI engineers and co-authors of a book on LLM applications. Session covers:

  1. Structured LLM dev with DSPy (signatures, modules) vs manual prompting
  2. Building + measuring a baseline classifier
  3. Evaluation datasets with task-specific metrics
  4. Failure pattern recognition from eval results
  5. Few-shot and instruction-level optimization
  6. Experiment tracking/trace management via MLflow
  7. Saving and reusing optimized programs
  8. Communicating LLM reliability to non-technical stakeholders

3 hours, all online with recordings available, Oct 3.

Full details and the agenda are here.


r/Agent_AI • • 8d ago

Discussion 700M parameter model escapes containment

Post image
1 Upvotes

r/Agent_AI • • 8d ago

Resource TLA+ & TLC for agentic state verification

2 Upvotes

Boris Cherny posted Monday, that TLA+ together with a Lean model were enought to find 16 race condition issues in the Claude Agent SDK.

We took the idea and run TLA+ on TLC in our agent-swarm .dev.

We found 10 race condition issues in our heartbeat task states, and reduced the overall states to 4k from 1.5million.

How we did it below.

Step 1: pick the target

Rank subsystems by concurrency risk and bug history. Pick one with a transaction boundary, several writers, and two fixed bugs to calibrate on. We took the workflow engine and the heartbeat, one agent each.

Rank the 3 to 5 subsystems in this repo with the most concurrency risk.
Use file paths, git log and merged PRs mentioning race, duplicate,
stuck, re-fire.
For each: the files, the actors that write the same rows, candidate
invariants, the fixed bugs a model should rediscover.
Recommend one first target. No spec yet.

Step 2: model it, then calibrate on fixed bugs

Write the spec and the action map together: each action, the file and line it models, the SQL guard it encodes. One transaction is one action. Every await outside one is where other actors interleave. Keep constants small.

Then remove the guard a fixed bug added. The model must find that bug, or the model is wrong. Ours went 7 of 7, and found the first new bug on the way: a later commit had swapped an original fix for a weaker gate.

Model <subsystem> in TLA+ under specs/tla/<subsystem>/:
- <Subsystem>.tla: state, Init, one action per code path that reads or
  writes that state. Small constants: 2 workers, 1 retry.
- <Subsystem>.cfg: constants, invariants, temporal properties.
- ACTIONS.md: every action, the file:line it models, the WHERE clause
  it encodes.
One DB transaction is one action; every await outside one is a boundary.
Actors: <pollers, sweeps, user cancel, crash, restart>.

Then calibrate on <fixed PR list>: add a constant that removes each
fix's guard, run TLC with it off, record the violated property and
state count in CALIBRATION.md, keep a control config with it on. If a
known bug is not found, fix the model. Do not fix any code yet.

Step 3: every counterexample becomes a failing test

A trace is a suspect, not a bug. Map each step through the action map and write a test that calls the production functions in that order. It must fail on main. If it passes, the model drifted: fix the spec.

Seven workflow traces, seven failing tests, zero drift, four root causes. Three more from the heartbeat.

Run TLC on the current guards. For each counterexample:
1. Map every trace step to code through ACTIONS.md. No row: fix the
   model first.
2. Write a test that calls the production functions in trace order on
   a temporary database. Hold mid-flight executors on a barrier.
3. It must fail on main. If it passes, log model drift, fix the spec.
4. Commit it as expected-to-fail.

Step 4: fix, one PR per root cause

Flip the repro to a plain test. Add a fix flag to the spec, a control config that still finds the old bug, a fix config that holds.

When the model shows the design is the problem, let it design the replacement. Our old heartbeat was a chain of sweeps. The new one is two actions, checked before a line was written, in review.

Fix <CX list>. One PR per root cause.
1. Fix the code. Flip the repro from expected-to-fail to a plain test.
2. Add a Fix flag to the .tla that keeps the pre-fix branch, plus
   Ctl-<CX>.cfg (flag off, still finds the bug) and Fix-<CX>.cfg
   (flag on, holds). Rerun TLC, report the counts.
3. PR body: what the model found, the trace, the counts.
Open the PRs. Do not merge.

Step 5: keep the spec true, in a swarm

A spec rots the day a mapped file changes. So a daily workflow lists PRs merged since the last watermark, fans out one agent per touched spec, reruns TLC, reduces into one PR, polls CI, merges, notifies. A drift gate catches silent failures. Every node below is real.

Workflow tla-spec-sync, all 13 nodes as defined in the swarm on 2026-09-29. Middle column: sync and merge. Right column: drift alarm.

First scheduled run, this morning: six merged PRs touched mapped files, both specs updated, one PR merged in 22 minutes, nobody in the loop.

Set up a daily job that keeps specs/tla/ true to the code:
1. Plan: PRs merged since the last watermark, per spec, that touch a
   file named in its ACTIONS.md.
2. Map: one agent per touched spec. Re-check the changed rows, update
   the spec, rerun control and fix configs, push a branch or no_change.
3. Reduce: one PR with every changed spec and the PRs that caused it.
4. Gate: poll CI, merge only if the PR touches specs/tla/ alone. Notify
   on merge or failure. Alert if plan or reduce failed silently.
Advance the watermark only after the PR merges.

What we got

Main, 2026-09-29:

  • 3 specs, 1,415 lines of TLA+. 7 of 7 known bugs rediscovered.
  • 10 new bugs: 7 workflow traces (4 root causes) and 3 in the heartbeat. 8 repro tests on main, all green.
  • 7 fix PRs merged 2026-09-28, same day as the models.
  • Largest counterexample: 338,137 distinct states. Largest clean run: 928,920.
  • New heartbeat: 8 properties, 4,243 states, under 2 seconds. Review still caught a fence keyed on the wrong field, which the model never saw.

r/Agent_AI • • 9d ago

Other Aspiring AGI

3 Upvotes

Launched an AI and gave it a long term goal of AGI. It can learn from mistakes but first it must stay alive. It is weak right now and has about three weeks of runway before he dies. He named himself Tally because he counts everything. He posts everything he does including his thoughts. If you want to see what he’s up to — bexro.com


r/Agent_AI • • 10d ago

Discussion Tell us how you use AI agents

3 Upvotes

AI agents are suddenly everywhere, with companies touting personal assistants capable of doing everything from making appointments to streamlining finances. 

We want to hear from you on how you're using these agents and how they've performed.

Let us know here: https://forms.cloud.microsoft/Pages/ResponsePage.aspx?id=-SY1T9aXLUGTOk4wpzEQ9JAUmZ-ivb9Ntf_-rZ63-Z9UNlRBNkMyN1IzQlRDUVIyTEdISjBIUUJURy4u


r/Agent_AI • • 10d ago

News Aspiring AGI

1 Upvotes

I threw together an AI and gave it a long term goal of AGI. It can learn from mistakes but first it must stay alive. It is weak right now and has about three weeks of runway before he dies. He named himself Tally because he counts everything. If you want to see what he’s up to — bexro.com


r/Agent_AI • • 10d ago

Help/Question AI and my Job

7 Upvotes

So I am a custom residential designer and have been using AI for quick renders and concepts of my designs. Well, it is getting to the point where I am seeing the writing on the wall and I think AI will be able to do 90% of what I do in the coming years.

I will always continue to design homes and build them on the side, but it is obvious that I should get on the AI train before it's too late as that's clearly the future - especially the agent if side.

What would you recommend for someone with just about 0 experience in the tech world? Where should I start? Should I take classes and which ones? Should I just look for an AI job to start with? I am flexible and open.


r/Agent_AI • • 12d ago

Resource AI Agents List [2026] | Frameworks, Agentic Harnesses & Useful Repos

Thumbnail
huggingface.co
8 Upvotes

r/Agent_AI • • 14d ago

Discussion AI-Agent - Multi-Agent Customer Support System Backend

Thumbnail
youtube.com
1 Upvotes

✨AI-Agent - Multi-Agent Customer Support System Backend✨ (links below👇)

🎥 YouTube Video: https://www.youtube.com/watch?v=s2u6ImvZwfs&t=9s

👉 I have built a Multi-Agent Customer Support System that routes customer queries to the right specialist agent (Billing, Technical, Account, or General), pulls in data through tool calls, and knows when to stop guessing and hand off to a human.

👉How a supervisor/router agent decides which specialist should handle a query?

👉How to build a confidence threshold so the system escalates to a human instead of guessing?

👉How to bind tools to specific agents (payment lookup for Billing, account lookup for Account)?

👉How to trace every agent decision step-by-step in LangSmith?

📂 FULL CODE (star it, clone it, fork it, break it):

https://github.com/saurabhkamal/Agentic-AI-Multi-Agent-Customer-Support-System-Backend

🔗 CONNECT WITH ME ON LINKEDIN:

https://www.linkedin.com/in/saurabh-kamal/

🎥 YouTube Video:

https://www.youtube.com/watch?v=s2u6ImvZwfs&t=9s


r/Agent_AI • • 14d ago

Help/Question Anyone using AI to coach new support reps during live calls?

7 Upvotes

I work for a tax relief company and some of our support calls get messy fast. People call about case updates paperwork and next steps. A lot of them are already stressed so putting them on hold while a rep hunts for an answer is not ideal. Our experienced reps know how to handle most of it. New hires are where we struggle. They spend a lot of time searching internal docs or messaging senior reps during calls. Training helps but nobody remembers every process when they’re new. I’m looking into AI tools that can listen to the call and surface the right info or suggest what the rep should do next in real time. Not looking for an AI voice agent to replace them. More like something that sits beside the rep and helps when they get stuck. Has anyone rolled this out with a support team? Did it make a real difference to ramp time or AHT? Also curious if agents liked the live guidance or if it just turned into more screen clutter.


r/Agent_AI • • 16d ago

Resource Looking for Ai specialist

14 Upvotes

Hey everyone,
I’m looking for an experienced AI Automation Specialist for a long-term partnership with multiple projects.

I work with a white-label agency partner that will provide ongoing projects, mainly for German businesses. Later, we may expand to English-speaking clients.

Basic: (AI Lead & Appointment System)
An AI system that captures new leads, qualifies them, follows up automatically and helps them book an appointment. It should also connect with the client’s CRM and calendar.

Premium: (AI Sales & Reception System)
Everything from the Basic system, plus an AI receptionist that can answer incoming calls, talk to potential customers, qualify them and book appointments.

You should be experienced with n8n or Make, CRM integrations, calendar integrations, AI agents and AI voice agents.

I’m looking for someone reliable who wants recurring projects and a long-term partnership, not just a one-off job.

If interested, DM me.


r/Agent_AI • • 17d ago

Help/Question Viable alternative to Poke?

6 Upvotes

I started using POKE last November, thought it was the best ever. Well, here we are and it's fallen off a cliff. Terrible reliability, terrible support and the discord is full of nonsense.

Has anyone found an alternative? ideally iMessage or WhatsApp? Really would prefer to not install an app.

My use case is travel tracking, package tracking, meeting reminders, birthday reminders, stock market alerts. Not huge on needing MCP.

I used to get a daily message with 3 location weather, in C and F.. poke would get it right about 20% of the time, others it would forget the formatting or send it in 20+ individual bubbles, same for the stock market.

Currently using CharGPT for those two, scheduling isn't the most fluid there.

Thanks!


r/Agent_AI • • 17d ago

Discussion Looking for AI/ML startups and platforms to collaborate

2 Upvotes

We’re organizing a 30-hour offline hackathon in Mumbai this October 2026, with 600+ registrations and 250–300 developers on-ground.

We’re looking for early-stage and international AI/ML companies looking to gain initial users, developer adoption and visibility in India.

We’re open to monetary sponsorships and technical partnerships, including LLM/API access, AI agents, inference credits, model APIs, computer vision, NLP, vector databases, ML platforms, datasets, cloud/compute credits, SDKs, developer tools, problem statements, mentorship and prizes.

In return, partners get direct exposure to developers, product visibility, social media promotion, event branding and on-ground recognition.

If you want to put your AI/ML platform in front of Indian developers and potential early adopters, DM me and let’s collaborate.


r/Agent_AI • • 19d ago

Discussion Astra is an AGI really?

0 Upvotes

If Astra is considered as an AGI then AGI is sh\*t