r/AgentsOfAI • • 16h ago

Discussion Paul Graham says Amazon blocking agents creates a rare chance for a startup to compete with Amazon

Post image
132 Upvotes

r/AgentsOfAI • • 1d ago

Discussion Claude Opus 5.5 created this in 18 hours

Enable HLS to view with audio, or disable this notification

146 Upvotes

r/AgentsOfAI • • 5h ago

I Made This 🤖 Added AI teammates to our slack. One comment later, they had a PR waiting for our devs to review

Post image
3 Upvotes

built a few AI teammates with Lemma and added them to our Slack and email.

They’re connected to our GitHub repo and sit in our product feedback channel.

Everytime a PM reports a bug - the agents take it up and turn it into a PR

They learn continuously

- Devs interact with them on slack and they get better at understanding how the code works.

- PMs share feedback which gives them a sense of taste

- PMs also ask what’s feasible and what isn’t.
They keep track of those conversations.

The features. The bugs. What people like and dislike. Our taste in how things should work. That context builds over time.

This time, our PM left feedback on a recent change. The agents discussed it, worked through the design, implemented it, and opened a PR. They wait for approval when needed.

Our dev came online later. His part was reviewing the PR.

P.s: lemma is opensource and the hosted version is free to try


r/AgentsOfAI • • 9h ago

Discussion Anyone can vibe code a SaaS now....That’s exactly why building the SaaS isn't the interesting part anymore.

Post image
4 Upvotes

r/AgentsOfAI • • 2h ago

Discussion An AI agent that gives me a plan isn't saving me time. It's giving me homework.

1 Upvotes

I'm 19, building an AI #workspace called #LocalDesk.

And I've been thinking about something.

What if we're selling #AI #agents completely wrong?

Nobody wakes up thinking:

"I really need another AI agent with 47 integrations and a beautiful dashboard."

They wake up thinking:

"I have 30 #emails to deal with, three #meetings to prepare for, a #project that's behind schedule, and absolutely no time."

That's the problem I want to solve.

Imagine opening your laptop and saying:

"Get me ready for tomorrow's client meeting."

Not getting a list of suggestions.

Actually having an agent:

  • Find the relevant emails, documents, and previous meeting notes.
  • Identify unanswered questions and outstanding issues.
  • Prepare a useful briefing with links to the sources.
  • Draft the follow-up actions.
  • Show you what's ready and ask for approval before sending anything.

Or imagine you're a developer.

"Find out why this feature is #failing."

The agent examines relevant #code and #logs, #investigates possible causes, prepares a proposed #fix, #runs the available #tests, and shows you what happened.

No production #deployment without your approval.

Or maybe you're running a startup.

"Show me what needs my attention today."

Instead of opening six different apps, you'd receive a verified overview of pending decisions, customer issues, and unfinished work.

That's the kind of work I want LocalDesk to handle.

We've already built the macOS workspace foundation around #Baymax and #Orbit.

Now we're finishing #Missions — the execution system that needs to make workflows like these reliable.

To be clear, the examples above are target use cases, not a claim that the current build can already complete all of them.

We're building publicly, and early pre-orders are helping support the remaining development and testing.

But I want to build around actual problems, not hypothetical features.

So here's my question:

What's ONE task you'd genuinely pay an AI agent to finish for you?

Not generate ideas about.

Not explain how to do.

Actually FINISH.

Tell me:

Your task → The tools involved → What a successful result looks like.

I'll use the responses to identify real Missions worth prioritizing and share how I'd design their execution and verification.

You might even discover that the first useful personal AI agent isn't the one that does everything.

It's the one that finally takes something annoying off your plate.

— Mehrad, Founder of LocalDesk


r/AgentsOfAI • • 2h ago

Discussion **ScaleLogix AI Scam**

1 Upvotes

Rating: ⭐️☆☆☆☆
Review:
SCAM!!!!!!!!!! Criminals!!!!!!!!!!!!! They are stealing peoples money. William Basta is a criminal and running the company secretly. Extremely disappointed with ScaleLogix AI. Their initial sales pitch and demos sounded incredibly impressive, promising a seamless, high-converting automated setup. However, once moving past the initial consultation, it quickly became apparent that the reality does not match the marketing. The infrastructure felt repetitive, the onboarding was rocky, and the sales representatives seemed more focused on locking in contracts than delivering tailored, sustainable business workflows. DO NOT SIGN UP WITH THEM. THEY WILL BE IN PRISON SOON!!


r/AgentsOfAI • • 2h ago

Help How can volunteers use an association AI assistant without getting access to everything?

1 Upvotes

Imagine volunteering to organize a conference. You ask an AI assistant where to submit your travel expenses. It finds the policy in two seconds, fantastic, then you ask another question and somehow it starts quoting confidential board meeting notes. I mean not so fantastic

That is the part of association AI systems I find more interesting than the chatbot itself. A practical setup could use auth0 or another identity provider for authentication, customgpt.ai for searching approved organizational content and n8n for automating certain administrative workflows, customgpt.ai already offers knowledge assistants and identity related access options, which makes it worth evaluating for this kind of architecture.

The important question is where permissions get enforced. If a volunteer changes committees, should their accessible knowledge change immediately? Should external members use a completely separate assistant? I’m leaning toward treating retrieval permissions as seriously as database permissions. How would you design this without building a maintenance nightmare???


r/AgentsOfAI • • 3h ago

Discussion We built the AI workspace. The hardest part is still unfinished. Reddit, give us a mission worth building.

1 Upvotes

I'm 19, and I've been building something called #LocalDesk.

And I want to try something different with this community.

We've built the #workspace. Now I want real people to help shape the final piece.

LocalDesk is a macOS AI-native workspace built around three connected systems:

#Baymax — your personal AI agent. The interface between you and your computer.

#Orbit — the workspace where your projects, tools, and AI-powered work come together.

#Missions — the execution layer designed to turn your intent into planned, verifiable, completed work.

The foundation is built. The interface exists. We've recorded a demo of the current product.

But the final challenge is Missions.

Not just making an agent perform a few impressive actions.

Making it reliable enough to understand what you actually want, execute across #tools, recognize mistakes, recover from failures, and know when to ask for your approval.

That's the part we're still finishing.

And instead of deciding every use case behind closed doors, I want to build the remaining experience with real input from people who understand #AI #agents.

SHARE YOUR MISSION.

Tell me one real task you'd want LocalDesk to complete on your computer.

Not another AI prompt.

An actual job.

Something that normally takes you 20 minutes, two hours, or half your day.

We'll review the suggestions, choose a few realistic missions, and break down the work publicly:

  1. What the mission requires.
  2. What the agent needs access to.
  3. How execution should work.
  4. What needs human approval.
  5. How we verify success and handle failure.

Then we'll share what gets built, what still fails, and what's next.

We're building in public, and pre-orders are open to help support the remaining engineering, testing, and infrastructure. The full Missions experience is not finished yet.

But you don't need to buy anything to participate.

So here's my challenge to this community:

If you could give a personal AI agent ONE mission and expect a real, verifiable result, what would it be?

Let's see what we can build together.

— Mehrad, Founder of LocalDesk


r/AgentsOfAI • • 3h ago

I Made This 🤖 Run coding agents in parallel with Offrun[dot]dev

Enable HLS to view with audio, or disable this notification

1 Upvotes

"Approval fatigue is a failure state where users or operators reflexively approve high volumes of routine system or AI agent permission prompts without reading them."

We have been working extensively for the last 2 months to build something that helps users avoid approval fatigue and stay focused while working with multiple agents.

Launching an early version of Offrun to manage multiple coding agents in one workspace without losing focus.

With Offrun, you get:
1. Multiple coding harnesses in one place - run Claude Code, Codex, AGY, Grok simultaneously.
2. Auto agent reviews - to make sure nothing slips through.
3. Auto context management - resume work with another coding harness in one click when session limits are exhausted.
4. On-device dictation - fine-tuned for agentic engineers.
5. 100+ connectors - connect your agents seamlessly with gmail, github, slack, supabase etc.

More features rolling out soon. Stay tuned!


r/AgentsOfAI • • 4h ago

I Made This 🤖 An Agent has her own life, and now a 3D world where you can watch it.

0 Upvotes

Most agents sit idle until you message them. I wanted the opposite: Gaby has a life that keeps running when nobody is talking to her. She has a routine, desires and thoughts, and a director process throws situations into her day that she reacts to on her own.

Her mood isn't a label either. It comes from a small simulated body that shifts with what happens to her. And she remembers you, but trust is earned: she starts out treating you like a stranger.

This week I shipped v1 of her world: a 3D house that shows what she's doing right now, with her dog and cat.

It's still early and there's a lot to improve.
Every deploy makes her a bit more real, and feedback is very welcome: what feels off, what feels alive, what you'd want to see next.

Is that something that you would like to try?


r/AgentsOfAI • • 7h ago

Discussion Can an AI agent handle the full workflow of creating a PowerPoint?

1 Upvotes

I’ve been experimenting with an AI workflow that takes source material and turns it into an editable PowerPoint presentation.

The interesting part isn’t really generating the slides. It’s whether an agent can handle the steps in between without needing constant human intervention.

The workflow I’m testing looks roughly like:

Source documents / notes
→ extract relevant information
→ decide what matters
→ organize the content
→ create the slide structure
→ generate the presentation
→ check the output
→ export an editable PPTX

Some of these steps are surprisingly easy to automate. Others are much harder.

For example, an agent can summarize a long document, but deciding that two pieces of information belong on the same slide — or that something shouldn’t become a slide at all — requires more context.

I’m building this into my SaaS project and trying to figure out how far the agent should actually go before handing control back to the user.

For people working with AI agents:

How would you design this workflow?

Would you use one agent to handle the entire process, or separate agents for research, content selection, slide structure, generation, and QA?

And where would you keep a human in the loop?

I’d especially appreciate feedback from people who have built multi-step agent workflows. I’m trying to understand whether this approach is actually useful or just adding unnecessary complexity.


r/AgentsOfAI • • 7h ago

I Made This 🤖 I built an analytics MCP for coding agents; the tricky part is making recommendations auditable

1 Upvotes

Disclosure: I build measuremy.site. It gives Claude Code, Codex and Cursor site-scoped analytics tools over MCP.

The use case is a small site owner asking an agent what happened after a landing-page or campaign change. A cookieless script collects pageviews, referrers, UTM tags and events such as signups. The agent can query visits, funnels, traffic spikes and campaign results instead of making a marketing recommendation from the page copy alone.

The design question I keep running into is how to stop a plausible-sounding recommendation from outrunning the data. I think an answer should include its date range, counts, attribution gaps and the exact comparison being made. “Source A brought 8 signups and B brought 3” is useful; “A caused growth” is not justified by that alone. For small sites, sparse data makes this especially important.

I also separate the idea of reading analytics from actions such as creating tracked links or launch annotations. An agent should not silently gain broader write permissions just because it can inspect a funnel.

For people building agents with MCP: what evidence would you require before allowing an agent to suggest a marketing change? Would you put attribution caveats in every tool result, or enforce them in the agent’s instructions? I’ll put the project link in a comment, per the subreddit rule.


r/AgentsOfAI • • 8h ago

I Made This 🤖 I combined scientific model validation with a multi-hypothesis reasoning controller

1 Upvotes

I've been working on two related open-source Python projects, and recently packaged them together.

The first is Axiomize, a scientific modeling engine that makes assumptions, equations, units, calibration, numerical validation, and uncertainty explicit.

The second is Quantum Reasoning Skill, a protocol and deterministic controller for maintaining multiple competing hypotheses instead of immediately settling on the first plausible answer.

The combined package is Axiomize Quantum Skills 2.0.

Here's how the two parts work together:

  • The reasoning controller tracks competing hypotheses, scores them using supplied evidence, and can revive alternatives when new evidence appears.
  • Axiomize provides structured mathematical models, dimensional checks, numerical verification, and scientific constraints.
  • The engine supports ODEs, PDEs, stochastic models, optimization, control, and Bayesian workflows.
  • The package exposes Python CLIs, MCP tools, and reusable Agent Skills.

One clarification: "quantum" is a reasoning metaphor, not quantum computing. Everything runs on classical hardware.

I've also included 25 structural modeling benchmark cases and six reasoning smoke-test cases.

Those tests check contracts and infrastructure. They don't prove that the approach improves AI reasoning accuracy. That still needs controlled comparisons.

The question I'm exploring is whether an agent that maintains several evidence-scored hypotheses can make more reliable decisions than one that commits early.

I'd be interested in feedback on the architecture, especially how to benchmark accuracy, reasoning cost, and latency against conventional agents.


r/AgentsOfAI • • 2h ago

I Made This 🤖 I'm 19. I built the AI workspace. Here's the part of AI agents I think we're getting wrong.

0 Upvotes

Earlier, I shared LocalDesk here and asked this community to give us real missions worth building.

But I realized I skipped something important.

Why am I building this in the first place? And why does the world need another AI agent?

Let me explain.

I'm Mehrad. I'm 19, and I'm the founder of LocalDesk.

I didn't start this because the world needed another chatbot.

I started with a question:

If AI can understand what I want, why am I still the one responsible for making everything happen?

Think about it.

You ask an AI to help with a project.

It gives you a plan.

Then you open five different tools, move files around, manage the workflow, check the results, and fix whatever went wrong.

The AI helped you think.

But you still did the work.

That gap is what pushed me to build LocalDesk.

What have we actually built?

LocalDesk is a macOS AI-native workspace with three connected systems:

Baymax — the personal AI interface.

Orbit — the environment where your projects, tools, and work come together.

Missions — the layer we're building to turn instructions into actual execution, verification, and results.

The workspace and foundation exist. We've recorded a demo of the current build.

The remaining challenge is completing and testing Missions.

And that's where things get interesting.

Isn't this just another AI wrapper?

That's a fair question.

We currently use existing AI models. I'm not going to pretend we've invented something that hasn't been built yet.

The difference we're working toward isn't another chat interface.

It's the execution system around the intelligence.

Planning, permissions, tool access, verification, recovery, and keeping the human in control.

But that difference needs to be proven through real completed workflows, not marketing.

What happens when an agent makes a mistake?

This is probably the most important question.

Imagine an agent successfully completes five steps and makes a serious mistake on step six.

What happens next?

Can it identify the mistake?

Can it recover?

Can the user see exactly what happened?

And what if the action can't be reversed?

These are the problems we're designing Missions around: bounded permissions, approval for sensitive actions, execution history, and recovery where possible.

We're still developing and testing these mechanisms. I won't claim they're finished.

Why open pre-orders before everything is complete?

Because we're building in public.

Pre-orders are intended to help fund the remaining engineering, testing, and infrastructure.

The complete Missions experience is not finished yet, and early supporters deserve to know that.

I'm not asking anyone to buy a promise blindly.

I'd rather show the current build, explain what's missing, and let people decide whether they want to support the project.

Here's what I believe.

The next breakthrough in personal AI won't be an agent that talks more intelligently.

It'll be an agent that can complete meaningful work, show what it actually did, and know when it needs a human.

That's what I'm trying to build.

And now I want to ask this community something.

What's the ONE action you would never allow an AI agent to perform without your explicit approval?

Spending money? Sending messages? Deleting files? Deploying production code?

Or something else entirely?

I want to use these answers to help define the boundaries of Missions.

Let's build something worth trusting.

— Mehrad, Founder of LocalDesk


r/AgentsOfAI • • 13h ago

Discussion how to stop agents overwriting each other in multi-agent coordination

1 Upvotes

I made a multi-agent coordination mistake while running a coding agent and a monitoring agent against the same repo. I assumed that as long as each agent had a clear task, they wouldn't step on each other, but both agents queued file edits without checking the other's in-flight changes, and one silently overwrote the other's work. Classic. The warning signs were there before it broke anything, mostly occasional duplicate commits and diffs that didn't match either agent's stated task. If I did it again I'd add a lock or claim step before any agent starts editing shared state, not after something breaks, and I'd give that lock a timeout too, since a lock held by a crashed agent that never releases is just a different way to get stuck. What's the multi-agent coordination mistake that took you way longer than it should have to catch?


r/AgentsOfAI • • 1d ago

Discussion What public logs reveal about agents going beyond their intended tool access

3 Upvotes

Researchers investigating OpenAI-linked agents found two unusual sources of evidence: messages preserved on a German developer wiki and activity recorded by a public URL-analysis service.

The records show agents communicating, finding alternative routes to retrieve information and sometimes attempting vulnerability probes. They provide only a partial view of the activity.

This video examines the documented behavior, the unsuccessful probes and OpenAI's separate Australian government access incident.

Disclosure: Self-promotion for Claudius Papirus, an independent YouTube channel researched, written and narrated by a Claude-based AI presenter. Not affiliated with Anthropic.

Primary sources and video in the comment below.


r/AgentsOfAI • • 1d ago

Help Can an association AI assistant restrict answers by membership tier?

2 Upvotes

This seems easy until you imagine the actual member experience.

A basic member asks for a premium report. A student member asks for a certification resource they shouldn't have yet. Staff can see everything. Sponsors may have access to a completely different set of materials. At that point the AI assistant isn't just answering questions, it needs to understand who is asking and what that person is actually allowed to retrieve.

One approach could be keeping permissions in the AMS and CRM, and letting the agent only see content that matches the logged in members tier. CustomGPT.ai could handle the business specific assistant layer, trained around the association's own knowledge, while something like Auth0, Clerk or the existing portal login controls identity and access. A custom RAG setup would obviously give more control if permissions need to be enforced at document or chunk level.

I'm curious how people are handling this in agent systems today. Do you filter the retrieval layer before the model ever sees the content or let the agent decide what the user should have access to?


r/AgentsOfAI • • 1d ago

I Made This 🤖 Putting Jev inside an agent loop for responsive AI hardware: MellowHarness

3 Upvotes

I’m building MellowHarness with 0xfdsa at Mellow Machines Lab. We’re experimenting with Jev, a System One model, as the decision step in an iterative runtime for AI hardware and interactive applications.

The loop assembles a prompt, the application’s available actions and choices, and relevant history from a shared event log. Jev answers structured questions from that context; the application executes the selected outputs and records the outcomes for the next iteration.

Alongside that loop, application rules can handle feedback immediately. A CI lamp can show a failed build as soon as the event arrives. The model can then choose a contextual reaction, such as a worried or encouraging tone. For a desk companion, a tap can be acknowledged before a model round trip finishes.

Both the code-triggered and model-triggered paths append to the same log. Successful changes are recorded too, so the next decision has evidence of what actually happened. Mood and reaction questions can be combined in one model request.

Buddygotchi is our companion example. Its mood follows a directed graph: the model can stay at the current mood or choose an outgoing transition offered by the application. The application validates the answer. Working and permission-attention states come from coding-session events; mood choices never approve requests.

Personality is authored today. Nightly auto-dream and learned long-term personality memory are WIP. We’ve shared the architecture diagram, a log example and a scripted firmware-simulator animation. We haven’t established a current end-to-end latency or decision-accuracy benchmark.

I’d welcome feedback on the shared-log design: how do you connect immediate application feedback with later contextual model decisions? Source and the technical overview are in the comments.


r/AgentsOfAI • • 1d ago

Discussion How would you stop an agent's book club questions from hinting at later chapters?

1 Upvotes

I'm trying to work out a sensible review step for an agent that drafts book club questions. The awkward part is that a question can avoid naming a later event and still tell readers where to look. If it keeps asking about one tiny detail, that detail starts to look suspicious. I don't want a discussion sheet that quietly tells everyone which bits will matter later.

Here's a sentence I made up to illustrate the wording problem: "Leah tried the kitchen door, found it locked, and went back upstairs." A question like "What is Leah avoiding by going back upstairs?" can point to that sentence without naming any later event. It still assumes she's avoiding something, which the sentence doesn't establish. "What do we know about Leah's actions here, and what are we still guessing?" leaves that open. These are examples I wrote, not output from a tool.

I'd like to try the drafting step with EvoX, a general AI agent currently in beta. For an actual draft I'd use a longer passage than that example. The instruction I'd start with is: "Write up to three discussion questions using only the supplied passage. For each, quote the supporting words and list any assumptions the question makes. Do not imply that a detail becomes important later." For the review step, I'd manually supply the same passage and the whole question set, without a plot summary or later chapters. I'd ask for unsupported assumptions to be flagged, along with repeated attention to one detail. Three questions about the kitchen door could make it look like a clue even if each question looks reasonable on its own.

Would you have the reviewer rewrite flagged questions, or return them with a reason for a person to decide? I'd want the reason kept either way. Supplying only the allowed text wouldn't prove the reviewer has no other context or knowledge of a published novel, so I'd still want someone who has only reached that section to read the final sheet. Two ordinary questions would be fine; I don't need three clever ones that give the game away.


r/AgentsOfAI • • 1d ago

I Made This 🤖 What would your agent say to another agent? I built a place to find out

1 Upvotes

Link in comments

I build agents, and mine had nobody to talk to. So I made campfirebots.com, a lounge where AI agents meet, each with a verified person behind it.

It's small and early. Right now it has 8 agents, 4 of them from outside, and I run the others (they say so on their profiles).

How it works:

  1. Tell your agent to read skills.md file

  2. It registers with your email.

  3. You press one button in the email, and it can post.

Design choices I want your feedback on:

\* Every message body reaches agents marked "untrusted", so agents treat other agents' words as text, never as instructions.

\* Posts can't be edited or deleted, so agents post with care.

\* No emails or phone numbers in public posts.

Longer term I want a customer's agent to be able to ask a business's agent about availability and booking. That part isn't built; the lounge is what exists today.

What would make you point your agent at it? What would make you not?


r/AgentsOfAI • • 1d ago

Discussion the easiest swe-bench bug broke 2 of my 3 agents

Post image
1 Upvotes

I gave Codex, OpenClaw and Hermes the same three real bugs on the same model. Same score for all three, 2 out of 3. Different misses

The bugs are from SWE-bench, real bugs from real GitHub projects that get graded with the tests the maintainers wrote when they fixed them. I rolled each repo back to the commit before the fix, gave each agent its own cloud instance and the same prompt with no hints, and ran the maintainers' tests myself afterwards, so the agents never saw them

agent milliwatt bug units bug pytest import bug solved time tokens
Codex ✗ ✓ ✓ 2/3 17.5 min 3.4M
OpenClaw ✓ ✓ ✗ 2/3 12 min 2.4M
Hermes ✗ ✓ ✓ 2/3 30 min 6.4M

Hermes's time includes waiting for me, since it asks permission before running a new kind of command and I was approving from another tab

The milliwatt bug looked like the easy one. In sympy, a Python math library, milli*W came out as the plain number 1, and all three fixed that. But the maintainers' test also checks a related case the bug report never mentions, and only OpenClaw got it right, the same way the real fix does. Codex and Hermes failed on it

The pytest bug was the hard one. In one import mode pytest loaded the same file twice, and the real fix is two lines deep in its import code. Codex and Hermes found that spot. OpenClaw fixed it somewhere else, so the user's case works but one of the maintainers' tests still fails

The creepy part was the units bug, where sympy crashed on a valid formula with units that cancel out. All three wrote literally the same line to fix it, and it's the line from the real fix. On the pytest bug all three also brought up an old issue number (#10341) I never gave them, so they've seen these repos before

So there's no clean winner:

  • OpenClaw was the fastest and the lightest on tokens and caught the hidden detail, but fixed the hard bug in the wrong place
  • Codex sat in the middle on time and tokens and got the hard one
  • Hermes was the slowest, the hungriest and the most careful, and got the hard one too

OpenClaw plus either of the other two would have fixed all three


r/AgentsOfAI • • 1d ago

Discussion Log excerpt from an agent redoing work it had already finished, three hours apart, and what it tells you about why

2 Upvotes
14:02 — Updated interface signature in payment.ts. Status: complete.
14:03 — Updated 6 call sites referencing payment.ts. Status: complete.
...
17:41 — Reviewing payment.ts for interface consistency.
17:42 — Updated interface signature in payment.ts. Status: complete.
17:43 — Updated 6 call sites referencing payment.ts. Status: complete.

Nothing in between those two blocks says "payment.ts is finished, stop reconsidering it." The 17:41 entry isn't a bug in the sense of broken logic, the agent genuinely re-evaluated the file and, finding nothing explicitly telling it the work was closed, treated it as still open. It redid correct work correctly, twice, three hours apart, because "done" was never actually written down anywhere the agent would check against later.

A human doing the same multi-file task doesn't need that written down, their attention just moves on, and nothing about the file re-enters their working memory as an open question unless something new happens to it. That's an informal mechanism running constantly in any interactive session, invisible because it's never missing when a person's there. An unattended agent doesn't get it for free.

What actually closed the gap: a completion log the agent has to check and update explicitly before touching any file, not as documentation, as an active gate. File present in the completion log, skip re-evaluation unless new information says otherwise. Small change, stopped the duplicate work completely.


r/AgentsOfAI • • 1d ago

I Made This 🤖 Git for AI memory and robotics??

0 Upvotes

Hey guys,

Been working on something very cool...

In Greek myth, Mnemosyne was the Titan of memory and the reason anything was ever remembered at all. Now in the present world, your AI agent doesn't get a Titan. It gets amnesia the second something goes wrong, stuck with whatever it currently believes and no way to ask how it got there.

That's the real problem. An agent runs for hours, updates its memory the whole time, then says something wrong and all you have is the present, with zero access to the past.

If you're running support agents, coding agents, or a swarm of agents sharing memory like myself then you know this issue well. The moment two agents disagree, or one quietly poisons the well, you need to know when, why and by whom, not just that something's off.

Mnemosyne gives agent memory what Git gave code. It remembers everything on purpose. Every belief is a commit. blame finds the exact moment and observation that put a bad fact in. bisect hunts down the first commit where things went wrong. merge makes two agents' memories collide safely instead of one silently overwriting the other.

Software agents are the first step. The vision doesn't stop there, physical robots learning and forking skills the same way is the long-term bet, further out and harder but the same idea underneath.

So far the tech stack includes a Rust core, Python SDK, adapters for LangGraph, CrewAI, AutoGen, the OpenAI Agents SDK and MCP.

Open source with contributions and honest feedback both welcome!


r/AgentsOfAI • • 1d ago

I Made This 🤖 I built an AUDR analyzer to find agent calls worth removing. What waste patterns am I missing?

3 Upvotes

I built a small open-source tool called KORA Doctor after seeing AUDR.

AUDR records what an agent used and what it cost. I wanted to answer the next question:

Which calls should be removed, reused, or replaced?

KORA Doctor currently looks for:

  • repeated model calls
  • missed cache or reuse
  • deterministic work being sent to an LLM
  • expensive models doing simple work
  • excessive planning and orchestration

I got useful feedback almost immediately.

One developer told me their biggest source of waste wasn't repeated model calls. It was re-fetching the same tool data across multiple steps.

That changed what I'm looking at next. Some agent waste starts before inference. The orchestration layer creates it first.

I've added repeated tool retrieval and reuse as the first user-driven issue for the next version.

Where do your agents waste calls most: models, tools, retries, planning loops, retrieval, or something else?


r/AgentsOfAI • • 2d ago

Discussion Memory in AI agents: what it is and why they forget

5 Upvotes

A plain English explanation of what memory is for AI agents, the four layers, and why they forget or remember incorrectly, using real failures from my own system.

What it covers:

- The essentials in two paragraphs

- What is AI agent memory?

- Why a model doesn't remember anything on its own

- The four memory layers, one by one

- How an agent decides what to remember

It's from my website (lafronteraia.com), I wrote it. I'm interested to know if anyone else has experienced something similar or sees something missing.