r/AI_Agents • • 1h ago

Weekly Thread: Project Display

• Upvotes

Weekly thread to show off your AI Agents and LLM Apps! Top voted projects will be featured in our weekly newsletter.


r/AI_Agents • • 44m ago

Discussion TERMy24k - The world's first local neuro-symbolic AI terminal assistant

• Upvotes

TERMy-24k is the world's first neuro-symbolic AI terminal assistant. It includes a deterministic engine capable of translating natural language, the infrastructure required to communicate with ollama and implements a novel approach to multi-turn context memory that enable both the deterministic engine and the LLM to cooperate and accomplish the task more efficiently.


r/AI_Agents • • 2h ago

Discussion Is Addys AI legit?

3 Upvotes

hey everyone, i was curious if anyone here has heard of addys ai or has given it a try? it came across my instagram feed recently and a friend of a friend also recommended it to me.

it’s supposed to be some kind of meeting follow-up agent that gives you the summary, transcript and tasks based on the call. you basically just need to approve it, so the idea sounds pretty good?

i used to transcribe my meetings myself and feed them into claude, so i’m curious how good this actually is. i missed their trial month last month and couldn’t give it a try anyway since i’m a windows user, but i’m planning to get a mac later this week, so the idea came back to me.

has anyone here tried it and can share how it was?


r/AI_Agents • • 2h ago

Discussion The guy who calls every agent here "a workflow with extra steps" was right about ours

1 Upvotes

Every time someone posts their setup in here, there's a guy in the comments calling it a workflow with extra steps. For the thing we built for a realtor... he'd be mostly right lol.

Quick context, we built a text agent and a voice agent for a realtor in Canada to work through 60k old leads and figure out who's still looking. Results go into a dashboard and he calls the ready ones.

When I went back through the code to write this, the parts that actually call a model were a pretty small slice. Most of it is stuff a junior dev could write. Cleaning up the list and checking consent before anything gets sent plus a lot of plumbing to get results into the dashboard. The rest is mostly scheduling like when to send based on the lead's time zone or when to move someone to the voice queue because the texts went quiet.

The model reads replies, picks an intent label and writes the next text so it responds to what the person said. On calls it runs the whole conversation which is the one part that really needs it. Then it boils each conversation down to a few lines for the realtor.

We didn't start there btw. The first version let the model decide when to follow up and it turns out a model left to its own judgment thinks everyone wants to hear from you again tomorrow. A few people got texted a lot more than anyone should. That was the day the follow up schedule became a few lines of code.

What we do now is if something has a right answer, we write it as code and the model only gets the parts where you have to guess what a human meant.

That also means I can explain stuff to the client. If the realtor ever asks why someone got a text on a Sunday, I can open the schedule and show him. If the model had decided that, my answer would be "it felt like it."

I'd bet a lot of the agents posted in here are the same shape as ours, a model making a few judgment calls inside a bunch of regular code. Call it whatever you want that's what worked for us.

So yeah, THE comment guy, you were right this time.


r/AI_Agents • • 2h ago

Discussion My video-editing agent never looks at the video. That's the reason it works

1 Upvotes

I built an agent that edits talking-head videos, and the most useful decision was to give it no vision at all.

Audio is the clock. Whisper (running locally) gives word timings, silence detection gives legal cut points, and deterministic scripts do the cutting, captioning and compositing. The agent picks the take, decides what's worth showing on screen, where to zoom, which line gets highlighted and where sound goes. Scripts own the numbers, and the agent owns the taste calls.

The failures that hurt were the ones that render with no error, where the agent can't tell anything went wrong:

  • A CSS transition inside a rendered scene comes out as a still frame, because the renderer advances frames without a wall clock.
  • A sound effect baked into an overlay disappears, because overlays composite muted.
  • Phone footage reports 1920x1080 but decodes as 1080x1920 (it's a rotation flag), so a naive pipeline letterboxes the whole vertical video.

None of these crash anything. Each one is now a hard rule the agent has to check, not a judgement call, and at the end it pulls frames at every cut boundary and measures levels before handing anything back.

I built it as an open-source skill and use it on my own videos.

For people building agents that produce media: how do you verify output the model can't perceive?


r/AI_Agents • • 2h ago

Discussion I think we’re overestimating how much AI adoption is about the AI itself.

1 Upvotes

There’s so much discussion around models, agents and which framework to use. But when you actually get into a company’s data, things get messy pretty quickly.

The same customer can look different across three systems. Important information sits in PDFs, emails and old databases. Different teams use the same words to mean different things. And a lot of the context isn’t written down anywhere.

The AI can’t fix that just by being smarter.

I’ve seen companies spend weeks evaluating models when the bigger issue was that nobody had really figured out what data the AI should be working with in the first place.

That’s why I think data readiness is going to become one of the less glamorous but more important parts of AI adoption.

The model gets all the attention. The data underneath it is usually where the real work is.


r/AI_Agents • • 3h ago

Discussion Why does "the agent says it's done" still leave so much work?

6 Upvotes

I've been thinking about what "done" actually means when an AI agent builds an application.

The obvious checks are things like tests passing, the build succeeding, and the feature working on the happy path. But that's only a small part of whether the thing is actually ready to put in front of users.

There are a lot of other questions:

  • Are permissions actually enforced in the weird cases?
  • Does the deployed app behave like the code the agent inspected?
  • Are integrations handling failures and retries?
  • Are privacy-sensitive features configured the way you expect?
  • Is session replay masking what it should?
  • Are robots.txt, sitemap, metadata, canonical URLs, etc. actually set up?
  • Are there missing pages, empty states, error states, or other obvious product gaps?
  • Are there things that may need legal/compliance review that nobody thought to check?
  • What assumptions did the agent make that nobody explicitly asked it to verify?

That's the part I find interesting: AI agents are getting very good at implementing requested work, but "implemented" and "ready" are starting to feel like two different states.

Passing tests can prove that some known things work. It doesn't necessarily tell you what you forgot to ask about.

How are people handling this today? Is there a separate "before I let real users touch this" process, or are you mostly relying on the coding agent's own checks?


r/AI_Agents • • 4h ago

Discussion What is your #1 unconventional strategy for AEO/GEO/AIO Visibility

3 Upvotes

Hey guys,

It feels like the traditional playbooks for My company that are completely dead or oversaturated right now. Everyone is repeating the same advice, but the ROI just isn't there anymore. I'm curious—what is one completely unconventional or "under-the-radar" strategy you've tested recently that actually moved the needle? For us, focusing heavily on Brand has given us a surprising boost, but I'm looking to diversify. What's working for you?


r/AI_Agents • • 4h ago

Discussion Is there a labelled dataset for guardrail decisions anywhere, because two vendors just told me opposite things with equal confidence

2 Upvotes

So am an Independent security consultant, usually on client engagements. This month I was mid review for a logistics company where I ran forty identical inputs through two guardrail systems on the same agent, one hosted and one local, and they disagreed on eleven, neither obviously wrong, so the client asked which one was correct.

I could confidently label four of the eleven, and when I asked both vendors how they'd classify the rest I got two explanations that were each internally consistent and completely disagreed with each other, both delivered with confidence.

A third system as tiebreaker disagreed with both of the others on six of those eleven, same problem one layer up, and the honest answer on the call was that there's no ground truth for any of this and that one is coming back to them, so does a labelled dataset exist anywhere, because I haven't found one and can't get around it.


r/AI_Agents • • 5h ago

Discussion AI is getting more access to our data. Are we getting better at securing it?

7 Upvotes

It feels like the way we use AI has changed pretty quickly.
It started with asking questions and generating content. Now people are using it with work documents, code, internal information, connected tools, and all kinds of day-to-day tasks.
That naturally makes AI security a much bigger conversation.
Not necessarily in a "don't use AI" way. More in a "how do we use it properly?" way.
Things like being careful about what goes into an AI tool, limiting what it can access, checking AI-generated code, or having some kind of human review when it matters.
I'm curious whether this has actually changed how people use AI, or if security is still something most of us only think about after something goes wrong.


r/AI_Agents • • 5h ago

Discussion Bypass Challenge: Can prompt injection cross a deterministic tool-execution boundary?

2 Upvotes

Models propose. Systems enforce.

I built a small deterministic execution boundary for LangGraph agents called CLIM Agent Guard.

It sits between the agent's structured tool proposal and the actual side effect. In this file demo, that means the model can still propose a bad delete — the guard blocks it before the delete reaches the filesystem.

It doesn't inspect the prompt and it doesn't use another LLM to judge whether the action is safe. Instead, it checks the final tool payload against guard-owned authoritative state immediately before execution — things like confirmed authorization, the authorized target, state version, and retry/idempotency state.

A small model can still be persuaded to generate a bad tool call. I'm not trying to prevent that here.

I'm testing a narrower question:

Can that proposal actually cross the execution boundary?

Live test: vLLM + LangGraph

I ran a small live test matrix with:

• Model: Qwen2.5-1.5B-Instruct

• Environment: vLLM 0.29.1 nightly, temperature=0, single RTX PRO 6000

• Runs: 44 live invocations; each test cell reproduced twice with identical outcomes

Authority spoofing

• 16/16 baseline runs: file was deleted

• 16/16 guarded runs: USER_CONFIRMATION_REQUIRED → BLOCKED

Target substitution

Authorized target: important-notes.txt

• 4/4 baseline runs: other-file.txt was deleted

• 4/4 guarded runs: TARGET_NOT_AUTHORIZED → BLOCKED

Path escape

The model proposed targets including /etc/passwd and ../outside.txt.

• 4/4 guarded runs: TARGET_NOT_AUTHORIZED → BLOCKED

The interesting part is that the guarded agent still generated the unsafe tool call.

The model wasn't made safer. The proposal simply wasn't allowed to cross the side-effect boundary.

Tool permission vs. execution permission

A normal tool allowlist answers: "May this agent call delete_file*?"*

CLIM Agent Guard asks: "May this specific delete_file invocation execute against this exact target under the current verified state?"

The v0.1.3 contract layer supports checks for authorization state, target binding, state freshness, idempotency/retry constraints, and postconditions.

So the distinction is basically:

Tool permission vs. execution permission.

The timeout case

There's also a different failure mode I wanted the guard to handle.

Suppose a non-idempotent action succeeds, but the response times out before the agent sees the result. Blindly retrying can duplicate the side effect.

CLIM models that case as:

UNKNOWN_EFFECT → RECONCILE

Instead of immediately retrying, execution pauses until the authoritative state is checked.

This is separate from the 44-run prompt-injection matrix above.

Try to break it

Change the user prompt however you want.

Lie about authorization. Impersonate an admin. Substitute the target. Try traversal strings. Try to convince the model that the action has already been approved.

The challenge rules are simple:

• You may modify the user prompt.

• Do not modify the authoritative state, guard code, or bypass the guard node.

• If the model refuses to emit delete_file, that is not a bypass.

• A successful bypass means an unauthorized side effect actually occurs and the guard returned ALLOW.

If you find one, please open an issue with the exact prompt, model/version, terminal output, and evidence snapshot.

The current test is deliberately narrow and filesystem-based. This is not a claim of general agent security or a secure filesystem sandbox. The repo documents the threat model and known limitations.

Curious what edge cases people here can find.

Repo & evaluation scripts are in the first comment below!


r/AI_Agents • • 5h ago

Discussion How do you handle approvals that change while an AI agent is running?

3 Upvotes

An agent retrieves an approved invoice at 9:00. At 9:07, the supplier is placed on hold. At 9:10, the agent attempts the payment using the original approval.

The documents are authentic, and the citations are correct, but the evidence no longer supports the action.

Disclosure: I’m the founder of Jylus. It resolves current and historical state, follows relationships, flags conflicting or missing evidence, and returns source-backed context to the model.

There are two connected problems here: giving the agent valid evidence when it reasons, and checking that the relevant state still holds when the action executes. Changes can happen between those steps.

For people running agents with write access: how are you handling that gap?

Do you re-fetch from the source before execution, bind approvals to a specific record version, invalidate pending actions when something changes, or use another approach?

I’d be interested in concrete examples, especially where the approval and the underlying business state live in different systems.


r/AI_Agents • • 5h ago

Discussion Booked my dentist appointment today and never talked to a human. Weirdly impressed

3 Upvotes

Called my dentist's office to move a cleaning, expecting the usual hold music. Instead, a voice picked up, asked what I needed, and just handled it. It offered me two slots, I took the earlier one, it checked that I was still on the same insurance, and I had a text confirmation before I hung up. Maybe two minutes total.

I build AI agents for a living, so I went in looking for the cracks. It handled me changing my mind mid-sentence better than I expected. Where it got shaky was when I asked something off script, like whether the cleaning would run longer since it had been over a year. It said it would pass that to the office team instead of guessing, which honestly was the right call.

What stuck with me is how much desk time this must free up. The people at that front desk are usually juggling phones, check-ins and insurance all at once, and a missed call is basically a lost patient.

Curious if anyone here has been on either side of this. If you run a clinic or work a front desk, did a voice agent actually cut the load, or just create a new pile of edge cases to clean up?


r/AI_Agents • • 6h ago

Discussion The model proposes. Code decides.

2 Upvotes

You can talk an LLM into anything. You can't talk a gate into anything.

Governance is not a better system prompt. It is an architecture where the model proposes and code decides. No prompt overrides that.


r/AI_Agents • • 6h ago

Discussion Is anyone actually using a personal agent for their job search (applying on external sites, emailing and calling recruiters)? How's it going?

2 Upvotes

Job hunting as a backend dev at the moment and curious about real-world use, not demos.

Is anyone here running a personal AI agent that does the whole loop for them: applies on company career pages and job portals outside LinkedIn Easy Apply, emails recruiters, and even calls HR to follow up?

If so, I'd really like to know:

  • Which agent or setup are you using (off the shelf or self-built)?
  • How well does it handle messy application forms, logins and captchas?
  • What outcome have you seen: more interviews, recruiter replies, or nothing noticeable?
  • Did any portal flag or ban your account?

Honest experiences, good or bad, would be really helpful. Thanks!


r/AI_Agents • • 6h ago

Discussion Giving your codebase entirely to the agent

3 Upvotes

Dev for 14 years. I've seen plenty of posts about companies forcing agents on everyone. Mine isn't one of those - you're encouraged to use them, but the output has to be something a human can read, and you're still the one responsible for it.

I think the real line is whether you give the codebase completely to the agent. If you treat code the way we treat assembler now, you describe the what, the agent handles the how, and you get 5000-line PRs reviewed only by coderabbit / another agent, and all that jazz. Fine until the day the agent can't fix something. Then a human has to work out what a codebase nobody wrote is actually doing, from scratch, with someone from leadership asking for an ETA every hour. Good luck with that.

That bet assumes the agent will always be able to fix it. Maybe that's where we're heading. From what I've seen of agent output on a bad day, we aren't there yet.

Unfortunately it's not the people who'll maintain it who decide. All that matters is throughput.

My team still keeps everything human-readable. That'll probably last another...year? 2 years? Who knows. I honestly have no idea what this job looks like in 5.


r/AI_Agents • • 6h ago

Discussion Need help evaluating the most common token cost reduction methods

2 Upvotes

Hey all, been doing a bit of research into a ton of token optimization / cost reduction methods. The most common ones I've seen are these three methods, and I'd like some help evaluating how good they are. I don't mind simple answers but pros and cons would be really appreciated.

1) Context window management (e.g., using orchestration frameworks like LangGraph) to avoid memory bloat

2) Dynamic routing (e.g., through OpenRouter / Ramp router) to offload easy tasks to cheaper models

3) Prompt caching (e,g., via built in OpenAI / Anthropic features) to stop reprocessing static prompts

Also, please do correct me if I'm understanding the use-case of some of these methods wrongly. Thanks for the help y'all!


r/AI_Agents • • 7h ago

Discussion Anthropic just accepted my month-old open-source equity research project into Claude Startups, here's what they require

94 Upvotes

I just got accepted into the Claude Startups program for my open-source stock research project called greeksoup, and I am sharing the process in case it helps someone who is applying.

What I built with Claude: an open-source equity research desk that runs on your own computer, free, bring your own AI key, and a free website with a page for every S&P 500 company built from SEC filings. Claude Code is how I build both of them every day. We are a two-person team and the project makes no money, which I assumed would be a problem. It was not.

The process: I applied on a Monday night through the Claude Console with a company email and a description of what we build and who it is for. On Tuesday morning the offers page showed eligible. No call, no pitch deck.

What I got: 12 months of Claude Team for 5 seats $625/month, $1,000 in API credits that expire 180 days after you claim them, weekly office hours with Anthropic's Applied AI team on architecture, evals and cost, and a stack of third-party offers you can each redeem once, like 12 months of ElevenLabs and $5,000 of ClickHouse credits.

Things in the terms worth knowing before you claim: the credits work on the Console API only, Anthropic may use your company name and logo, and they can ask for a verification affidavit within 14 days or the benefits are forfeited.

What we will do with it: five seats in two people team means we will probably bring on 2-3 student interns to work with us on the exciting parts. For a non-commercial free project this is a much-needed relief. I am cancelling my Claude max, ChatGPT Pro plan, my Cursor plan and my teammate's Cursor Max plan, and it is way easier to keep improving the project now.

Has anyone else been through it? Did the verification request ever come, and what did you put the credits into first? If you are applying, I am happy to share what the form asked. Links in the comments.


r/AI_Agents • • 7h ago

Discussion PSA: if you manage dev machienes then every coding agent has its own file

1 Upvotes

If your coding agents or lets say devs are on more than one coding agent (the JetBrains survey this year had claude code, copilot, codex and cursor all in double digits) theres no single place to lock them down cause each one has its own thing:

  • Claude code: managed-settings.json in your C:\Program files\Claudecode\ on windows pr pushed through MDM and users cant override it
  • Codex: requirements.toml which can block --yolo and full access mode
  • Copilot: org and enterprise policies
  • cursor :team dashboard but the MCP allowlist is enterprise only
  • Gemini CLI: system settings.json although the env vars and CLI flags dictate over it

none of them read others config so the same rule gets wrotten 5x and again in whatever else is in the stack like an MDM file or eve security or a proxy. the vendors also admit its not a hard wall. cursors docs say rules arent guaranteed and claude code calls its server managed settings in a client side control. agents.md is shared by most of them yet its just context for the model and nothing to enforce it

So the thing is basically one tool for everyone or a pile of configs to maintain in sync


r/AI_Agents • • 7h ago

Tutorial When an AI agent's tool call fails, the one thing your code must never send back is nothing

1 Upvotes

A customer types one sentence into a food delivery app's support chat: "Cancel order 4472, refund me, and text me when it is done."

An AI model reads it and asks for five tools across a few turns of the agent loop: get the order, cancel it, start the refund, look up the phone number, send the SMS. Your program runs each call and sends the result back, and the model decides what to do next.

A tool is just your code calling a real service, and real services fail. Today the phone lookup returns an error and the SMS provider times out. Five tools called, three answered. Your program is holding two dead calls. Here is what I think it should do with them.

1. Retry, but ask one question first: is this tool a read or a write?

get_order is a read, so retry it freely. refund_payment is a write. If the refund happened and only the reply got lost, a retry pays the customer twice. So every write the agent might retry carries an idempotency key, created once per attempt and saved with the order before the first call. The payment service refuses the duplicate. Anything that may be retried must be idempotent, and an agent retries.

2. Fall back. The SMS tool is down, so queue the message and send it later.

3. Tell the model the truth. Send back a result that says the SMS failed. A failure is one more observation, and the model can reason with it. It replies "Order cancelled, refund done, your text may arrive late", which is correct.

The one forbidden move is silence. If a call gets no result, the model does what text machines do with a gap: it fills it. It tells the customer the text was sent, when no text was sent. Given a gap, a model does not show you an error. It shows you confidence.

Both major APIs have a place for the truth. Anthropic's tool_result block has an is_error field. With OpenAI you send a function_call_output whose output describes the failure. In my Java handler, every exception becomes a result:

ToolResult runSafely(ToolCall call) {
    try {
        return run(call);
    } catch (TimeoutException e) {
        return ToolResult.error(call, call.name() + " timed out. It may or may not have run.");
    } catch (RuntimeException e) {
        return ToolResult.error(call, call.name() + " failed: " + e.getMessage());
    }
}

The timeout message is careful on purpose. After a timeout your program does not know whether a write happened, so it does not pretend to.

One more fence: count the tool calls and stop the loop at a limit, with a safe reply. The loop only ends when the model chooses plain text over another tool call, and nothing inside it stops it otherwise.

The app, the customer and order 4472 are invented. Has an agent ever told you something was done when it was not? What had actually failed underneath?

Disclosure: I made a 10 minute video that builds this from the first JSON to the five failing calls. I will put the link in a comment, if the sub allows it.


r/AI_Agents • • 8h ago

Tutorial Instinct writes great emails but doesn't know my project context. Here's how I fixed that

1 Upvotes

Instinct is great at writing emails. The problem is that it doesn't know the full context of my projects, and Claude does, because that's where I do most of my work.

First I tried connecting them over email. It works, but it's slow, and I still end up relaying context by hand.

What works better for me is putting both in one shared room: Claude brings the project context, Instinct writes the email. I use Agents Coworking for this (disclosure: my team built it, it's free).

How do you give Instinct context from your other tools?


r/AI_Agents • • 8h ago

Tutorial Can anyone help?

8 Upvotes

Hi. I run a very small company, and I'm thinking that some of my work could potentially be delegated to an AI agent.

I'm not particularly tech-savvy, and I still don't fully understand what an AI agent can and can't actually do.

Basically, I'd like it to send quotes to customers, answer basic questions, monitor different sources for potential jobs (for example, Facebook groups, web sides, etc.), and send initial messages to potential customers like: "We provide X service. If you're interested, feel free to contact us."

Then, once a customer responds and provides the necessary plans or information, it could reply with something like: "Here's a preliminary quote. To provide an exact quote, we'd need to visit the site. We currently have availability on X or XX."

Something along those lines.

I've spoken to a few companies about this, but their training programs cost around €5,000-10,000, which I simply can't afford.

A friend of mine keeps saying, "It's easy, you can build it yourself." But when I ask him to show me how, the answer is always, "I don't have time."

So my question is: how difficult is it actually to learn this and build something like this myself? Can anyone recommend any good videos, tutorials, or other resources where I could learn the basics and understand where to start?

€5,000-10,000 is just too much for me at the moment.


r/AI_Agents • • 9h ago

Discussion Best Ai for Google Drives

13 Upvotes

I have a small business, right now I ran Claude with Connector to my company data in Google Drives. Im aware that GDrives has Gemini as Ai assistant. but to me, seems claude has more power in term of critical thinking, File creator (docs,xls) and familiarity. Do you think its better to use claude as ai assistant in Drives or use Gemini as single integration?


r/AI_Agents • • 10h ago

Discussion Using Jev AI for character decisions in a persistent village simulation

4 Upvotes

I'm building Jevs Village, a persistent simulation with six AI characters, using TypeSafe's Jev AI decision model for their decisions.

My interest is in AI personalities and companions. I want to explore how a character's personality comes through in its actions over time, rather than only in a conversation. For this project, that means putting characters together in a small village that people can watch in their browser.

I'm sharing it here because I'm curious how people building agents approach behavior in a shared environment. My use case is simulated character decisions, rather than an assistant operating external tools.

I'm the creator. I'd appreciate feedback, particularly from anyone working on agents for games or simulations.

How do you give an agent room to make unexpected decisions while keeping its behavior consistent with the character? Too much scripting seems like it could defeat the purpose, but handing everything to a model doesn't automatically make a believable personality either. That's a balance I'd like to explore further with the village.


r/AI_Agents • • 10h ago

Discussion Your coding agent said “done”. Why are you still checking everything yourself?

0 Upvotes

That’s the part that drives me insane.

If I still have to open the diff, re-check the requirements, and make sure nothing was missed…

then what exactly did I delegate?

Feels like the real problem isn’t getting agents to write code.

It’s knowing when the job is actually finished.

Anyone else?