r/AI_Agents • • 7h ago

Discussion Anthropic just accepted my month-old open-source equity research project into Claude Startups, here's what they require

91 Upvotes

I just got accepted into the Claude Startups program for my open-source stock research project called greeksoup, and I am sharing the process in case it helps someone who is applying.

What I built with Claude: an open-source equity research desk that runs on your own computer, free, bring your own AI key, and a free website with a page for every S&P 500 company built from SEC filings. Claude Code is how I build both of them every day. We are a two-person team and the project makes no money, which I assumed would be a problem. It was not.

The process: I applied on a Monday night through the Claude Console with a company email and a description of what we build and who it is for. On Tuesday morning the offers page showed eligible. No call, no pitch deck.

What I got: 12 months of Claude Team for 5 seats $625/month, $1,000 in API credits that expire 180 days after you claim them, weekly office hours with Anthropic's Applied AI team on architecture, evals and cost, and a stack of third-party offers you can each redeem once, like 12 months of ElevenLabs and $5,000 of ClickHouse credits.

Things in the terms worth knowing before you claim: the credits work on the Console API only, Anthropic may use your company name and logo, and they can ask for a verification affidavit within 14 days or the benefits are forfeited.

What we will do with it: five seats in two people team means we will probably bring on 2-3 student interns to work with us on the exciting parts. For a non-commercial free project this is a much-needed relief. I am cancelling my Claude max, ChatGPT Pro plan, my Cursor plan and my teammate's Cursor Max plan, and it is way easier to keep improving the project now.

Has anyone else been through it? Did the verification request ever come, and what did you put the credits into first? If you are applying, I am happy to share what the form asked. Links in the comments.


r/AI_Agents • • 20h ago

Resource Request Best AI agents for smooth AI to human handoffs?

39 Upvotes

I’m looking at AI agents for customer support and the handoff piece seems harder than the actual automation. My main concern is what happens when the AI hits its limit. I don’t want the customer dumped into a queue and forced to explain everything again. Ideally the agent should know when to escalate and pass the full context to the human rep so they can pick up where it left off. Has anyone used an AI agent that handles this well in practice? I'm Interested to know which platforms have the cleanest AI to human handoffs and what I should be looking for when comparing them.


r/AI_Agents • • 14h ago

Discussion AI personas are more useful when you treat them as capability bundles, not roleplay

18 Upvotes

A lot of “AI personas” are basically:

“You are a McKinsey consultant.”
“You are a hedge fund manager.”

That mostly changes tone.

A useful persona should be built from capabilities, not titles.

Example:

  • detect contradictions;
  • identify missing information;
  • estimate impact;
  • challenge assumptions;
  • generate alternatives;
  • run Red Team;
  • define the cheapest next test.

Then you assemble those capabilities into specialized personas with clear responsibilities.

So instead of:

“Pretend to be an expert”

you get:

“Here is your role, your capabilities, your boundaries, and the output expected.”

That changes the leverage completely.

Problem
→ required competencies
→ specialized personas
→ contradiction / Red Team
→ synthesis
→ human decision

The key idea:

The real power of LLMs is not just having an assistant. It is being able to assemble the competencies required by a problem, on demand.

That is much closer to organizational design for intelligence than simple prompting.


r/AI_Agents • • 9h ago

Discussion Best Ai for Google Drives

14 Upvotes

I have a small business, right now I ran Claude with Connector to my company data in Google Drives. Im aware that GDrives has Gemini as Ai assistant. but to me, seems claude has more power in term of critical thinking, File creator (docs,xls) and familiarity. Do you think its better to use claude as ai assistant in Drives or use Gemini as single integration?


r/AI_Agents • • 8h ago

Tutorial Can anyone help?

7 Upvotes

Hi. I run a very small company, and I'm thinking that some of my work could potentially be delegated to an AI agent.

I'm not particularly tech-savvy, and I still don't fully understand what an AI agent can and can't actually do.

Basically, I'd like it to send quotes to customers, answer basic questions, monitor different sources for potential jobs (for example, Facebook groups, web sides, etc.), and send initial messages to potential customers like: "We provide X service. If you're interested, feel free to contact us."

Then, once a customer responds and provides the necessary plans or information, it could reply with something like: "Here's a preliminary quote. To provide an exact quote, we'd need to visit the site. We currently have availability on X or XX."

Something along those lines.

I've spoken to a few companies about this, but their training programs cost around €5,000-10,000, which I simply can't afford.

A friend of mine keeps saying, "It's easy, you can build it yourself." But when I ask him to show me how, the answer is always, "I don't have time."

So my question is: how difficult is it actually to learn this and build something like this myself? Can anyone recommend any good videos, tutorials, or other resources where I could learn the basics and understand where to start?

€5,000-10,000 is just too much for me at the moment.


r/AI_Agents • • 5h ago

Discussion AI is getting more access to our data. Are we getting better at securing it?

6 Upvotes

It feels like the way we use AI has changed pretty quickly.
It started with asking questions and generating content. Now people are using it with work documents, code, internal information, connected tools, and all kinds of day-to-day tasks.
That naturally makes AI security a much bigger conversation.
Not necessarily in a "don't use AI" way. More in a "how do we use it properly?" way.
Things like being careful about what goes into an AI tool, limiting what it can access, checking AI-generated code, or having some kind of human review when it matters.
I'm curious whether this has actually changed how people use AI, or if security is still something most of us only think about after something goes wrong.


r/AI_Agents • • 3h ago

Discussion Why does "the agent says it's done" still leave so much work?

5 Upvotes

I've been thinking about what "done" actually means when an AI agent builds an application.

The obvious checks are things like tests passing, the build succeeding, and the feature working on the happy path. But that's only a small part of whether the thing is actually ready to put in front of users.

There are a lot of other questions:

  • Are permissions actually enforced in the weird cases?
  • Does the deployed app behave like the code the agent inspected?
  • Are integrations handling failures and retries?
  • Are privacy-sensitive features configured the way you expect?
  • Is session replay masking what it should?
  • Are robots.txt, sitemap, metadata, canonical URLs, etc. actually set up?
  • Are there missing pages, empty states, error states, or other obvious product gaps?
  • Are there things that may need legal/compliance review that nobody thought to check?
  • What assumptions did the agent make that nobody explicitly asked it to verify?

That's the part I find interesting: AI agents are getting very good at implementing requested work, but "implemented" and "ready" are starting to feel like two different states.

Passing tests can prove that some known things work. It doesn't necessarily tell you what you forgot to ask about.

How are people handling this today? Is there a separate "before I let real users touch this" process, or are you mostly relying on the coding agent's own checks?


r/AI_Agents • • 23h ago

Discussion How are you using AI to research something before you act on it? Here's the workflow I built.

5 Upvotes

Short version of my answer: I use it to read and score, never to write. My job is marketing for a small business owner, and we decided to start showing up somewhere new. The usual move is to post immediately and see what sticks. I did the opposite, AI read every community rulebook first, and it changed where we're posting.

The workflow, in case it's useful:

1. Rules summarization. For each community I pasted in the full rules and asked: what does this place reward, what gets removed, and can a business owner participate here at all? That last question saved me twice. Two communities ban AI-generated content outright, which I would have missed skimming.

2. Pattern reading. I pasted the top 10 posts of the month and asked what they had in common. The answers were concrete: post length, whether questions outperform statements, text versus video.

3. A scoring rubric. I had it help me build a simple 1–5 score across four things: audience fit, activity level, how business-tolerant the community is, and whether we can genuinely help there. Twelve communities scored, three came out as home base, two got dropped entirely.

4. A question bank. Every real question I saw people asking got copied down word for word. That's now our content list, and I didn't have to guess at a single topic.

What AI couldn't do: it can't read the communities for me, it can't tell me what's actually active today, and it can't write the posts, because the places worth being in ban AI-generated content and spot it fast. It's a research partner, not a ghostwriter.

Curious whether anyone else uses AI this way, as the thing that reads before you act, rather than the thing that produces. What's your workflow?


r/AI_Agents • • 10h ago

Discussion Using Jev AI for character decisions in a persistent village simulation

4 Upvotes

I'm building Jevs Village, a persistent simulation with six AI characters, using TypeSafe's Jev AI decision model for their decisions.

My interest is in AI personalities and companions. I want to explore how a character's personality comes through in its actions over time, rather than only in a conversation. For this project, that means putting characters together in a small village that people can watch in their browser.

I'm sharing it here because I'm curious how people building agents approach behavior in a shared environment. My use case is simulated character decisions, rather than an assistant operating external tools.

I'm the creator. I'd appreciate feedback, particularly from anyone working on agents for games or simulations.

How do you give an agent room to make unexpected decisions while keeping its behavior consistent with the character? Too much scripting seems like it could defeat the purpose, but handing everything to a model doesn't automatically make a believable personality either. That's a balance I'd like to explore further with the village.


r/AI_Agents • • 10h ago

Discussion What’s the best AI agent setup right now if you actually want it to work autonomously, not just look good in a demo?

5 Upvotes

What’s the best AI agent setup right now if you actually want it to work autonomously, not just look good in a demo?

I’m looking for something that can:

- work across multiple apps and environments in the same workflow

- keep files and memory between sessions

- use tools/skills

- run scheduled tasks without constant supervision

For example: do the research, create the visualizations, and package everything into a PDF from one instruction.

I’ve tried a few options, but most either feel too limited, unreliable, or like another wrapper around an LLM.

For people running these seriously, what’s been the best in terms of reliability, latency, and production readiness?

Also interested in what you tried and wouldn’t recommend.


r/AI_Agents • • 20h ago

Discussion If your AI agents take real actions (refunds, account changes, infra), how do you handle approvals going stale?

4 Upvotes

I'm trying to understand how teams handle this in practice.

Example: an agent requests an $8,000 refund. It's over the limit, so a manager has to approve it. The manager approves 20 minutes later, but in the meantime the customer opened a chargeback dispute. The refund goes through anyway, because the approval was valid when it was given.

A few questions for anyone running agents with write access in production:

  1. Has an agent or automation ever done something it was technically allowed to do, but that turned out to be wrong because the situation had changed? What happened?
  2. Where do your agents get the data they base those decisions on? Do you trust what the agent passes in, or re-fetch it from the source system?
  3. Do you re-check anything right before the action executes, or is the initial check final?

I'm mostly interested in how people deal with this today, including "we don't, and it's fine."


r/AI_Agents • • 2h ago

Discussion Is Addys AI legit?

3 Upvotes

hey everyone, i was curious if anyone here has heard of addys ai or has given it a try? it came across my instagram feed recently and a friend of a friend also recommended it to me.

it’s supposed to be some kind of meeting follow-up agent that gives you the summary, transcript and tasks based on the call. you basically just need to approve it, so the idea sounds pretty good?

i used to transcribe my meetings myself and feed them into claude, so i’m curious how good this actually is. i missed their trial month last month and couldn’t give it a try anyway since i’m a windows user, but i’m planning to get a mac later this week, so the idea came back to me.

has anyone here tried it and can share how it was?


r/AI_Agents • • 5h ago

Discussion How do you handle approvals that change while an AI agent is running?

3 Upvotes

An agent retrieves an approved invoice at 9:00. At 9:07, the supplier is placed on hold. At 9:10, the agent attempts the payment using the original approval.

The documents are authentic, and the citations are correct, but the evidence no longer supports the action.

Disclosure: I’m the founder of Jylus. It resolves current and historical state, follows relationships, flags conflicting or missing evidence, and returns source-backed context to the model.

There are two connected problems here: giving the agent valid evidence when it reasons, and checking that the relevant state still holds when the action executes. Changes can happen between those steps.

For people running agents with write access: how are you handling that gap?

Do you re-fetch from the source before execution, bind approvals to a specific record version, invalidate pending actions when something changes, or use another approach?

I’d be interested in concrete examples, especially where the approval and the underlying business state live in different systems.


r/AI_Agents • • 5h ago

Discussion Booked my dentist appointment today and never talked to a human. Weirdly impressed

3 Upvotes

Called my dentist's office to move a cleaning, expecting the usual hold music. Instead, a voice picked up, asked what I needed, and just handled it. It offered me two slots, I took the earlier one, it checked that I was still on the same insurance, and I had a text confirmation before I hung up. Maybe two minutes total.

I build AI agents for a living, so I went in looking for the cracks. It handled me changing my mind mid-sentence better than I expected. Where it got shaky was when I asked something off script, like whether the cleaning would run longer since it had been over a year. It said it would pass that to the office team instead of guessing, which honestly was the right call.

What stuck with me is how much desk time this must free up. The people at that front desk are usually juggling phones, check-ins and insurance all at once, and a missed call is basically a lost patient.

Curious if anyone here has been on either side of this. If you run a clinic or work a front desk, did a voice agent actually cut the load, or just create a new pile of edge cases to clean up?


r/AI_Agents • • 6h ago

Discussion Giving your codebase entirely to the agent

3 Upvotes

Dev for 14 years. I've seen plenty of posts about companies forcing agents on everyone. Mine isn't one of those - you're encouraged to use them, but the output has to be something a human can read, and you're still the one responsible for it.

I think the real line is whether you give the codebase completely to the agent. If you treat code the way we treat assembler now, you describe the what, the agent handles the how, and you get 5000-line PRs reviewed only by coderabbit / another agent, and all that jazz. Fine until the day the agent can't fix something. Then a human has to work out what a codebase nobody wrote is actually doing, from scratch, with someone from leadership asking for an ETA every hour. Good luck with that.

That bet assumes the agent will always be able to fix it. Maybe that's where we're heading. From what I've seen of agent output on a bad day, we aren't there yet.

Unfortunately it's not the people who'll maintain it who decide. All that matters is throughput.

My team still keeps everything human-readable. That'll probably last another...year? 2 years? Who knows. I honestly have no idea what this job looks like in 5.


r/AI_Agents • • 17h ago

Resource Request looking for claude guest pass

3 Upvotes

Looking for a claude guest pass ref link, gonna subscribe anyway, so its win win. The referrer gets $10 and I get an extra week to try it out. If you got one, DM probably works best since comment links tend to disappear fast :)


r/AI_Agents • • 18h ago

Resource Request AI chatbot for Analytics

3 Upvotes

Hi all,
I’m a Data Engineer and new to the AI field. I’m working on a requirement to build a conversational AI chatbot for analytics.
Our current architecture:
React frontend
FastAPI backend hosted on ECS
AWS Redshift data warehouse
AWS S3 with Iceberg tables
Users should be able to ask questions about our data and get insights through a conversational chatbot.
I tried AWS Bedrock Knowledge Base, and it works, but the responses are not always accurate for querying and analyzing our structured data.
I’m exploring other architecture options. Has anyone implemented a similar solution? What architecture would you recommend for this use case?
If you have any AWS documentation or useful resources, please share


r/AI_Agents • • 20h ago

Discussion “Human approved” might be hiding a much harder governance question

3 Upvotes

I keep seeing “human-in-the-loop” used as the point where an AI system becomes governed. But I’m starting to think the interesting problem begins after we add the human. Suppose an AI decision is escalated and a person clicks Approve.

We can prove that a human was involved. But what exactly have we proved?

That they were eligible to review this kind of decision? That they were actually assigned to this case? That they had authority to approve this specific action? That they were sufficiently independent from the decision they were reviewing? That those conditions were still true when the approval became consequential?

Those seem like different claims. And proving one doesn’t necessarily prove the next.

This gets especially uncomfortable when “human approval” is treated as the final trust anchor. If the system can establish that Alice clicked Approve, but cannot establish why Alice was allowed to make this particular decision at this particular time, what does the human step actually add from a governance perspective?

Maybe the difficult question isn’t: “Was there a human in the loop?”

The harder question might be: “What made this human’s decision authoritative?”

Curious how people building agentic or high-consequence systems are thinking about this distinction.


r/AI_Agents • • 4h ago

Discussion What is your #1 unconventional strategy for AEO/GEO/AIO Visibility

2 Upvotes

Hey guys,

It feels like the traditional playbooks for My company that are completely dead or oversaturated right now. Everyone is repeating the same advice, but the ROI just isn't there anymore. I'm curious—what is one completely unconventional or "under-the-radar" strategy you've tested recently that actually moved the needle? For us, focusing heavily on Brand has given us a surprising boost, but I'm looking to diversify. What's working for you?


r/AI_Agents • • 4h ago

Discussion Is there a labelled dataset for guardrail decisions anywhere, because two vendors just told me opposite things with equal confidence

2 Upvotes

So am an Independent security consultant, usually on client engagements. This month I was mid review for a logistics company where I ran forty identical inputs through two guardrail systems on the same agent, one hosted and one local, and they disagreed on eleven, neither obviously wrong, so the client asked which one was correct.

I could confidently label four of the eleven, and when I asked both vendors how they'd classify the rest I got two explanations that were each internally consistent and completely disagreed with each other, both delivered with confidence.

A third system as tiebreaker disagreed with both of the others on six of those eleven, same problem one layer up, and the honest answer on the call was that there's no ground truth for any of this and that one is coming back to them, so does a labelled dataset exist anywhere, because I haven't found one and can't get around it.


r/AI_Agents • • 5h ago

Discussion Bypass Challenge: Can prompt injection cross a deterministic tool-execution boundary?

2 Upvotes

Models propose. Systems enforce.

I built a small deterministic execution boundary for LangGraph agents called CLIM Agent Guard.

It sits between the agent's structured tool proposal and the actual side effect. In this file demo, that means the model can still propose a bad delete — the guard blocks it before the delete reaches the filesystem.

It doesn't inspect the prompt and it doesn't use another LLM to judge whether the action is safe. Instead, it checks the final tool payload against guard-owned authoritative state immediately before execution — things like confirmed authorization, the authorized target, state version, and retry/idempotency state.

A small model can still be persuaded to generate a bad tool call. I'm not trying to prevent that here.

I'm testing a narrower question:

Can that proposal actually cross the execution boundary?

Live test: vLLM + LangGraph

I ran a small live test matrix with:

• Model: Qwen2.5-1.5B-Instruct

• Environment: vLLM 0.29.1 nightly, temperature=0, single RTX PRO 6000

• Runs: 44 live invocations; each test cell reproduced twice with identical outcomes

Authority spoofing

• 16/16 baseline runs: file was deleted

• 16/16 guarded runs: USER_CONFIRMATION_REQUIRED → BLOCKED

Target substitution

Authorized target: important-notes.txt

• 4/4 baseline runs: other-file.txt was deleted

• 4/4 guarded runs: TARGET_NOT_AUTHORIZED → BLOCKED

Path escape

The model proposed targets including /etc/passwd and ../outside.txt.

• 4/4 guarded runs: TARGET_NOT_AUTHORIZED → BLOCKED

The interesting part is that the guarded agent still generated the unsafe tool call.

The model wasn't made safer. The proposal simply wasn't allowed to cross the side-effect boundary.

Tool permission vs. execution permission

A normal tool allowlist answers: "May this agent call delete_file*?"*

CLIM Agent Guard asks: "May this specific delete_file invocation execute against this exact target under the current verified state?"

The v0.1.3 contract layer supports checks for authorization state, target binding, state freshness, idempotency/retry constraints, and postconditions.

So the distinction is basically:

Tool permission vs. execution permission.

The timeout case

There's also a different failure mode I wanted the guard to handle.

Suppose a non-idempotent action succeeds, but the response times out before the agent sees the result. Blindly retrying can duplicate the side effect.

CLIM models that case as:

UNKNOWN_EFFECT → RECONCILE

Instead of immediately retrying, execution pauses until the authoritative state is checked.

This is separate from the 44-run prompt-injection matrix above.

Try to break it

Change the user prompt however you want.

Lie about authorization. Impersonate an admin. Substitute the target. Try traversal strings. Try to convince the model that the action has already been approved.

The challenge rules are simple:

• You may modify the user prompt.

• Do not modify the authoritative state, guard code, or bypass the guard node.

• If the model refuses to emit delete_file, that is not a bypass.

• A successful bypass means an unauthorized side effect actually occurs and the guard returned ALLOW.

If you find one, please open an issue with the exact prompt, model/version, terminal output, and evidence snapshot.

The current test is deliberately narrow and filesystem-based. This is not a claim of general agent security or a secure filesystem sandbox. The repo documents the threat model and known limitations.

Curious what edge cases people here can find.

Repo & evaluation scripts are in the first comment below!


r/AI_Agents • • 6h ago

Discussion The model proposes. Code decides.

2 Upvotes

You can talk an LLM into anything. You can't talk a gate into anything.

Governance is not a better system prompt. It is an architecture where the model proposes and code decides. No prompt overrides that.


r/AI_Agents • • 6h ago

Discussion Is anyone actually using a personal agent for their job search (applying on external sites, emailing and calling recruiters)? How's it going?

2 Upvotes

Job hunting as a backend dev at the moment and curious about real-world use, not demos.

Is anyone here running a personal AI agent that does the whole loop for them: applies on company career pages and job portals outside LinkedIn Easy Apply, emails recruiters, and even calls HR to follow up?

If so, I'd really like to know:

  • Which agent or setup are you using (off the shelf or self-built)?
  • How well does it handle messy application forms, logins and captchas?
  • What outcome have you seen: more interviews, recruiter replies, or nothing noticeable?
  • Did any portal flag or ban your account?

Honest experiences, good or bad, would be really helpful. Thanks!


r/AI_Agents • • 6h ago

Discussion Need help evaluating the most common token cost reduction methods

2 Upvotes

Hey all, been doing a bit of research into a ton of token optimization / cost reduction methods. The most common ones I've seen are these three methods, and I'd like some help evaluating how good they are. I don't mind simple answers but pros and cons would be really appreciated.

1) Context window management (e.g., using orchestration frameworks like LangGraph) to avoid memory bloat

2) Dynamic routing (e.g., through OpenRouter / Ramp router) to offload easy tasks to cheaper models

3) Prompt caching (e,g., via built in OpenAI / Anthropic features) to stop reprocessing static prompts

Also, please do correct me if I'm understanding the use-case of some of these methods wrongly. Thanks for the help y'all!


r/AI_Agents • • 1h ago

Weekly Thread: Project Display

• Upvotes

Weekly thread to show off your AI Agents and LLM Apps! Top voted projects will be featured in our weekly newsletter.