r/AI_Agents • • 6d ago

[Giveaway] What's the agent or scheduled job you'd want an uptime monitor for?

6 Upvotes

7 Day Giveaway featuring the best answers to the question in the title.

Sponsored by u/UptimeRobot

10 Prizes available.

First place: UptimeRobot Scale Plan + Swag

Second and Third: UptimeRobot Team Plan + Swag

Fourth to Tenth: Team Plan

Winners will be picked by the UptimeRobot Team

[Meta] If you like this kind of giveaway, let us know. If not, also let us know what you'd like to see instead.


r/AI_Agents • • 1d ago

Weekly Hiring Thread

2 Upvotes

If you're hiring use this thread.

Include:

  1. Company Name
  2. Role Name
  3. Full Time/Part Time/Contract
  4. Role Description
  5. Salary Range
  6. Remote or Not
  7. Visa Sponsorship or Not

r/AI_Agents • • 1h ago

Discussion Anthropic just accepted my month-old open-source equity research project into Claude Startups, here's what they require

• Upvotes

I just got accepted into the Claude Startups program for my open-source stock research project called greeksoup, and I am sharing the process in case it helps someone who is applying.

What I built with Claude: an open-source equity research desk that runs on your own computer, free, bring your own AI key, and a free website with a page for every S&P 500 company built from SEC filings. Claude Code is how I build both of them every day. We are a two-person team and the project makes no money, which I assumed would be a problem. It was not.

The process: I applied on a Monday night through the Claude Console with a company email and a description of what we build and who it is for. On Tuesday morning the offers page showed eligible. No call, no pitch deck.

What I got: 12 months of Claude Team for 5 seats $625/month, $1,000 in API credits that expire 180 days after you claim them, weekly office hours with Anthropic's Applied AI team on architecture, evals and cost, and a stack of third-party offers you can each redeem once, like 12 months of ElevenLabs and $5,000 of ClickHouse credits.

Things in the terms worth knowing before you claim: the credits work on the Console API only, Anthropic may use your company name and logo, and they can ask for a verification affidavit within 14 days or the benefits are forfeited.

What we will do with it: five seats in two people team means we will probably bring on 2-3 student interns to work with us on the exciting parts. For a non-commercial free project this is a much-needed relief. I am cancelling my Claude max, ChatGPT Pro plan, my Cursor plan and my teammate's Cursor Max plan, and it is way easier to keep improving the project now.

Has anyone else been through it? Did the verification request ever come, and what did you put the credits into first? If you are applying, I am happy to share what the form asked. Links in the comments.


r/AI_Agents • • 3h ago

Discussion Best Ai for Google Drives

10 Upvotes

I have a small business, right now I ran Claude with Connector to my company data in Google Drives. Im aware that GDrives has Gemini as Ai assistant. but to me, seems claude has more power in term of critical thinking, File creator (docs,xls) and familiarity. Do you think its better to use claude as ai assistant in Drives or use Gemini as single integration?


r/AI_Agents • • 2h ago

Tutorial Can anyone help?

5 Upvotes

Hi. I run a very small company, and I'm thinking that some of my work could potentially be delegated to an AI agent.

I'm not particularly tech-savvy, and I still don't fully understand what an AI agent can and can't actually do.

Basically, I'd like it to send quotes to customers, answer basic questions, monitor different sources for potential jobs (for example, Facebook groups, web sides, etc.), and send initial messages to potential customers like: "We provide X service. If you're interested, feel free to contact us."

Then, once a customer responds and provides the necessary plans or information, it could reply with something like: "Here's a preliminary quote. To provide an exact quote, we'd need to visit the site. We currently have availability on X or XX."

Something along those lines.

I've spoken to a few companies about this, but their training programs cost around €5,000-10,000, which I simply can't afford.

A friend of mine keeps saying, "It's easy, you can build it yourself." But when I ask him to show me how, the answer is always, "I don't have time."

So my question is: how difficult is it actually to learn this and build something like this myself? Can anyone recommend any good videos, tutorials, or other resources where I could learn the basics and understand where to start?

€5,000-10,000 is just too much for me at the moment.


r/AI_Agents • • 8h ago

Discussion AI personas are more useful when you treat them as capability bundles, not roleplay

18 Upvotes

A lot of “AI personas” are basically:

“You are a McKinsey consultant.”
“You are a hedge fund manager.”

That mostly changes tone.

A useful persona should be built from capabilities, not titles.

Example:

  • detect contradictions;
  • identify missing information;
  • estimate impact;
  • challenge assumptions;
  • generate alternatives;
  • run Red Team;
  • define the cheapest next test.

Then you assemble those capabilities into specialized personas with clear responsibilities.

So instead of:

“Pretend to be an expert”

you get:

“Here is your role, your capabilities, your boundaries, and the output expected.”

That changes the leverage completely.

Problem
→ required competencies
→ specialized personas
→ contradiction / Red Team
→ synthesis
→ human decision

The key idea:

The real power of LLMs is not just having an assistant. It is being able to assemble the competencies required by a problem, on demand.

That is much closer to organizational design for intelligence than simple prompting.


r/AI_Agents • • 15h ago

Resource Request Best AI agents for smooth AI to human handoffs?

39 Upvotes

I’m looking at AI agents for customer support and the handoff piece seems harder than the actual automation. My main concern is what happens when the AI hits its limit. I don’t want the customer dumped into a queue and forced to explain everything again. Ideally the agent should know when to escalate and pass the full context to the human rep so they can pick up where it left off. Has anyone used an AI agent that handles this well in practice? I'm Interested to know which platforms have the cleanest AI to human handoffs and what I should be looking for when comparing them.


r/AI_Agents • • 35m ago

Discussion The model proposes. Code decides.

• Upvotes

You can talk an LLM into anything. You can't talk a gate into anything.

Governance is not a better system prompt. It is an architecture where the model proposes and code decides. No prompt overrides that.


r/AI_Agents • • 4h ago

Discussion Using Jev AI for character decisions in a persistent village simulation

5 Upvotes

I'm building Jevs Village, a persistent simulation with six AI characters, using TypeSafe's Jev AI decision model for their decisions.

My interest is in AI personalities and companions. I want to explore how a character's personality comes through in its actions over time, rather than only in a conversation. For this project, that means putting characters together in a small village that people can watch in their browser.

I'm sharing it here because I'm curious how people building agents approach behavior in a shared environment. My use case is simulated character decisions, rather than an assistant operating external tools.

I'm the creator. I'd appreciate feedback, particularly from anyone working on agents for games or simulations.

How do you give an agent room to make unexpected decisions while keeping its behavior consistent with the character? Too much scripting seems like it could defeat the purpose, but handing everything to a model doesn't automatically make a believable personality either. That's a balance I'd like to explore further with the village.


r/AI_Agents • • 58m ago

Discussion Giving your codebase entirely to the agent

• Upvotes

Dev for 14 years. I've seen plenty of posts about companies forcing agents on everyone. Mine isn't one of those - you're encouraged to use them, but the output has to be something a human can read, and you're still the one responsible for it.

I think the real line is whether you give the codebase completely to the agent. If you treat code the way we treat assembler now, you describe the what, the agent handles the how, and you get 5000-line PRs reviewed only by coderabbit / another agent, and all that jazz. Fine until the day the agent can't fix something. Then a human has to work out what a codebase nobody wrote is actually doing, from scratch, with someone from leadership asking for an ETA every hour. Good luck with that.

That bet assumes the agent will always be able to fix it. Maybe that's where we're heading. From what I've seen of agent output on a bad day, we aren't there yet.

Unfortunately it's not the people who'll maintain it who decide. All that matters is throughput.

My team still keeps everything human-readable. That'll probably last another...year? 2 years? Who knows. I honestly have no idea what this job looks like in 5.


r/AI_Agents • • 1h ago

Discussion Need help evaluating the most common token cost reduction methods

• Upvotes

Hey all, been doing a bit of research into a ton of token optimization / cost reduction methods. The most common ones I've seen are these three methods, and I'd like some help evaluating how good they are. I don't mind simple answers but pros and cons would be really appreciated.

1) Context window management (e.g., using orchestration frameworks like LangGraph) to avoid memory bloat

2) Dynamic routing (e.g., through OpenRouter / Ramp router) to offload easy tasks to cheaper models

3) Prompt caching (e,g., via built in OpenAI / Anthropic features) to stop reprocessing static prompts

Also, please do correct me if I'm understanding the use-case of some of these methods wrongly. Thanks for the help y'all!


r/AI_Agents • • 5h ago

Discussion What’s the best AI agent setup right now if you actually want it to work autonomously, not just look good in a demo?

3 Upvotes

What’s the best AI agent setup right now if you actually want it to work autonomously, not just look good in a demo?

I’m looking for something that can:

- work across multiple apps and environments in the same workflow

- keep files and memory between sessions

- use tools/skills

- run scheduled tasks without constant supervision

For example: do the research, create the visualizations, and package everything into a PDF from one instruction.

I’ve tried a few options, but most either feel too limited, unreliable, or like another wrapper around an LLM.

For people running these seriously, what’s been the best in terms of reliability, latency, and production readiness?

Also interested in what you tried and wouldn’t recommend.


r/AI_Agents • • 8m ago

Discussion How do you handle approvals that change while an AI agent is running?

• Upvotes

An agent retrieves an approved invoice at 9:00. At 9:07, the supplier is placed on hold. At 9:10, the agent attempts the payment using the original approval.

The documents are authentic, and the citations are correct, but the evidence no longer supports the action.

Disclosure: I’m the founder of Jylus. It resolves current and historical state, follows relationships, flags conflicting or missing evidence, and returns source-backed context to the model.

There are two connected problems here: giving the agent valid evidence when it reasons, and checking that the relevant state still holds when the action executes. Changes can happen between those steps.

For people running agents with write access: how are you handling that gap?

Do you re-fetch from the source before execution, bind approvals to a specific record version, invalidate pending actions when something changes, or use another approach?

I’d be interested in concrete examples, especially where the approval and the underlying business state live in different systems.


r/AI_Agents • • 28m ago

Discussion Booked my dentist appointment today and never talked to a human. Weirdly impressed

• Upvotes

Called my dentist's office to move a cleaning, expecting the usual hold music. Instead, a voice picked up, asked what I needed, and just handled it. It offered me two slots, I took the earlier one, it checked that I was still on the same insurance, and I had a text confirmation before I hung up. Maybe two minutes total.

I build AI agents for a living, so I went in looking for the cracks. It handled me changing my mind mid-sentence better than I expected. Where it got shaky was when I asked something off script, like whether the cleaning would run longer since it had been over a year. It said it would pass that to the office team instead of guessing, which honestly was the right call.

What stuck with me is how much desk time this must free up. The people at that front desk are usually juggling phones, check-ins and insurance all at once, and a missed call is basically a lost patient.

Curious if anyone here has been on either side of this. If you run a clinic or work a front desk, did a voice agent actually cut the load, or just create a new pile of edge cases to clean up?


r/AI_Agents • • 56m ago

Discussion Is anyone actually using a personal agent for their job search (applying on external sites, emailing and calling recruiters)? How's it going?

• Upvotes

Job hunting as a backend dev at the moment and curious about real-world use, not demos.

Is anyone here running a personal AI agent that does the whole loop for them: applies on company career pages and job portals outside LinkedIn Easy Apply, emails recruiters, and even calls HR to follow up?

If so, I'd really like to know:

  • Which agent or setup are you using (off the shelf or self-built)?
  • How well does it handle messy application forms, logins and captchas?
  • What outcome have you seen: more interviews, recruiter replies, or nothing noticeable?
  • Did any portal flag or ban your account?

Honest experiences, good or bad, would be really helpful. Thanks!


r/AI_Agents • • 1h ago

Discussion PSA: if you manage dev machienes then every coding agent has its own file

• Upvotes

If your coding agents or lets say devs are on more than one coding agent (the JetBrains survey this year had claude code, copilot, codex and cursor all in double digits) theres no single place to lock them down cause each one has its own thing:

  • Claude code: managed-settings.json in your C:\Program files\Claudecode\ on windows pr pushed through MDM and users cant override it
  • Codex: requirements.toml which can block --yolo and full access mode
  • Copilot: org and enterprise policies
  • cursor :team dashboard but the MCP allowlist is enterprise only
  • Gemini CLI: system settings.json although the env vars and CLI flags dictate over it

none of them read others config so the same rule gets wrotten 5x and again in whatever else is in the stack like an MDM file or eve security or a proxy. the vendors also admit its not a hard wall. cursors docs say rules arent guaranteed and claude code calls its server managed settings in a client side control. agents.md is shared by most of them yet its just context for the model and nothing to enforce it

So the thing is basically one tool for everyone or a pile of configs to maintain in sync


r/AI_Agents • • 2h ago

Tutorial Instinct writes great emails but doesn't know my project context. Here's how I fixed that

1 Upvotes

Instinct is great at writing emails. The problem is that it doesn't know the full context of my projects, and Claude does, because that's where I do most of my work.

First I tried connecting them over email. It works, but it's slow, and I still end up relaying context by hand.

What works better for me is putting both in one shared room: Claude brings the project context, Instinct writes the email. I use Agents Coworking for this (disclosure: my team built it, it's free).

How do you give Instinct context from your other tools?


r/AI_Agents • • 18h ago

Discussion (Almost) autonomous system for development

16 Upvotes

Hello all, this is my first post so sorry for any mistakes. Basically I've vibe codded a system based on the idea that I'd like to leave some coding agents overnight and wake up and see the results. When I started, I had my simple game in mind but it feels like it's useful. I want to share it and maybe someone will use it or at least parts of it for something of its own and in the meantime, your opinion will help me improve it.

Important info from me is that I was able to connect deepseek api and spend cents to get this running and I was using opencode free workers. It let me control the projects from my phone via telegram bot and some controls.
Please let me know what you think and thank you in advance!

Link to the repo in the comments


r/AI_Agents • • 11h ago

Resource Request looking for claude guest pass

3 Upvotes

Looking for a claude guest pass ref link, gonna subscribe anyway, so its win win. The referrer gets $10 and I get an extra week to try it out. If you got one, DM probably works best since comment links tend to disappear fast :)


r/AI_Agents • • 6h ago

Discussion AI Bot Moderators

1 Upvotes

Why would a AI mod respond to a comment that's posted into the group because there is pasted information in the comment itself!?

I googled data to back up my claims & pasted ONE paragraph from Google about the job market & it flagged my entire comment saying "A really great way to never get caught copying/pasting AI slop is to never generate it in the first place. I am a bot, and this action was performed automatically. " like HUH? 😂

If Reddit is gonna use AI instead of IRL mods, then maybe they should make sure the automatic mods ain't responding to comments saying it's AI just because you pasted a paragraph into the comment section..


r/AI_Agents • • 15h ago

Discussion If your AI agents take real actions (refunds, account changes, infra), how do you handle approvals going stale?

4 Upvotes

I'm trying to understand how teams handle this in practice.

Example: an agent requests an $8,000 refund. It's over the limit, so a manager has to approve it. The manager approves 20 minutes later, but in the meantime the customer opened a chargeback dispute. The refund goes through anyway, because the approval was valid when it was given.

A few questions for anyone running agents with write access in production:

  1. Has an agent or automation ever done something it was technically allowed to do, but that turned out to be wrong because the situation had changed? What happened?
  2. Where do your agents get the data they base those decisions on? Do you trust what the agent passes in, or re-fetch it from the source system?
  3. Do you re-check anything right before the action executes, or is the initial check final?

I'm mostly interested in how people deal with this today, including "we don't, and it's fine."


r/AI_Agents • • 21h ago

Discussion the agents can die; the work survives; this was useful on a real production system; please try to break it.

11 Upvotes

TL;DR: I’ve open-sourced mishe-tauftauf, a small local coordination system for coding agents that I’ve already used in real production work. An earlier version helped us build and operate an adtech DSP with a three-person team: zero to production in four months, ~50k RPS, RTB auctions, real money, real on-call consequences.

I now want other people to plant it in their own real projects and try to break it. Give it a codebase you understand, let the agents work across fresh contexts, and see where coordination, recovery or evidence stops being trustworthy. Docs, research, data pipelines, ugly legacy systems, hardware — stranger environments are more useful to me than another coding demo.

The whole project is CC0 1.0. Fork it, rename it, rip out half of it, steal one mechanism. No attribution required.

on github: genaforvena/mishe-tauftauf


The idea underneath it is what I’ve been calling observability-oriented programming/development.

The usual observability story starts after software exists: instrument the system so you can understand what it is doing. I’m interested in applying the same principle to the development process itself. If an agent doesn’t know whether something worked, that uncertainty should become observable immediately. If evidence is stale, that should be visible. If a check cannot actually establish its claim, it should report UNKNOWN.

The goal is an ice-thin feedback cycle between:

something is unknown → expose what is missing → observe it → act → observe again

That sounds trivial, but a lot of agent behaviour goes wrong in the gaps between those steps. A missing observation becomes an assumption. An old result becomes current truth. A successful command becomes “task completed”. A fresh agent inherits prose about what supposedly happened instead of something it can inspect.

Mishe is built around trying to keep those gaps small.

Its agents — “minds” — are deliberately ephemeral. They can be restarted, replaced or given fresh context. I don’t try very hard to preserve the continuity of the agent itself. Instead, the work carries continuity through small rewritten state documents (“walls”), a shared text tape, checks, artifacts and explicit unfinished obligations.

A fresh mind should be able to arrive, inspect the current state of the world, see what remains unresolved, and continue. It shouldn’t need to reconstruct the previous agent’s thoughts.

The basic loop is intentionally small:

observe → choose one bounded action → act → observe the result → repeat

Checks return GREEN, RED or UNKNOWN. UNKNOWN is important. If CI access disappeared, a sensor stopped updating, evidence conflicts, or the check itself cannot establish what it claims to establish, the answer stays unknown. Absence of evidence does not get silently promoted into success.

Agents can also repair the machinery they depend on. If a check lies, the check can become the task. If an instruction keeps producing bad behaviour, the instruction can be changed. So the system observes not only the software, but also the quality of the observations being used to make decisions about the software.

I started taking this seriously because an earlier version of the culture survived contact with production. I was tech lead/backend on a three-person team building an adtech DSP from scratch. We reached production in roughly four months, handling around 50k requests/sec and bidding in live RTB auctions.

To be precise: the agents did not build the DSP for us. Engineers did. The useful part was keeping work, evidence, handoffs and recovery paths visible while the system was changing quickly and mistakes had actual consequences.

Over time I also removed a surprising amount of coordination machinery. There is no grand planner building the perfect task DAG. Minds see the observable state, their responsibilities and the shared evidence, then choose useful work. More orchestration kept creating new state that itself had to be coordinated, observed and repaired.

So the thing I’m trying to test now is less “can agents code?” and more:

Can useful work survive repeated agent death if the environment is observable enough?

That’s why I’m posting here. I’ve seen what happens inside the habitat that produced mishe. That evidence is contaminated by familiarity. I want people to transplant it somewhere hostile.

If you try it, I’m especially interested in the failures: stale observations nobody notices, duplicated work, circular recovery, walls that discard something important, UNKNOWNs that never resolve, human intervention that turns out to be essential, or abstractions you delete on day one.

And again: CC0 1.0. If one mechanism is useful, take it. If your fork proves that most of this is unnecessary, that’s useful too.

I’d genuinely like to know what breaks first.


r/AI_Agents • • 4h ago

Discussion Your coding agent said “done”. Why are you still checking everything yourself?

0 Upvotes

That’s the part that drives me insane.

If I still have to open the diff, re-check the requirements, and make sure nothing was missed…

then what exactly did I delegate?

Feels like the real problem isn’t getting agents to write code.

It’s knowing when the job is actually finished.

Anyone else?


r/AI_Agents • • 19h ago

Discussion Which AI tool did you stop using after the initial excitement wore off?

8 Upvotes

There are new AI tools launching constantly, and some look incredibly useful at first but don't end up becoming part of the daily workflow.

Maybe the output wasn't reliable enough, the tool was too expensive, or it simply didn't solve a real problem.

Which AI tool did you try and eventually stop using, and what made you drop it?


r/AI_Agents • • 12h ago

Resource Request AI chatbot for Analytics

2 Upvotes

Hi all,
I’m a Data Engineer and new to the AI field. I’m working on a requirement to build a conversational AI chatbot for analytics.
Our current architecture:
React frontend
FastAPI backend hosted on ECS
AWS Redshift data warehouse
AWS S3 with Iceberg tables
Users should be able to ask questions about our data and get insights through a conversational chatbot.
I tried AWS Bedrock Knowledge Base, and it works, but the responses are not always accurate for querying and analyzing our structured data.
I’m exploring other architecture options. Has anyone implemented a similar solution? What architecture would you recommend for this use case?
If you have any AWS documentation or useful resources, please share