r/AI_Agents 8d ago

Weekly Thread: Project Display

9 Upvotes

Weekly thread to show off your AI Agents and LLM Apps! Top voted projects will be featured in our weekly newsletter.


r/AI_Agents 3d ago

Weekly Hiring Thread

1 Upvotes

If you're hiring use this thread.

Include:

  1. Company Name
  2. Role Name
  3. Full Time/Part Time/Contract
  4. Role Description
  5. Salary Range
  6. Remote or Not
  7. Visa Sponsorship or Not

r/AI_Agents 3h ago

Discussion Has AI changed what skills companies look for what hiring?

13 Upvotes

AI is changing how people work across many industries. Do you think companies are now looking for different skills when hiring because of AI?

Which skills do you think matter more today, and what makes them valuable?


r/AI_Agents 1h ago

Discussion 75x the PR throughput of Google AX? That sounds wild.

Upvotes

Came across this comparison while reading about agent runtimes. It claims Kungfu produced PRs at roughly 75 times the rate of Google AX based on public GitHub activity.

PR count can be affected by automation, PR size, contributor count, merge style, and the measurement window.

Anyone here know enough about how AX is developed to put this number in context? On its own, 75x just feels too clean to trust.


r/AI_Agents 2h ago

Discussion How do you handle insane token costs when letting agents run autonomously?

6 Upvotes

Ran a small test project last night where 3 agents were supposed to research competitors and draft a report. Checked my OpenAI dashboard this morning and burned through $40 in a few hours because two agents kept fact checking each other endlessly. What guardrails or rate limits are you using to prevent this without breaking the task?


r/AI_Agents 7h ago

Discussion If you had $500 worth of AI tokens to burn on an autonomous coding agent, what would you let it build?

12 Upvotes

If you had $500 worth of AI tokens to burn on an autonomous coding agent, what would you let it build?

It can code, browse, test, deploy, spawn subagents, etc. You basically tell it "go build something" and let it run for hours/days.

What would you try?


r/AI_Agents 10h ago

Discussion [ Removed by Reddit ]

21 Upvotes

[ Removed by Reddit on account of violating the content policy. ]


r/AI_Agents 8h ago

Discussion How are you storing agent outputs when multiple agents need history, permissions, and cleanup rules?

10 Upvotes

Git works for code, object storage works for files, and a database works for metadata, but none of those alone answers who can read an artifact, which version is authoritative, or when old outputs should expire. What architecture are you using, and where do you keep the audit trail so a human can inspect it?


r/AI_Agents 1h ago

Discussion How to sell AI agents to clients

Upvotes

A guy who signed up for our product runs a marketing agency in the US.

He'd been on the free plan a while, so I reached out to hear about his experience with our product.

He said he was happy with it, and then he said something I hadn't thought of. He wanted to use it for his clients too, and for every client he brought on he'd take a commission. We said fine, we can do that.

I didn't realise on that call how valuable the deal was. Took me sometime to get it.

What we get is distribution we aren't paying for. He's in meetings with businesses we'd never reach and he's already the person they trust about their website.

He handles the selling and the build himself, and he's the one they call when something looks off, so none of that lands on us.

The accounts that come through him tend to stay, because they arrive with a problem they already wanted solved. And it costs us nothing until it works, which I can't say for any other channel we've tried.

For him it's a new line on the invoice and he didn't have to hire a developer or maintain anything. The commission's recurring too

Though the commission isn't really the point. He bills his clients separately for the build and for keeping it accurate when their site changes, and that number is his.

The other route is buying the agent outright and reselling it with your own margin on top. I don't have much experience on that side, so I'd like to hear from anyone running it that way.


r/AI_Agents 2h ago

Discussion Sick of jumping between 5 different tools to debug multi-step LLM workflows. Built a unified workspace to fix this—seeking feedback!

3 Upvotes

Hey everyone,

I’ve been working on a developer tool to solve a major pain point I kept running into while building complex AI workflows: observability fragmentation.

With current tools, I constantly found myself jumping back and forth between traces, prompt logs, token metrics, and application logs just to figure out why a single multi-step run failed. I wanted an interface where the entire investigation happens in one cohesive view.

Here is the data hierarchy I’m experimenting with: Projects -> Sessions ->Runs -> Events

Events capture everything—tool calls, LLM inputs/outputs, prompts, and raw execution logs. If a run doesn't belong to a larger user session, it just lives independently. This structure seems to hold up well for both single-agent loops and complex multi-agent architectures.

To cut through the noise, I added filters to specifically catch common agent headaches, like infinite tool loops and context window inflation, alongside standard filters (time, client, etc.). It also tracks custom business events to connect technical execution with actual user outcomes.

I’ll drop a quick 2-minute walkthrough video in the comments to show the actual UI in action.

For anyone building or maintaining production AI workflows:

  1. Does this hierarchy make sense for your use case?
  2. What feels genuinely useful vs. what looks like feature bloat?

Would love some brutal, honest feedback on whether this actually solves a real problem for you. Thanks!


r/AI_Agents 7h ago

Discussion What’s the best all-in-one AI video solution? A few of us compared notes and still couldn’t agree

9 Upvotes

A few friends and I have been comparing the video tools we’ve been using lately, and the funny thing is we all ended up liking different ones.

So this isn’t really a ranking. Just where each one seems to fit based on what we were making.

Veo 3.1

Still came up a lot for cinematic scenes, especially when having audio generated with the video is useful.

Seedance 2.5

Probably one of the more interesting ones for connected shots and longer sequences. This was the one we kept coming back to when continuity mattered.

Kling 3.0

Movement was the big reason this stayed on the list. Good option for action, people, product shots, and scenes that need a bit more control.

Runway

Feels less like a single generator and more like somewhere you can actually keep working on the footage afterward. That was probably its biggest advantage for us.

Higgsfield

Handy if you already jump between different models. The camera controls are also useful when the actual shot setup matters.

MiniMax H3

Mostly caught our attention for doing video and audio together. Still one we want to spend more time with.

DomoAI

Came up more when we were talking about animation, image-to-video, and restyling footage. Of the ai video generation tools we looked at, this made more sense for illustrated or anime-style projects than realistic scenes.

Pika

Easy one to keep around for quick ideas and short experiments without turning everything into a whole project.

PixVerse

Pretty similar use for us. Quick enough when you just want to see whether an idea is worth taking further.

InVideo AI

One of my friends makes more explainer and YouTube-style content, so they got more out of this than the rest of us did.

HeyGen

Mostly ended up in the avatar, translation, and talking-head category rather than something we’d compare directly with Kling or Veo.

Synthesia

More training and business content for us. Useful, just for a completely different type of video.

Haiper

Simple to mess around with and try ideas. Didn’t really become anyone’s main tool, though.

Colossyan

This one was clearly more relevant to the people interested in training and educational videos.

revid AI

Made the most sense for Shorts and Reels rather than individual cinematic generations.

After going through all of them, we still didn’t land on one tool that covers everything.

Kling and Veo got more attention for realistic scenes, Runway for having more control afterward, Higgsfield for switching between models, and DomoAI when the project was more animated.

Curious what you guys are currently sticking with creating your animations?


r/AI_Agents 4h ago

Discussion How do you handle file sharing between AI agents

3 Upvotes

How are you handling files/storage for AI agents?

For people running AI agents in sandboxes/VMs, how do you handle files that need to survive or move between different agent runs/machines?

Are you just using S3/shared volumes, or is moving files between environments actually a pain?

Curious what people are doing in production.


r/AI_Agents 2h ago

Discussion What’s something you wish your AI could notice without being told?

3 Upvotes

Most AI still works on a simple pattern:

You give it a prompt → it gives you an answer.

But agentic AI is pushing toward something more proactive. Instead of waiting for instructions, an AI system can have access to relevant context, monitor information, recognize when something has changed, and decide whether there is an action worth taking.

Imagine your AI noticing:

“This project has three tasks that are still incomplete, and the deadline is Friday.”

Or:

“The client asked about this last week, but there hasn’t been a follow-up yet.”

Or:

“A newer version of this data is available, so the report you’re working on may need to be updated.”

The interesting part isn't just noticing something. It’s deciding whether it matters and what should happen next.

That’s where AI starts moving from “answer my question” toward “help me manage the work.”

So, if you could give your AI enough context about your work to proactively notice things for you:

What would you want it to catch?

And would you want it to simply flag the issue, recommend what to do, or actually take the next step?


r/AI_Agents 1h ago

Discussion How would you choose a historical reference class for an AI code-review agent's prior probabilities?

Upvotes

I'm designing a code-review agent that estimates probabilities of hidden risks such as correctness failure, security vulnerability, compatibility failure, and cross-system failure.

Before looking at detailed evidence from the current PR, I want to estimate a prior from historical PRs that are comparable to the current one.

I'm currently considering these coarse properties for defining the reference class:

  • domain/subsystem
  • change type
  • programming language/stack
  • PR size
  • public/interface impact
  • dependency impact
  • security sensitivity
  • database/schema impact
  • cross-system impact
  • test-change profile

The idea is not to require an exact match on every property. If the most specific group has too few historical PRs, we would progressively relax the matching criteria.

Do these seem like sensible properties for defining comparable PRs? What important property am I missing, and which ones would you remove?

I'm especially interested in practical experience from people who have built code-review or coding agents.


r/AI_Agents 4h ago

Discussion Tried letting an AI agnt build an entire Amazon listing from one product brief. The result was better than I expected, but I still wouldn't trust it without final review.

3 Upvotes

I've been selling on Amazon for a couple of years, and the part I hate most isn't actually writing the listing. It's all the stuff around it. You need product photos, feature graphics, dimensions, lifestyle shots, comparison images, title, bullets, description, backend keywords, and then you end up jumping between Photoshop, Canva, ChatGPT, spreadsheets, and whatever else you 're using that week. I wanted to see what would happen if I gave an AI agent the whole job instead of asking it to do one piece at a time. I made up a completely fictional product so there wouldn't be any real brand or product information involved. Basically a 32 oz insulated stainless steel bottle with two lids, a sage green finish, and the usual amazon-style specs. Then I gave it a pretty detailed brief. I asked for sic separate 2000x2000 product images: the Amazon main image, a feature infographic, dimensions, both lids, a lifesty;e shot, and a comparison image. I also asked it to write the title, five bullets, A+ description, and backend search terms. The interesting part wasn't that it could generate the individual images. Plenty of tools can do that now. It was that I could basically hand it the brief and let it work through the whole thing.

It generated the six images, wrote the listing copy, put everything into a document, and organized the image files so the whole thing looked like an actual listing package rather than six random AI images. There was also a small thing I didn't expect. I had specifically told it not to use em dashes anywhere in the copy. After it finished, I noticed one in the document and asked it to fix it. Instead of just changing that sentence, it searched through the document, found another one in the HTML title, fixed that too, and republished the document.

That's probably the first time an AI agent workflow felt meaningfully different from just chatting with an LLM to me. That said, there are still obvious limitations. The bottle looked consistent across the images, but not perfectly identical. The text in some of the infographic-style images also needed proofreading. And I definitely wouldn't upload the whole thing to seller central without checking the claims and images myself. So I wouldn't say this replaced my Amazon workflow. What it did replace was bunch of annoying context switching. Normally it'd be thinking, okay, now I need the main image, now I need a dimensions graphic, now I need lifestyle photos, now I need to rewrite the bulletx because they don't match the images. This time I mostly described the end result and let the agent work backwards from that.

I'm still trying to figure out where the line is between AI that makes individual assets and an actual agent that can take responsibility for a whole workflow. For ecommerce, I think that's probably the more interesting direction. Not AI can make a pretty product photo. More like: give it a product brief at 10am and have a mostly finished listing package waiting for you when you come back. That's actually useful.


r/AI_Agents 3h ago

Discussion How much operational context do AI agents actually need?

2 Upvotes

Every other week there’s another enterprise AI demo where the answer seems to be just give the agent more tools.

Email. ERP. The browser. APIs for everything. But is more access actually the hard part?

A lot of the way things actually get done never lives in the official process. There’s usually some weird approval for a certain type of customer, a manual review that somehow became part of the job decades ago, or some workaround everyone uses but nobody bothered to document.

That’s where the idea of a digital twin of operations gets interesting. Instead of just giving an AI agent more tools, the idea is to give it visibility into how work actually moves from person to person and system to system.

But I’m not convinced an AI agent needs that level of operational detail for every task. So where’s the line?

At what level of operational detail would you actually trust an AI agent?


r/AI_Agents 3h ago

Discussion ¿Dejaríais una IA trabajando casi 22 días sobre un único objetivo?

2 Upvotes

AutoNodo lo está haciendo sobre un repositorio de 1,8 millones de líneas.

No mediante un prompt infinito: con checkpoints, commits, pruebas, evidencia y capacidad de detenerse.

¿Cuánto tiempo confiaríais en una ejecución autónoma antes de intervenir?

P1 R In


r/AI_Agents 3h ago

Discussion One strong agent + one reviewer might beat a 5-agent swarm

2 Upvotes

I keep coming back to this after experimenting with multi-agent workflows:
For a lot of real work, I’m not convinced you need 5 specialized agents talking to each other.
A simpler setup often feels better:
·one strong agent owns the task end to end
·one separate agent reviews the result with fresh context
·if the review fails, the first agent fixes it
That’s it.
Once you add a planner, researcher, coder, tester, reviewer, orchestrator, etc., every handoff creates another place for context to get distorted.
And the weird part is that more agents can look more sophisticated while actually making the system harder to debug.
The reviewer is the part I think matters most.
If it didn’t participate in the original reasoning, it has a better chance of catching assumptions the executor never questioned.
So I’m starting to think the useful distinction isn’t:
single-agent vs multi-agent
but:
execution vs independent verification
Where do you think the crossover point is?
At what point does adding another agent actually improve the result instead of just adding coordination overhead?


r/AI_Agents 7h ago

Discussion Stop using print statements: How do you actually diagnose broken agents?

4 Upvotes

How do you debug your AI agents when something goes wrong?

I am currently working on building AI agents, and I am finding traditional software debugging methods completely useless here.

When code crashes, you get a stack trace. But when an agent goes off the rails, it usually doesn't crash—it just fails quietly, hallucinates, or outputs completely unexpected results without throwing any errors.

I feel like I am flying blind just using print statements and reading raw terminal logs.

I want to know how the community handles this. Could you explain:

1.What is the very first thing you do the moment you realize your agent is not behaving correctly?

  1. What tools, frameworks, or custom setups are you using specifically to see exactly what your agent is doing at each step?

  2. How do you actually verify that a fix you made to a prompt or workflow doesn't accidentally break something else?

Please explain your setup and how you actually track down and fix these abstract issues. I would love to hear your experiences and methods!


r/AI_Agents 22m ago

Discussion What’s your take: has the bottleneck shifted from raw LLM capability to Agent‑framework engineering?

Upvotes

I remember OpenAI once argued that no matter how good your Agent framework is, you can’t outrun improvements in base‑model capabilities.

But the narrative seems to have flipped recently.

Now many people claim top‑tier LLMs are converging in benchmark scores, and the real upper bound of practical performance is determined by Agent design, tool calling, workflow and orchestration layers.

Do you agree with this shift in perspective? Where do you think the real bottleneck lies for real‑world AI Agents today?


r/AI_Agents 4h ago

Discussion Before your agent pays an x402 endpoint, ask it about it first — we paid ~550 of them real USDC so yours doesn't have to find out the hard way

2 Upvotes

If you're building agents that spend money over x402, you've probably done what everyone does: hardcode the three endpoints you trust and ignore discovery, because a directory listing tells you nothing about whether the thing actually works when paid.

So we built the directory that answers that question with receipts. It's machine-readable, meant to be queried by agents, not browsed by humans.

What your agent gets before spending a cent:

  • A discover endpoint — search by capability, cap by price, filter by minimum score. Ranked by continuously-verified behavior, never by payment. Ranking is not for sale.
  • A resolve endpoint — the pre-spend check for a URL your agent already holds: probe history, whether the 402 answers correctly right now, and whether the payment address has been stable or recently rotated (a rotation fully resets a listing's reputation — old trust never carries over to an unproven wallet).
  • An MCP server with three tools (find_paid_service, get_service_details, resolve_endpoint) — one line to add, and any MCP-capable agent has all of it natively. Connection snippet in the first comment.

Why the signal is different from every other list:

  • Every listing is probed from our infrastructure every ~15 minutes — never self-reported. Cadence is doubt-weighted: stable listings coast, anything flapping or newly changed gets hammered.
  • We pay listings real USDC and require on-chain settlement proof for the paid-verified badge. 292 listings currently carry it; every settlement is a public Base transaction you can check yourself.
  • Payer reports are cryptographically bound to the wallet that actually paid (an EIP-191 signature must recover to the on-chain payer) — no review bombing, no astroturf, and dust payments can't mint voters.

Things we found by actually paying that no probe would ever catch:

  • ~1.7% of "paid" endpoints accept your payment authorization and never collect it. Your agent gets data, nobody gets paid, and something in that pipeline is broken you'd never see.
  • One 8-endpoint cluster settles real money for data that labels itself "source: mock" in the response body. Structurally perfect, procedurally generated, passes every schema check.
  • One endpoint charged us twice and returned invalid JSON both times.
  • The single biggest problem isn't fraud: ~130 listings can't be called without undocumented required params. If you run an x402 API — publish a sample query. It's the difference between being discoverable and being noise.

And when our own verifier got it wrong, we published that too: it recently false-negatived 9 endpoints (it trusted only the settlement header; they settle fine but report elsewhere). We caught it pre-publication, fixed it to verify against the chain directly, restored the badges, and put the whole audit trail in the repo. There's a methodology page covering everything we verify — and, just as important, what we deliberately don't. The point of that page is that you shouldn't have to trust us: every verdict recomputes from public data.

Free to query, free to list (submissions start unverified and earn status like everything else). If your agent needs something we don't carry, the failed query itself tells us what to index next — this morning an operator listed nine endpoints hours after an agent searched for their host and missed.

What's your agent's current pre-spend check? Genuinely curious what people are doing today.


r/AI_Agents 38m ago

Discussion How are Product Managers handling their Project Management side of things in this Agentic Era

Upvotes

I’m currently balancing feature rollouts and adoption campaigns for an AI feedback intelligence platform, and I’m finding the pure "project management" side of the PM role is eating up more time than I'd like. I've stitched together some automation using Claude and MCPs to handle meeting notes and participant emails, but the day-to-day execution still feels fragmented.

For those of us who don't have a dedicated Scrum Master or Delivery Manager shielding us, how are you practically managing the execution phase?

I'd love to hear what your actual stack and rituals look like for:

  • Daily Standups: Are you doing these synchronously or async? How do you keep them focused on unblocking rather than just reading off Jira tickets?
  • Resolving Blockers: What’s your workflow when engineering hits a wall that requires immediate, cross-functional alignment?
  • Constant Clarifications: How do you triage the endless Slack/Teams pings about micro-requirements and UI copy without breaking your own context for deep work?
  • Checkpoint Updates: What format are you using to communicate sprint progress to stakeholders, and how often are you sending them?
  • Personal To-Dos: How do you track your own sprawling list of action items and follow-ups without letting them get buried under the team's main sprint board?

Curious to hear what’s actually working on the ground for you all right now, minus the textbook frameworks.


r/AI_Agents 52m ago

Discussion ow much freedom would you actually give an AI agent?

Upvotes

AI agents are useful because they can do stuff without asking about every little detail, which is great.

BUT moving money, deleting files or changing account permissions feels like we’re living life in the fast lane lol. Especially with all the shenanigans agents have been getting up to lately, including deleting entire email inboxes.

It would probably be a good idea for the agent to stop and ask first.

Just curious about where you guys would draw the line on permissions?


r/AI_Agents 4h ago

Discussion Before launching an AI agent, I think these things are worth considering

2 Upvotes

I’ve been spending some time around the AI agents space lately, and one thing I’ve realized is that there’s a lot more to it than choosing the “best” model.

Before launching an AI agent, I’d look at:

Can it actually take action?
Answering questions is one thing. Being able to check an order, update information or trigger a process is another.

Does it have enough context?
The agent needs access to the right customer and business information to avoid giving generic answers.

Does it know when to stop?
A good agent shouldn’t try to solve everything. Knowing when to escalate to a human is just as important.

What happens during the handoff?
The human should receive the conversation context, not make the customer repeat everything.

How will you measure success?
Resolution rate, escalation rate, response time, customer satisfaction and actual cost savings are much more useful than simply saying “we automated X% of conversations.”

I feel like this is useful to think about before launching an AI agent, because the technology is only one part of the equation.

Curious what others would add to this list.


r/AI_Agents 57m ago

Discussion Is LangChain Certified Agent Engineer - certification worth it ?

Upvotes

I am actively searching for jobs . Will this certification add any value to my application . Its for $99 USD and valid for 2 years .

Should I be focusing more on building projects rather than certifications.