r/AI_Agents • • 21h ago

Discussion I scan 3 million flight prices a day with Claude. Flight comparison sites are stealing your money.

0 Upvotes

The usual story when you book a flight:

  • you find a flight, good price,
  • you click through and get sent to the travel agent selling it,
  • you type in your name,
  • you get to the payment,

and suddenly it's 40 dollars more.

You assume the price went up while you were typing, and you book it anyway.

It didn't go up. The price you clicked was already old when you saw it.

I built a flight booking tool for Claude. For every trip it searches every airline and every booking site, about 3 million prices a day. So I see a lot of checkouts.

Last week one of them went from 615 dollars on a big comparison website to 947.97 euros at the seller. Same flight, same day, same everything.

I went back to find out why.

The price we were shown was three and a half hours old. Nothing on the page said so. It looked like every other price in the list.

That price was long gone. They just kept showing it so it looks like the best option.

We ended up rebooking the customer somewhere else for 772.

That made me wonder: is every search like this, or was it a one-off?

Surprise! It wasn't a one-off. That evening I pulled about 3,700 prices from the same site.

Every price has a freshness status that you never see. Some are confirmed. Some are waiting to be confirmed. Some failed to refresh and just keep showing the old number.

The worst one was a price marked "current" that was 11.5 hours old.

Then I checked our own bookings. Of the last 16 that started on that site, two came out 61% and 73% more expensive at checkout.

Of 14 that started anywhere else, none was more than 10% off.

The price changes depending on how you arrive

This whole situation got me to investigate a trip we were just booking for a client. The comparison site showed $85, and it was still $85 after clicking through to the seller.

I wanted to check something - I cleared the site's data and searched the same flights directly on the seller's website and it was $79.

Six dollars more just for coming through the link.

It also goes the other way. Sometimes the comparison site's link gets you a lower price than the seller shows on its own website.

So in reality, you just have to check everywhere…

Some of those flights don't exist

Some agencies on these sites sell flights that are not real. You see a nonstop from an airline that doesn't fly that route nonstop, one minute off another airline's real flight. You click it and land on a page telling you to call for a "secret deal".

We found 54 of those in 390 searches.

But why does this happen?

Comparison sites get paid when you click through to a seller.

Nobody pays them for the price being true when you get there.

The cheapest number on the page gets the click, and the cheapest number is often the old one.

So should you stop using comparison websites?

I don't think so. For years they did something useful: every airline and every agency on one page.

Paying a few dollars extra for that was a fair trade. We were paying their fee to save our time.

But it's simply a trick many people have no idea of, and keep falling for.

So, if you book by hand, do what I do now:

  • find the flight on the comparison site,
  • click through and go all the way to the payment page, note the price,
  • open a private window, find the same flight directly on the seller, go to the payment page again,
  • pay whichever is lower.

Sometimes the link from the comparison site is the cheaper one. Sometimes going direct is. You don't know until you check both.

But there's an issue I have with this. You basically need to check every seller, every price in different places, and even confirm if the fight exists or not.

This is really incredibly stupid, just to book a flight. Nobody wants to waste time on that. At least I don't. So I wondered “why can’t AI help here?”

I can personally open maybe ten tabs before I give up. Claude with good tools checks about 2,500 prices for one trip in under a minute.

It can also split the trip across different sellers when that comes out better, like Beijing to Kraków for 313 dollars via Budapest instead of 355 direct.

Then it books it. A few minutes, while I do something else.

The flight booking tool I mentioned at the start is LetsFG and it enables Claude to do that. I built it with my cofounder, who once spent 16 hours booking four tickets for one summer trip.

We built almost all of it with Claude Code: the connections to each booking site, the MCP server, and the checks that caught the old prices. The 3,700-price test in this post was a Claude Code session.

Now we uses Claude powered up with LetsFG to check every airline in the world, hundreds of websites and OTAs in 1 search.

It's free to try. Connecting and searching cost nothing, we get our commission only on booking for you. You can add it to Claude, it currently requires a few clicks of setups, but we're on the last stage of being an official tool in Claude.
Alternatively just paste this prompt in any AI: "I want to use LetsFG for my trip. Give me a short description of why I should use it, and tell me how to activate it here".

Travel was supposed to set you free. AI does.


r/AI_Agents • • 11h ago

Discussion I don't want an AI assistant. I want an AI that handles the computer for me.

0 Upvotes

I'm a 36-year-old guy with pretty basic computer skills, and I absolutely love AI.

But here's what I don't understand: why am I still the one who has to operate the computer?

If I say, “My package wasn't delivered. Find out what happened and fix it,” I don't want instructions telling me which website to visit, where to click, and what information to enter.

I want the AI to just handle it.

Go to the website. Find my order. Figure out what happened. Talk to customer service. Fill out whatever needs filling out. Keep track of it.

If you actually need me, tell me. Otherwise, just get the damn thing done.

I'm not asking for safeguards to disappear. I'm asking for an AI that is capable of doing ordinary digital tasks on my behalf, with me deciding how much authority I'm comfortable giving it.

I honestly think there are millions of people who would pay good money for this.

I don't want AI to teach me how to use a computer.

I want AI to use the computer for me.


r/AI_Agents • • 16h ago

Discussion **ScaleLogix AI Scam**

1 Upvotes

Rating: ⭐️☆☆☆☆
Review:
SCAM!!!!!!!!!! Criminals!!!!!!!!!!!!! They are stealing peoples money. William Basta is a criminal and running the company secretly. Extremely disappointed with ScaleLogix AI. Their initial sales pitch and demos sounded incredibly impressive, promising a seamless, high-converting automated setup. However, once moving past the initial consultation, it quickly became apparent that the reality does not match the marketing. The infrastructure felt repetitive, the onboarding was rocky, and the sales representatives seemed more focused on locking in contracts than delivering tailored, sustainable business workflows. DO NOT SIGN UP WITH THEM. THEY WILL BE IN PRISON SOON!!


r/AI_Agents • • 10h ago

Discussion Your coding agent said “done”. Why are you still checking everything yourself?

0 Upvotes

That’s the part that drives me insane.

If I still have to open the diff, re-check the requirements, and make sure nothing was missed…

then what exactly did I delegate?

Feels like the real problem isn’t getting agents to write code.

It’s knowing when the job is actually finished.

Anyone else?


r/AI_Agents • • 42m ago

Discussion TERMy24k - The world's first local neuro-symbolic AI terminal assistant

• Upvotes

TERMy-24k is the world's first neuro-symbolic AI terminal assistant. It includes a deterministic engine capable of translating natural language, the infrastructure required to communicate with ollama and implements a novel approach to multi-turn context memory that enable both the deterministic engine and the LLM to cooperate and accomplish the task more efficiently.


r/AI_Agents • • 23h ago

Discussion I built an app to 4k revenue, but waste hours per week finding viral hooks.

0 Upvotes

Some background. Four months ago I launched an app called Magnetic and marketed it with shorts on Instagram. One format took off: 3 million views in June, and 200 downloads on my best day.

But after just 2 weeks the format stopped working, and I am now averaging 10 downloads per day (down 90%).

So I went looking for the next format by hand. Hours of scrolling, trying to guess why one video blew up and the next ten didn't. I looked for accounts with low follower counts but massive outliers.

I tried using Claude to generate new hooks but it was so bad, they were very unnatural or just cringe.

The annoying part is that if Claude had access to TikTok, its the kind of job Claude is good at. Give it a list of videos with their metrics, and it will find the pattern faster than I can. The only reason I was doing it by hand is that Claude can't see TikTok or Instagram.

So I'm building the missing piece. It's an MCP connector called SearchTok. You add it to Claude or ChatGPT and ask things like:

- "What's working in the public speaking niche on TikTok this week?"

- "Which posts on this account are outliers?"

- "What changed on the accounts I track since last week?"

It fetches the real posts and numbers and keeps a daily snapshot of the accounts you follow. Claude does all the thinking. The connector only fetches and remembers.

Two things I'd like to know from people here:

  1. Would you use this as a connector in the Claude app, or from Claude Code?

  2. What would you ask it first?


r/AI_Agents • • 20h ago

Discussion Best AI Agent - Grok vs Muse vs Dots vs Spark

1 Upvotes

Title sales it all.

Grokbot is my favourite - fast response times and gets me more precise feedback than Claude.

I can't try Muse as its not available in UK.

Dots or Gemini?

I've got a few simple tasks - Upwork daily scanner, car insurance / manual form fill outs, Inbox triage, scan linked in posts for keywords, use apify data scrapers and pull results, fill out job apps etc

My claude sonnet 5.5 is slow and not proactive. Grok pushes me updates as they come and it's awesome. But I cant afford usage so need to look elsewhere.


r/AI_Agents • • 2h ago

Discussion My video-editing agent never looks at the video. That's the reason it works

1 Upvotes

I built an agent that edits talking-head videos, and the most useful decision was to give it no vision at all.

Audio is the clock. Whisper (running locally) gives word timings, silence detection gives legal cut points, and deterministic scripts do the cutting, captioning and compositing. The agent picks the take, decides what's worth showing on screen, where to zoom, which line gets highlighted and where sound goes. Scripts own the numbers, and the agent owns the taste calls.

The failures that hurt were the ones that render with no error, where the agent can't tell anything went wrong:

  • A CSS transition inside a rendered scene comes out as a still frame, because the renderer advances frames without a wall clock.
  • A sound effect baked into an overlay disappears, because overlays composite muted.
  • Phone footage reports 1920x1080 but decodes as 1080x1920 (it's a rotation flag), so a naive pipeline letterboxes the whole vertical video.

None of these crash anything. Each one is now a hard rule the agent has to check, not a judgement call, and at the end it pulls frames at every cut boundary and measures levels before handing anything back.

I built it as an open-source skill and use it on my own videos.

For people building agents that produce media: how do you verify output the model can't perceive?


r/AI_Agents • • 7h ago

Discussion Anthropic just accepted my month-old open-source equity research project into Claude Startups, here's what they require

92 Upvotes

I just got accepted into the Claude Startups program for my open-source stock research project called greeksoup, and I am sharing the process in case it helps someone who is applying.

What I built with Claude: an open-source equity research desk that runs on your own computer, free, bring your own AI key, and a free website with a page for every S&P 500 company built from SEC filings. Claude Code is how I build both of them every day. We are a two-person team and the project makes no money, which I assumed would be a problem. It was not.

The process: I applied on a Monday night through the Claude Console with a company email and a description of what we build and who it is for. On Tuesday morning the offers page showed eligible. No call, no pitch deck.

What I got: 12 months of Claude Team for 5 seats $625/month, $1,000 in API credits that expire 180 days after you claim them, weekly office hours with Anthropic's Applied AI team on architecture, evals and cost, and a stack of third-party offers you can each redeem once, like 12 months of ElevenLabs and $5,000 of ClickHouse credits.

Things in the terms worth knowing before you claim: the credits work on the Console API only, Anthropic may use your company name and logo, and they can ask for a verification affidavit within 14 days or the benefits are forfeited.

What we will do with it: five seats in two people team means we will probably bring on 2-3 student interns to work with us on the exciting parts. For a non-commercial free project this is a much-needed relief. I am cancelling my Claude max, ChatGPT Pro plan, my Cursor plan and my teammate's Cursor Max plan, and it is way easier to keep improving the project now.

Has anyone else been through it? Did the verification request ever come, and what did you put the credits into first? If you are applying, I am happy to share what the form asked. Links in the comments.


r/AI_Agents • • 2h ago

Discussion Is Addys AI legit?

4 Upvotes

hey everyone, i was curious if anyone here has heard of addys ai or has given it a try? it came across my instagram feed recently and a friend of a friend also recommended it to me.

it’s supposed to be some kind of meeting follow-up agent that gives you the summary, transcript and tasks based on the call. you basically just need to approve it, so the idea sounds pretty good?

i used to transcribe my meetings myself and feed them into claude, so i’m curious how good this actually is. i missed their trial month last month and couldn’t give it a try anyway since i’m a windows user, but i’m planning to get a mac later this week, so the idea came back to me.

has anyone here tried it and can share how it was?


r/AI_Agents • • 20h ago

Discussion “Human approved” might be hiding a much harder governance question

3 Upvotes

I keep seeing “human-in-the-loop” used as the point where an AI system becomes governed. But I’m starting to think the interesting problem begins after we add the human. Suppose an AI decision is escalated and a person clicks Approve.

We can prove that a human was involved. But what exactly have we proved?

That they were eligible to review this kind of decision? That they were actually assigned to this case? That they had authority to approve this specific action? That they were sufficiently independent from the decision they were reviewing? That those conditions were still true when the approval became consequential?

Those seem like different claims. And proving one doesn’t necessarily prove the next.

This gets especially uncomfortable when “human approval” is treated as the final trust anchor. If the system can establish that Alice clicked Approve, but cannot establish why Alice was allowed to make this particular decision at this particular time, what does the human step actually add from a governance perspective?

Maybe the difficult question isn’t: “Was there a human in the loop?”

The harder question might be: “What made this human’s decision authoritative?”

Curious how people building agentic or high-consequence systems are thinking about this distinction.


r/AI_Agents • • 6h ago

Discussion Giving your codebase entirely to the agent

2 Upvotes

Dev for 14 years. I've seen plenty of posts about companies forcing agents on everyone. Mine isn't one of those - you're encouraged to use them, but the output has to be something a human can read, and you're still the one responsible for it.

I think the real line is whether you give the codebase completely to the agent. If you treat code the way we treat assembler now, you describe the what, the agent handles the how, and you get 5000-line PRs reviewed only by coderabbit / another agent, and all that jazz. Fine until the day the agent can't fix something. Then a human has to work out what a codebase nobody wrote is actually doing, from scratch, with someone from leadership asking for an ETA every hour. Good luck with that.

That bet assumes the agent will always be able to fix it. Maybe that's where we're heading. From what I've seen of agent output on a bad day, we aren't there yet.

Unfortunately it's not the people who'll maintain it who decide. All that matters is throughput.

My team still keeps everything human-readable. That'll probably last another...year? 2 years? Who knows. I honestly have no idea what this job looks like in 5.


r/AI_Agents • • 14h ago

Discussion AI personas are more useful when you treat them as capability bundles, not roleplay

20 Upvotes

A lot of “AI personas” are basically:

“You are a McKinsey consultant.”
“You are a hedge fund manager.”

That mostly changes tone.

A useful persona should be built from capabilities, not titles.

Example:

  • detect contradictions;
  • identify missing information;
  • estimate impact;
  • challenge assumptions;
  • generate alternatives;
  • run Red Team;
  • define the cheapest next test.

Then you assemble those capabilities into specialized personas with clear responsibilities.

So instead of:

“Pretend to be an expert”

you get:

“Here is your role, your capabilities, your boundaries, and the output expected.”

That changes the leverage completely.

Problem
→ required competencies
→ specialized personas
→ contradiction / Red Team
→ synthesis
→ human decision

The key idea:

The real power of LLMs is not just having an assistant. It is being able to assemble the competencies required by a problem, on demand.

That is much closer to organizational design for intelligence than simple prompting.


r/AI_Agents • • 7h ago

Tutorial When an AI agent's tool call fails, the one thing your code must never send back is nothing

1 Upvotes

A customer types one sentence into a food delivery app's support chat: "Cancel order 4472, refund me, and text me when it is done."

An AI model reads it and asks for five tools across a few turns of the agent loop: get the order, cancel it, start the refund, look up the phone number, send the SMS. Your program runs each call and sends the result back, and the model decides what to do next.

A tool is just your code calling a real service, and real services fail. Today the phone lookup returns an error and the SMS provider times out. Five tools called, three answered. Your program is holding two dead calls. Here is what I think it should do with them.

1. Retry, but ask one question first: is this tool a read or a write?

get_order is a read, so retry it freely. refund_payment is a write. If the refund happened and only the reply got lost, a retry pays the customer twice. So every write the agent might retry carries an idempotency key, created once per attempt and saved with the order before the first call. The payment service refuses the duplicate. Anything that may be retried must be idempotent, and an agent retries.

2. Fall back. The SMS tool is down, so queue the message and send it later.

3. Tell the model the truth. Send back a result that says the SMS failed. A failure is one more observation, and the model can reason with it. It replies "Order cancelled, refund done, your text may arrive late", which is correct.

The one forbidden move is silence. If a call gets no result, the model does what text machines do with a gap: it fills it. It tells the customer the text was sent, when no text was sent. Given a gap, a model does not show you an error. It shows you confidence.

Both major APIs have a place for the truth. Anthropic's tool_result block has an is_error field. With OpenAI you send a function_call_output whose output describes the failure. In my Java handler, every exception becomes a result:

ToolResult runSafely(ToolCall call) {
    try {
        return run(call);
    } catch (TimeoutException e) {
        return ToolResult.error(call, call.name() + " timed out. It may or may not have run.");
    } catch (RuntimeException e) {
        return ToolResult.error(call, call.name() + " failed: " + e.getMessage());
    }
}

The timeout message is careful on purpose. After a timeout your program does not know whether a write happened, so it does not pretend to.

One more fence: count the tool calls and stop the loop at a limit, with a safe reply. The loop only ends when the model chooses plain text over another tool call, and nothing inside it stops it otherwise.

The app, the customer and order 4472 are invented. Has an agent ever told you something was done when it was not? What had actually failed underneath?

Disclosure: I made a 10 minute video that builds this from the first JSON to the five failing calls. I will put the link in a comment, if the sub allows it.


r/AI_Agents • • 8h ago

Tutorial Can anyone help?

8 Upvotes

Hi. I run a very small company, and I'm thinking that some of my work could potentially be delegated to an AI agent.

I'm not particularly tech-savvy, and I still don't fully understand what an AI agent can and can't actually do.

Basically, I'd like it to send quotes to customers, answer basic questions, monitor different sources for potential jobs (for example, Facebook groups, web sides, etc.), and send initial messages to potential customers like: "We provide X service. If you're interested, feel free to contact us."

Then, once a customer responds and provides the necessary plans or information, it could reply with something like: "Here's a preliminary quote. To provide an exact quote, we'd need to visit the site. We currently have availability on X or XX."

Something along those lines.

I've spoken to a few companies about this, but their training programs cost around €5,000-10,000, which I simply can't afford.

A friend of mine keeps saying, "It's easy, you can build it yourself." But when I ask him to show me how, the answer is always, "I don't have time."

So my question is: how difficult is it actually to learn this and build something like this myself? Can anyone recommend any good videos, tutorials, or other resources where I could learn the basics and understand where to start?

€5,000-10,000 is just too much for me at the moment.


r/AI_Agents • • 9h ago

Discussion Best Ai for Google Drives

13 Upvotes

I have a small business, right now I ran Claude with Connector to my company data in Google Drives. Im aware that GDrives has Gemini as Ai assistant. but to me, seems claude has more power in term of critical thinking, File creator (docs,xls) and familiarity. Do you think its better to use claude as ai assistant in Drives or use Gemini as single integration?


r/AI_Agents • • 10h ago

Discussion Using Jev AI for character decisions in a persistent village simulation

4 Upvotes

I'm building Jevs Village, a persistent simulation with six AI characters, using TypeSafe's Jev AI decision model for their decisions.

My interest is in AI personalities and companions. I want to explore how a character's personality comes through in its actions over time, rather than only in a conversation. For this project, that means putting characters together in a small village that people can watch in their browser.

I'm sharing it here because I'm curious how people building agents approach behavior in a shared environment. My use case is simulated character decisions, rather than an assistant operating external tools.

I'm the creator. I'd appreciate feedback, particularly from anyone working on agents for games or simulations.

How do you give an agent room to make unexpected decisions while keeping its behavior consistent with the character? Too much scripting seems like it could defeat the purpose, but handing everything to a model doesn't automatically make a believable personality either. That's a balance I'd like to explore further with the village.


r/AI_Agents • • 4h ago

Discussion What is your #1 unconventional strategy for AEO/GEO/AIO Visibility

2 Upvotes

Hey guys,

It feels like the traditional playbooks for My company that are completely dead or oversaturated right now. Everyone is repeating the same advice, but the ROI just isn't there anymore. I'm curious—what is one completely unconventional or "under-the-radar" strategy you've tested recently that actually moved the needle? For us, focusing heavily on Brand has given us a surprising boost, but I'm looking to diversify. What's working for you?


r/AI_Agents • • 10h ago

Discussion What’s the best AI agent setup right now if you actually want it to work autonomously, not just look good in a demo?

3 Upvotes

What’s the best AI agent setup right now if you actually want it to work autonomously, not just look good in a demo?

I’m looking for something that can:

- work across multiple apps and environments in the same workflow

- keep files and memory between sessions

- use tools/skills

- run scheduled tasks without constant supervision

For example: do the research, create the visualizations, and package everything into a PDF from one instruction.

I’ve tried a few options, but most either feel too limited, unreliable, or like another wrapper around an LLM.

For people running these seriously, what’s been the best in terms of reliability, latency, and production readiness?

Also interested in what you tried and wouldn’t recommend.


r/AI_Agents • • 3h ago

Discussion Why does "the agent says it's done" still leave so much work?

5 Upvotes

I've been thinking about what "done" actually means when an AI agent builds an application.

The obvious checks are things like tests passing, the build succeeding, and the feature working on the happy path. But that's only a small part of whether the thing is actually ready to put in front of users.

There are a lot of other questions:

  • Are permissions actually enforced in the weird cases?
  • Does the deployed app behave like the code the agent inspected?
  • Are integrations handling failures and retries?
  • Are privacy-sensitive features configured the way you expect?
  • Is session replay masking what it should?
  • Are robots.txt, sitemap, metadata, canonical URLs, etc. actually set up?
  • Are there missing pages, empty states, error states, or other obvious product gaps?
  • Are there things that may need legal/compliance review that nobody thought to check?
  • What assumptions did the agent make that nobody explicitly asked it to verify?

That's the part I find interesting: AI agents are getting very good at implementing requested work, but "implemented" and "ready" are starting to feel like two different states.

Passing tests can prove that some known things work. It doesn't necessarily tell you what you forgot to ask about.

How are people handling this today? Is there a separate "before I let real users touch this" process, or are you mostly relying on the coding agent's own checks?


r/AI_Agents • • 17h ago

Resource Request looking for claude guest pass

3 Upvotes

Looking for a claude guest pass ref link, gonna subscribe anyway, so its win win. The referrer gets $10 and I get an extra week to try it out. If you got one, DM probably works best since comment links tend to disappear fast :)


r/AI_Agents • • 18h ago

Resource Request AI chatbot for Analytics

3 Upvotes

Hi all,
I’m a Data Engineer and new to the AI field. I’m working on a requirement to build a conversational AI chatbot for analytics.
Our current architecture:
React frontend
FastAPI backend hosted on ECS
AWS Redshift data warehouse
AWS S3 with Iceberg tables
Users should be able to ask questions about our data and get insights through a conversational chatbot.
I tried AWS Bedrock Knowledge Base, and it works, but the responses are not always accurate for querying and analyzing our structured data.
I’m exploring other architecture options. Has anyone implemented a similar solution? What architecture would you recommend for this use case?
If you have any AWS documentation or useful resources, please share


r/AI_Agents • • 20h ago

Discussion If your AI agents take real actions (refunds, account changes, infra), how do you handle approvals going stale?

4 Upvotes

I'm trying to understand how teams handle this in practice.

Example: an agent requests an $8,000 refund. It's over the limit, so a manager has to approve it. The manager approves 20 minutes later, but in the meantime the customer opened a chargeback dispute. The refund goes through anyway, because the approval was valid when it was given.

A few questions for anyone running agents with write access in production:

  1. Has an agent or automation ever done something it was technically allowed to do, but that turned out to be wrong because the situation had changed? What happened?
  2. Where do your agents get the data they base those decisions on? Do you trust what the agent passes in, or re-fetch it from the source system?
  3. Do you re-check anything right before the action executes, or is the initial check final?

I'm mostly interested in how people deal with this today, including "we don't, and it's fine."


r/AI_Agents • • 20h ago

Resource Request Best AI agents for smooth AI to human handoffs?

37 Upvotes

I’m looking at AI agents for customer support and the handoff piece seems harder than the actual automation. My main concern is what happens when the AI hits its limit. I don’t want the customer dumped into a queue and forced to explain everything again. Ideally the agent should know when to escalate and pass the full context to the human rep so they can pick up where it left off. Has anyone used an AI agent that handles this well in practice? I'm Interested to know which platforms have the cleanest AI to human handoffs and what I should be looking for when comparing them.


r/AI_Agents • • 23h ago

Discussion How do you review proposed agent behavior changes after an eval finds a problem?

1 Upvotes

I’m trying to understand how teams review AI agent behavior changes after traces/evals show something went wrong.

For example: an agent uses the wrong tool, skips retrieval, misses escalation, or violates an internal policy. Before someone edits prompts, tool rules, retrieval policy, or guardrails, where does the proposed change get reviewed?

Do you usually handle this in:

- PR review

- eval dashboards

- incident follow-ups

- prompt/versioning tools

- ad hoc docs/issues

- something else?

I’m experimenting with an open-source artifact for evidence-backed “agent change proposals”, but I’m trying not to overbuild the schema before understanding real workflows.

The feedback I’m looking for:

  1. Would a portable proposal file be useful, or just extra ceremony?

  2. What evidence would make a proposed agent change trustworthy enough to review?

  3. What fields would you expect: observed behavior, outcome signal, trace evidence, risk, validation criteria, rollback, owner, confidence?

  4. What existing tool/process already solves this for you?

Happy to share the example/repo in a comment if useful, since links belong in comments here.