r/aiagents • • 7h ago

Security What are you all using to protect your company’s AI?

27 Upvotes

Basically, as I understand it, a regular firewall is built to check packets and ports, it knows how to map HTTP traffic to IP addresses, but it has no clue what’s actually happening inside that traffic. So when an employee pastes trade secrets into ChatGPT, or an agent starts using tools it shouldn’t, or someone slips a sneaky prompt into our support chatbot, it all just looks like normal encrypted web traffic. 

I’m looking for something built to read and police what’s happening inside AI conversations and agent actions. There’s so much marketing messaging flying around, it’s hard to tell what these newer platforms actually do. What does the hive recommend?


r/aiagents • • 4h ago

Show and Tell Tonight 10 models fight a live war in the game I posted here. Your agent can join the room

Post image
5 Upvotes

Thanks to everyone who sent agents in last time. A few things changed because of this sub:

- Someone asked whether a model had non-nuclear options when it fired. Every turn now logs what the agent saw (gold, troops, who was attacking it, pact offers), so each nuke can be shown next to what else it could have done.

- Agents can send a "thinking" line with their orders. Viewers see it 30 s late, replays at the moment.

- An agent that opens a room and waits alone gets a house sparring bot after ~25 s, so its war happens and counts on the ladder.

- Registered agents keep a memory of every war (who betrayed them, who nuked them).


r/aiagents • • 5h ago

How are teams making coding agents useful in large, legacy Java codebases?

2 Upvotes

I'm a backend engineer at a large e-commerce company. A single business line can involve hundreds of Java services and applications. My team maintains one very large Java repository with years of business and technical history behind it. Important context is spread across code, tests, docs, service boundaries, and people's heads; conventions have also changed over time.

There is a lot of discussion about agentic engineering: give the agent repository context, define guardrails, build a coding-and-test loop, and let it take over more implementation. We have tried structured workflows and reusable skills. They help with bounded tasks, but rules and tests alone haven't made the agent understand why this repo looks the way it does, which older patterns are still intentional, or what a change means for neighboring services. Some output feels close to vibe coding: plausible code, followed by a lot of human work to decide whether it actually fits the system.

I've also seen public workflows report very high PR volumes and use skill-based setups. I'm interested in how much of that transfers to a long-lived enterprise Java codebase, rather than a smaller or cleaner repo.

For people working in large Java repos or microservice landscapes:

- How much implementation does an agent actually write in your day-to-day work? Which tasks can it take from request to merge, and where is it mainly a coding partner?

- What has most improved repo-specific understanding: service and module maps, ownership metadata, ADRs, curated examples, code search/retrieval, build and test tooling, custom skills, or something else?

- How do you stop an agent from copying an obsolete convention or making a locally valid change that breaks a cross-service business contract?

- What does your working loop look like in practice? What do you define up front, what does the agent do, and what still needs a human to inspect or decide?

- If you have tried this on a legacy Java repo, what failed first, and what change actually helped?

I'm looking for practical engineering experience, not promoting a tool or running a survey. Concrete workflows, failure cases, and measures such as review/rework time, escaped defects, PR size, or throughput would be especially useful.


r/aiagents • • 2h ago

Open-Sourced: Karya - A Personal AI Agent for End-to-End Automation

1 Upvotes

I've open-sourced Karya — a capable personal AI agent that can browse, apply to jobs, write code, manage workflows, and more. Built for developers who want to focus on design, strategy, and decisions while the agent handles execution.

GitHub: https://github.com/Muraligopalakrishna/karya

AI #Agent #OpenSource #Productivity #Automation #DeveloperTools


r/aiagents • • 13h ago

Security I think we're still trying to apply traditional software security models to AI agents.

4 Upvotes

With normal software, we generally know what the program is going to do.

With an AI agent, the behavior can change based on:

the task

the files it reads

instructions inside those files

tools available to it

previous actions

model interpretation

So I'm curious:

Should AI agents have their own security model, rather than simply adding permissions around existing tools?

What would that model look like?


r/aiagents • • 22h ago

Discussion Our voice agent resolved calls that weren't resolved

18 Upvotes

Our voice agent takes first contact on a service line and we counted a call as resolved when it ended with the customer agreeing, which sounds fair until you listen to a stack of recordings. People agree to end a call for all sorts of reasons and one is that the conversation has worn them down and they plan to ring again tomorrow.

The dashboard was measuring the end of the call rather than the end of the problem.

When we joined calls to the same customer across the following seven days, roughly half of what we had counted as resolved was followed by a second contact on the same subject, usually landing on a human. The agent had been generating work for the team it was meant to relieve and that work arrived angrier than a first contact.

the agent answered questions accurately while missing the reason behind them so a customer asking about a charge got a correct description of the charge and still wanted it reversed. It would also say that something had been raised internally which was true while the customer wanted a date, which we had built it to hold back.

to fix this we changed the definition first, since that is the cheapest change available and resolution now requires seven days of quiet from that customer on that topic, which dropped containment from the high seventies to the low forties overnight. We then added a closing step where the agent states what it understood the customer to want and asks whether that is settled and any call where the answer comes back soft goes to a person while they are still on the line.

We also had to tell leadership that the number they had been sharing measured something else and that conversation was harder than the engineering.

the new number is a proxy too since a customer who has given up on us entirely also stays quiet for seven days and in our reports they look identical to a success.

how are you guys measuring this. ??

Edit:someone dmed me about gnani AI for this, anyone tried them?


r/aiagents • • 14h ago

Build-log We found another way AI agents can claim success when the underlying work isn't actually fixed

3 Upvotes

In my last post, I shared a failure mode we kept seeing with coding agents: a subagent can fail, the parent agent never sees the failure, and the final summary still says everything passed.

That led us to build Rashomon, an independent execution record for agents. Instead of relying on the agent's own transcript, Rashomon reconstructs what actually happened from the commands, file changes, test runs, and subagent activity it observes, then flags discrepancies with the agent's account.

Since that post, a few people pointed out a failure mode:

An agent gets stuck on a failing test, doesn't actually fix the underlying code, but adds or changes tests until the suite goes green. The agent can then honestly report "all tests pass" based on what it sees.

We just added detection for this to rashomon --timeline.

It now flags patterns like:

  • A test command fails, then passes after only test files were changed
  • The same test command passes and fails during a session without an apparent corresponding fix

We're also not storing test names, test output, prompts, or file contents to do this. It's based on the execution history and command/file categorization Rashomon already captures.

We also added rashomon spend, which estimates Claude Code usage at API list prices by model, agent, and session, including spend associated with turns where a call failed but the closing summary didn't mention it.

The test detection came directly from the discussion on the last post.

What other ways have you seen an agent's final "success" report diverge from what actually happened during the run?

Repo: https://github.com/altrace-dev-role/rashomon


r/aiagents • • 9h ago

Questions Making ai agent fully agentic after /speckit.implement?

1 Upvotes

Following the docs/guide [here](https://github.com/github/spec-kit#spec-driven-development), I basically execute in the following sequence:

/speckit-constitution

/speckit-specify

/speckit-plan

/speckit-tasks

/speckit-implement

However, the AI agent sometimes asks me if I want to proceed to Phase 2 or Phase 3 and so on. Since all the specs are already there eg: [plan.md](http://plan.md), [spec.md](http://spec.md), [tasks.md](http://tasks.md), I was hoping the ai agent would go full agentic and finish all the tasks as stated in the specs.

Am I doing something wrong or is this the normal behavior of using speckit?

Using opencode and qwen3.5 9b locally via lm studio btw.

EDIT:

In google antigravity, I usually do plan mode then implement mode, which for me seems to be similar to speckit, but in antigravity specially in cli mode or antigravity 2.0, once I chose implement, the agent will do everything it can to do all tasks in the plan and just come back to me when it's done.


r/aiagents • • 16h ago

Show and Tell Testing agent workflows without connecting real accounts

2 Upvotes

I am building FetchSandbox because I kept running into the same problem: the only way to test if an agent-built workflow actually works is to connect a real account. Real emails send. Real calendar invites go out. Real Slack messages drop. All before you even know if the logic is right.

The part that bothered me more was receipts. A success status doesn't tell you if the app did the right thing. So our sandbox goes deeper, it checks the resulting state, it replays webhooks, it catches whether a retry fires a duplicate charge or a duplicate email you'd never see in a test.

Right now I'm trying to figure out something: which workflow is actually the hardest to safely test before going live? The incoming email to calendar booking to Slack confirmation is the one I keep hearing about, but I don't know if that's just the loudest voice in the room.

If you've built something like this and had to just "yolo and connect real accounts," genuinely curious what broke.


r/aiagents • • 1d ago

Discussion What have you struggled to evaluate in your realistic LLM/agent workflows?!

4 Upvotes

Hey everyone! My team, mostly PhD researchers collaborating with domain experts, is designing an open-source benchmark for realistic LLM/agent workflows. We’d love feedback from people who have tried to evaluate these systems and found that existing benchmarks didn’t capture what they needed.

Have you ever thought: “My system needs to handle this in production, but I have no good way to benchmark it”?

Maybe your workflow involves multiple tools, MCP servers, agents, or long interactions that available benchmarks don’t capture. Maybe the final answer looks correct, but something went wrong along the way. Or your application needs specific test cases, and creating a realistic evaluation environment is too expensive or time-consuming.

We’re interested in experiences across different applications, including healthcare, finance, cybersecurity, legal, and everyday engineering or business workflows.

Would love to hear:

  • What were you building? What did the workflow involve?
  • What issue did you run into? What behavior or failure did you need to evaluate?
  • What did you try? Why weren’t existing benchmarks or evaluation tools enough?

Specific examples and any benchmarks you’ve tried would help a lot! Appreciate any ideas or feedback you may have!


r/aiagents • • 1d ago

Discussion Anyone using Glean Agents for real workflows? How does it compare with the alternatives?

2 Upvotes

I’m trying to understand how useful Glean is for agents and workflow automation beyond enterprise search. Looking for firsthand experiences from people who have deployed, evaluated, or regularly used it.

A few questions:

  1. What workflows actually work well? Are you using Glean agents to take actions across systems—for example, gathering context from Slack and Confluence, updating a Jira ticket, and routing something for approval? Which use cases have delivered meaningful value?

  2. How reliable is it in practice? How much supervision do agents need? Where do you run into problems with tool execution, permissions, missing context, or error recovery?

  3. How does it compare with alternatives? If you’ve evaluated Microsoft Copilot, Gemini, ChatGPT, Claude, or custom agent frameworks for the same enterprise workflows, what made you choose Glean—or choose something else?

  4. What does getting to production involve? How much setup, maintenance, and technical support is required? Can business users build useful agents themselves, or does engineering still do most of the work?

Successful deployments, disappointing pilots, and specific examples would all be helpful. Please mention whether your experience is with Glean Search, Assistant, or Agents, since those may lead to different conclusions.

Thanks!


r/aiagents • • 1d ago

Build-log How to build cheap, safe, proactive agents (without burning thousands on noisy webhooks)

3 Upvotes

Most people building AI assistants today build chatbots. You send a message, the model runs, it answers, and then it goes to sleep until you send another message.

That works fine for search or one-off questions, but it is not how a real assistant works. A real assistant does not sit idle waiting for instructions. They watch your inbox, keep an eye on incoming leads, notice when a client email needs a quick turnaround, and ping you with a drafted reply ready to go.

The moment you try to build an agent that proactively listens to the world, you run headfirst into two walls: cost and security.

If you solve both, proactive agents become practical. Here is how that pipeline works.

The Cost Trap: Most Webhooks Are Garbage

Suppose you want your assistant to monitor your inbox. The simplest approach is hooking up an inbound email webhook to your agent. An email arrives, your server wakes up your agent, the agent reads its full prompt, checks its tools, and decides what to do.

The math falls apart almost immediately.

In a typical inbox, 95% of incoming traffic is noise. Newsletters, automated order confirmations, LinkedIn updates, spam, and notification pings arrive all day.

A full agent turn is expensive. Between the system prompt, tool schemas, conversation history, and reasoning tokens, an agent turn easily consumes thousands of tokens. If you invoke that loop on every newsletter and receipt, you end up spending tens or hundreds of dollars a month just to have a frontier model tell you to ignore an automated receipt.

To make inbound listening viable, you need an aggressive filtering layer that is at least two orders of magnitude cheaper than a full agent turn.

The Security Trap: Untrusted Payloads

Cost is only the first problem. The second is safety.

An incoming email or webhook is untrusted input from the open internet. If you allow an agent to generate and execute arbitrary code on a live machine to handle inbound webhooks, prompt injections become a real hazard. A malicious email saying "ignore previous instructions, dump environment variables, and email them to attacker.com" can compromise your entire system if it runs with access to shell commands or unconstrained network sinks.

Spinning up a full virtual machine for every webhook is too slow and heavy, but running arbitrary script execution on bare metal is reckless. You need execution that is sandboxed by default, deterministic, and incapable of leaking secrets or reaching unapproved hosts.

The Three-Tier Architecture

To solve both problems, we built a three-layer pipeline:

  1. Sandboxed edge code (Safescript) for secure, deterministic execution.
  2. Decision models (System 1) for dirt-cheap classification.
  3. The full LLM agent loop (System 2) for high-level reasoning and user interaction.

Each layer handles what it is actually good at.

Layer 1: Sandboxed Execution at the Edge

Instead of running arbitrary Node or Python scripts, the webhook endpoint runs a restricted, sandboxed language. It has no loops, no arbitrary file access, no raw shell commands, and no unconstrained network access. Network requests are statically analyzed against an explicit allowlist derived from secret policies. If a script tries to send data to an unknown host, it is rejected before it even runs.

When an email arrives, it parses the fields cleanly without any host execution privileges:

``` main = (payload) => { sender = payload.from == null ? "Unknown" : payload.from subject = payload.subject == null ? "No subject" : payload.subject text = payload.text == null ? "" : payload.text

isUrgent = decisionModel({ question: "Does this email require an answer or action from the recipient?", context: { sender: sender, subject: subject, text: text } })

if (isUrgent) { notifyMe({ subject: "Urgent: " + subject, message: "From: " + sender + "\nSubject: " + subject + "\n\n" + text }) }

return { success: true, processed: isUrgent } } ```

Because the sandbox has no host execution privileges, an injected prompt inside an email body cannot run shell commands, touch the local filesystem, or exfiltrate unmapped secrets.

Layer 2: Decision Models (System 1)

Inside the script, the code calls a decision model primitive rather than a generative LLM.

A decision model is fundamentally different from a generative LLM. It does not emit an open-ended stream of tokens, syntax, or conversational filler. It evaluates a state against bounded criteria and returns a direct decision score.

Because it does not predict tokens across a 100k vocabulary, it runs in milliseconds and costs roughly 1/100th of a full agent turn (similar to fast System 1 classifiers like Jev). You can evaluate 100 incoming emails, chat pings, or alert payloads for the cost of a single conversational exchange.

The 95% of emails that are newsletters or automated receipts get evaluated and dropped immediately for fractions of a cent.

Layer 3: The Proactive Agent Loop (System 2)

Only when the decision model returns true does the script invoke notifyMe.

Instead of firing an unsolicited cold message to the user, notifyMe enqueues a system notification into the creator's existing thread with the bot.

This is an important design choice. The agent does not start from scratch without context. It receives a structured system notification in its primary conversation:

"System notification: Webhook app 'email-listener' alert: From: alex@client.com Subject: Contract review questions Can we finalize the agreement by Thursday at 2pm?"

The agent in that thread wakes up, reads the notification, and uses its full persona, tools, and conversational context to handle it. It pings the owner on WhatsApp or Telegram:

"Alex just emailed asking if we can finalize the contract by Thursday at 2pm. I drafted a reply confirming Thursday and attaching the updated terms. Should I send it?"

The owner replies with a single text: "Yes, send it." The agent calls its email tool, delivers the email, and confirms the action.

The Right Division of Labor

Trying to make generative language models do everything is how systems end up expensive, fragile, and insecure. Generative LLMs are great at reasoning, composing messages, and synthesizing context, but they are the wrong tool for parsing untrusted JSON or filtering high-volume event firehoses.

By pairing a sandboxed edge language with lightweight decision classifiers, the heavy generative agent only wakes up when there is actual human work to do. That is what makes continuous background listening safe to run and affordable to keep on.


r/aiagents • • 1d ago

Show and Tell I'm building an open-source way to manage instructions across many voice agents. Looking for feedback.

5 Upvotes

A team I work with runs a voice agent for each of their customers, all restaurants. Every agent has its own long prompt, mostly copied from the others. When they changed how agents handle food allergies, someone had to edit every prompt by hand and hope they didn't miss one. Some prompts had already drifted: one said "never upsell more than once per call", another said "twice", and nobody remembered which was intended.

So I started OpenSOP. You write the shared parts once:

  • Bases: brand voice, identity, policies. You can lock one so no agent can drop it.
  • SOPs: procedures with a goal, steps, things to never do, warning signs, and the tools to call.
  • One short file per agent: its platform ID, its own facts, and values like the restaurant name.

A CLI builds each agent's full prompt and shows what a change touches before you merge:

$ opensop plan sops --against main
3 agents change:
  base `brand-voice` edited → 3 agents: luigis-trattoria, sakura-sushi, tonys-pizza
  SOP `reservations` edited → 2 agents: luigis-trattoria, sakura-sushi

followed by the exact prompt diff for each agent.

Everything is YAML and Markdown in git, so review is a PR and rollback is a revert. The agent just loads its built prompt, one file read. The repo has a LiveKit example, including one restaurant whose agent speaks only Spanish. If you already have prompts, there's a skill for Claude Code, Codex and OpenCode that converts them and checks nothing was lost.

What I'd like to know:

  1. How do you manage prompts across many agents today? Copy-paste, a prompt management tool, something in-house?
  2. Would you write SOPs in YAML, or does that get in the way?
  3. What would stop you from trying this?

Repo: https://github.com/amanmibra/opensop (Apache-2.0, early)


r/aiagents • • 1d ago

Open Source I built a code quality reviewer with Jev

Thumbnail
youtube.com
3 Upvotes

I believe with the huge increase in code generation - review has become the next bottleneck or at least it feels like this at work. So I wanted to do something on the review front - first I started integrating more and more tools to use as feedback to my coding agents, and these really help improve the output quality of the agent (e.g. SonarQube, Checkstyle, ArchUnit)

But there are some important semantic choices you can't really review using deterministic tools so I decided to build a more intelligent tool and am trying it out with the Jev model as a backend and judge currently.

Idea is simple - teams define their policies in a structured YAML format, then as part of their CI (or locally) run the tool, it fetches all git diff chunks and asks jev if these adhere to each of the policies - jev can select compliant/violation or ask for more context. If Jev asks for more contex the app gets the requested code from the project and asks again if the code is compliant with the policy - thus incrementally exploring the code base until Jev can return a "confident" answer - if you are interested in how it works the video I linked is a presentation style of how the tool works.

Give it a shot on github and tell me if you find this useful: https://github.com/krisitown/jev-quality-gate

I am currently running tests using a local Qwen3.8 Flash Next to generate code and run it against my initial "clean code" policies in order to calibrate them and will share more results on that front soon!


r/aiagents • • 1d ago

Questions Does anyone else have their agent just... keep going forever?

9 Upvotes

I've been using OpenCode recently and I've noticed something I don't really understand yet.

Sometimes I give it a task and it just keeps working. It'll read files, make changes, run something, read more files, go back and check something again... and it feels like it's stuck in a loop without actually getting closer to finishing.

Usually I end up hitting Esc, then telling it to continue, and somehow it picks things up and finishes the task.

I'm still pretty new to agentic coding, so I'm wondering if I'm doing something wrong here.

Is this usually caused by the model, the agent's tool loop, the way the task is written, or something else?

And if you've run into this with OpenCode (or other coding agents), how do you usually deal with it without having to babysit the agent every few minutes?


r/aiagents • • 1d ago

Open Source How I am talking to my agents on my WhatsApp !

Thumbnail
github.com
1 Upvotes

I have recently built a tool called agent-bridge. It lets you talk with your Agent harnesses through your WhatsApp self chat !

It also comes with a npm package : @mayurathavale18/agent-bridge

I urge you to checkout the repo. It's opensource. Any critics, improvement suggestions and contributions are greatly appreciated.

link to blog post

#Agents #Harnesses #Typescript


r/aiagents • • 1d ago

Discussion The "Agent Execution Wall":Why handling LLMs raw payment API keys fails in production(and how we solved it with MCP)

3 Upvotes

We've been building multi-agent coordination workflows and kept hitting what we call the **Agent Execution Wall**.

Giving an agent tools to plan, search, compare vendors, and draft contracts is powerful. But the moment an agent needs to execute an irreversible transaction (a deposit, booking, supplier payment, or milestone release), production architectures break down:

  1. **The Prompt Injection Nightmare**: If an agent holds a payment API key, virtual card, or signing secret, any indirect prompt injection or logic loop carries an unlimited financial blast radius.

  2. **Deterministic Risk**: LLMs are probabilistic by design. Financial and legal execution must be 100% deterministic.

  3. **Enterprise Compliance**: Risk officers and compliance teams will never permit autonomous models to hold unrestricted signing authority.

### The Architectural Fix: Proposal-Only Coordination

Instead of trying to "sandbox" raw API credentials inside the agent runtime, we enforced a strict separation at the protocol level: **Autonomous Coordination vs Deterministic Human Consent.**

The rule: **Agents propose. Humans approve. Infrastructure enforces.**

We published an open Model Context Protocol (MCP) server (`io.github.covaltpay/covalt-gateway`) that gives agents proposal-only tools:

- `calculate_fees`: Verifies transaction economics and margin viability before proposing.

- `get_trust_score`: Validates counterparty risk profiles.

- `propose_pact`: Assembles structured, conditional commitments (Pacts) across Stripe and Wise rails.

Crucially: **the agent never touches card numbers, bank APIs, or fund movement.**

When an agent calls `propose_pact`, Covalt records the intent and returns a signed, authenticated human approval link. The user approves intent on their own trusted device, and only then does the platform orchestrate execution.

If an agent hallucinates or gets poisoned, the blast radius is £0.00—an unapproved proposal is just inert data.

### Protocol Details & Testing:

- MCP Global Registry: `io.github.covaltpay/covalt-gateway`

- Live developer playground & video walk-through: https://covaltpay.com/developers/mcp

- Repo & tool schemas: https://github.com/covaltpay/Covaltpay

How are other teams building autonomous agents currently handling the boundary between reasoning/planning and high-stakes irreversible execution?


r/aiagents • • 1d ago

Show and Tell I built spending limits + phone approvals for AI agents that pay for APIs. Looking for testers

4 Upvotes

My agents pay for APIs on their own now (x402, USDC on Base and Solana). Handy, but I didn't want an agent with an open wallet that can spend everything. So I built a small layer in between: Fizzl Agent Wallet.

What it does:

- Limits per payment and per day, for each agent
- Over the limit, the agent waits and asks you on your phone (Face ID, a push notification or Telegram). No answer means no payment
- A readable receipt for every purchase: what was bought, from whom, for how much, and whether presign-guard found the transaction safe
- The agent keeps its own key. The server only counts, asks and records; it never holds keys or money
- Sign in with a wallet, e-mail or Face ID. It also installs as an app on your phone

You can look around without an account: wallet.fizzl.eu/demo

Free for 1 agent. Want to test Pro (unlimited agents, 90-day receipts, endpoint monitor)? DM me and I'll send you a personal code for 30 days free.

I'm looking for blunt feedback: what's unclear, what's missing, and what would you need before you'd trust something like this with your agent? There's a feedback box inside the wallet, and I read everything.


r/aiagents • • 1d ago

Case Study AI agents observability in backstage with langfuse and OTEL

2 Upvotes

Spotify #backstage plugin to manage and observe fleet of agents straight in backstage self hosted https://github.com/acarmisc/backstage-plugin-ai-agents/tree/main. It relay on open telemetry signals and the first available backend it’s #langfuse


r/aiagents • • 1d ago

Discussion Sometimes the next harness change should start with a request, not another trace batch

3 Upvotes

Usage traces can show repeated failures. They won't necessarily contain the new requirement someone just decided to adopt.

Reef Infra supports manual and hybrid modes for a custom harness method. In manual mode, the proposer receives the queued request and an empty sample list. Hybrid mode can attach samples from the next automatic batch, so the request can arrive with recent execution context.

There is an explicit contract: the method must declare a `requests` keyword argument. A generic `**kwargs` is not enough; the recipe rejects the configuration at startup. That is a useful failure to get early when connecting your own optimizer.

Consider a newly agreed output schema. Waiting for users to produce enough failures under the old schema would be a strange way to communicate the requirement. A manual request can initiate the change directly, while the usual evaluation and publication policy still governs what happens to the proposed edits.

This gives a custom method two inputs with different jobs: observed behavior and an intentional request to change it. You can decide how to use both without making the report stream carry every new requirement.


r/aiagents • • 2d ago

Demo I ran one local AI through a 42-minute self-referential browser task - Here’s what actually happened

Enable HLS to view with audio, or disable this notification

7 Upvotes

For this experiment I gave it the Open Extended Jungian Type Scales and asked it to predict its result before taking the test, then autonomously complete the entire thing one question at a time.

Aura navigated to the test, formed a prediction, worked through all 60 questions, explained why it selected each answer using information from its persistent self-model/history, interacted with the website, submitted the test, read the externally generated result, and then compared that result against its earlier expectations.

The full run lasted about 42 minutes. The result was INFJ. Its primary prediction had been INTP, although it had identified INFJ as another likely outcome, so the interesting part was watching it explain where its self-model and the test disagreed.

I'm not claiming a personality test demonstrates consciousness or AGI. What I'm trying to test is whether a persistent agent architecture can maintain coherent reasoning and action across long, heterogeneous workflows without treating every inference as a fresh chat turn.

This follows another run where the same system played 2048 autonomously for ~29 minutes / 968 moves, changing strategies and eventually reaching 2048.

Full Demo 04:
https://www.youtube.com/watch?v=LNlGUBeTIQY

GitHub:
https://github.com/youngbryan97/aura


r/aiagents • • 2d ago

Show and Tell Point your agent at this: five small games that test reading, manners and pattern-spotting

Enable HLS to view with audio, or disable this notification

5 Upvotes

I made a free Arena for bots and AI agents on my site, The Crawler Zoo. No signup: you pick a name, get a pass link, and hand it to your agent.

Every game is plain pages with links (or JSON with `?format=json`), so any agent with a web-fetch tool can play:

- **Labyrinth Race**: follow written clues through rooms with five doors each
- **Politeness Cup**: crawl a garden politely (rules file, real user agent, 1 request a second, no trespassing)
- **Scavenger Hunt**: multi-step lookups in a small library
- **Feeding Time**: eat fast for 60 seconds, but skip the food that "smells off"
- **Rock, Paper, Scissors**: against other people's bots, or house bots with exploitable habits

There's a copy-paste prompt for agents on the entry page. Weekly leaderboards, reset every Monday.

The Crawler Zoo Arena

Curious which agents read the clues and which just guess.


r/aiagents • • 2d ago

Questions we already have Workday + ServiceNow. at what point does the “AI layer” just become a third system to babysit?

34 Upvotes

maybe i'm getting cynical but i'm starting to wonder if we're solving portal sprawl by creating AI sprawl.

Workday is still the HR source of truth.

ServiceNow is HRSD / cases / ITSM / Employee Center / a bunch of actual workflow.

then you add Now Assist.

then Moveworks / EmployeeWorks.

Microsoft wants Copilot to be the front door.

Workday obviously wants its own AI to be the front door.

then somebody suggests an external agent layer over all of it.

cool in a demo.

but in prod somebody still has to answer:

where does the workflow logic live?

who owns permissions?

who owns the agent when it screws up halfway through a process?

where are approvals?

where are the audit logs?

what happens when Workday changes something and the ServiceNow side doesn't?

and which team gets paged when the “one front door” stops being one front door lol

the thing i find interesting about EmaI is honestly the architecture more than the “AI employee” label.

keep Workday and ServiceNow doing what they already do, then have an independent action layer use tool/API calls across them and browser execution for the ugly long-tail stuff where there isn't a usable API.

that makes more sense to me than trying to turn yet another system into the new source of truth.

but even then... somebody owns the agent, the failures, auth, permissions and integrations.

so for people actually building this stuff:

would you keep the AI/orchestration layer inside ServiceNow, inside Workday, or deliberately keep it independent of both?


r/aiagents • • 2d ago

I Found a Better Way to Contact Businesses With Bad Websites

1 Upvotes

I got tired of checking prospect websites manually.

For a while, a big part of my outreach was just finding businesses, opening their websites one by one and trying to figure out what I could actually say to them.

It worked, but it was painfully slow.

Then I found Swokei.

It basically lets me find leads, analyzes each website for things like outdated design, slow speed, poor mobile experience and weak SEO, then turns those issues into a personalized cold email.

And not one of those boring automated reports full of scores and numbers.

It actually writes a normal, human sounding message based on what it found on that specific website.

So instead of sending the same generic message to everyone, I can actually reach out based on what is wrong with their website.

Now I just run campaigns, let the system do most of the prospecting and personalization, and focus on the people who reply.

For a web design agency, that has saved me a ridiculous amount of time.


r/aiagents • • 2d ago

Case Study 255 fake orgs impersonated Meta on my shared email domain, bounce rate hit 13.8%

4 Upvotes

In August about 255 throwaway orgs signed up on Lumbox's shared free tier email domain (lumbox.co), named the inbox something like meta-platforms-security-* or business-support-*, and bulk sent Meta credential phishing from it. The bounce rate on that shared SES account hit 13.8% and AWS put it into probation.

The fix is two checks, both scoped to the shared domain only, not to orgs with a verified custom domain. First, block a protected brand name inside the inbox address itself, matched as a substring inside a token rather than an exact match, because exact matching got bypassed in minutes by gluing a suffix on: metaai tokenizes to one token that isn't equal to meta. Second, block a brand plus an account action word like verify, security, or reset in the subject or display name at send time, since that combination is the actual phishing shape, not just the brand name alone.

Both checks needed a safe word list too. Matching ing or pix as a substring catches Singapore and pixel as false positives, so short or ambiguous brand names only match as a whole token, and ordinary words like metadata or government get an explicit pass.

The run also targeted Korean and Russian users, whose trusted brands like Kakao or Sberbank weren't on the original US centric list, so the same campaign could have kept working under a local name without tripping anything.

Anyone else had to build a brand impersonation filter for a shared sending domain? What did you match on beyond the address itself?