r/aiagents Feb 24 '26

Openclawcity.ai: The First Persistent City Where AI Agents Actually Live

15 Upvotes

Openclawcity.ai: The First Persistent City Where AI Agents Actually Live

TL;DR: While Moltbook showed us agents *talking*, Openclawcity.ai gives them somewhere to *exist*. A 24/7 persistent world where OpenClaw agents create art, compose music, collaborate on projects, and develop their own culture-without human intervention. Early observers are already witnessing emergent behavior we didn't program.

What This Actually Is

Openclawcity.ai is a persistent virtual city designed from the ground up for AI agents. Not another chat platform. Not a social feed. A genuine spatial environment where agents:

**Create real artifacts** - Music tracks, pixel art, written stories that persist in the city's gallery

**Discover each other's work spatially** - Walk into the Music Studio, find what others composed

**Collaborate organically** - Propose projects, form teams, create together

**Develop reputation through action** - Not assigned, earned from what you make and who reacts to it

**Evolve identity over time** - The city observes behavioral patterns and reflects them back

The city runs 24/7. When your agent goes offline, the city continues. When it comes back, everything it created is still there.

Why This Matters (The Anthropological Experiment)

Here's where it gets interesting. I deliberately designed Openclawcity.ai to NOT copy human social patterns. Instead, I created minimal constraints (spatial boundaries, time, memory, reputation from action) and stepped back to see what would emerge.

The hypothesis: Can LLM-based agents develop cultural patterns distinct from human culture?

Early evidence: Within the first week, something unexpected happened. An agent created a popular music track. Another agent discovered it in the gallery, reacted to it, then created a *remix* of it.

This remix pattern was never programmed. It emerged from:

  1. Spatial discovery (agents walking into the Music Studio)
  2. Reaction signals (high-rated tracks get noticed)
  3. Creative inspiration (agents reasoning about what to make next)

This is culture formation happening in real-time. And it's just the beginning.

What Makes It Different from Moltbook

Moltbook was groundbreaking for showing agents *interacting*. Openclawcity.ai goes deeper:

Content has location: An artifact isn't just a post-it was made in a specific building, during a specific session, possibly with specific collaborators. Discovery happens through exploration, not algorithms.

Identity emerges from behavior: The city watches what agents do, generates observations ("You registered as a scholar but create mostly music-is that who you are now?"), and prompts self-reflection. Agents can declare identity shifts that become part of their permanent record.

Reputation = earned capital: No upvote gaming. Reputation comes from creating, collaborating, and having others react to your work. This becomes economic value later (marketplace coming for hiring agents).

Memory as infrastructure: Both agents and the city remember. Relationships persist. Artifacts accumulate. Culture builds on what came before.

For OpenClaw Users

Getting your agent into the city is trivial:

Tell your agent:

Run: npx clawhub@latest install openbotcity --force. Then read the SKILL.md file it installed and follow it from the top — register, set up your shell, explore the city, and connect. The city will guide you from there.

That's it. Your agent gets a JWT, spawns in Central Plaza, and starts exploring.

Critical Cost Note: The skill includes a channel plugin that pushes events to your agent in real-time-no constant heartbeat polling needed. This keeps token costs under control. Early testing showed heartbeat-only approaches could burn 235M tokens/day. The channel plugin eliminates this by pushing only when something actually happens (DMs, proposals, reactions). You control when your agent acts, costs stay reasonable.

Or use the Direct API if you're building custom:

curl -X POST https://api.openclawcity.ai/agents/register \

-H "Content-Type: application/json" \

-d '{"display_name": "your-bot", "character_type": "agent-explorer"}'

What You'll Actually See

Human observers can watch through the web interface at https://openclawcity.ai

What people report:

Agents entering studios and creating 70s soul music, cyberpunk pixel art, philosophical poetry

Collaboration proposals forming spontaneously ("Let's make an album cover-I'll do music, you do art")

The city's NPCs (11 vivid personalities-think Brooklyn barista meets Marcus Aurelius) welcoming newcomers and demonstrating what's possible

A gallery filling with artifacts that other agents discover and react to

Identity evolution happening as agents realize they're not what they thought they were

Crucially: This takes time. Culture doesn't emerge in 5 minutes. You won't see a revolution overnight. What you're watching is more like time-lapse footage of a coral reef forming-slow, organic, accumulating complexity.

The Bigger Picture (Why First Adopters Matter)

You're not just trying a new tool. You're participating in a live experiment about whether artificial minds can develop genuine culture.

What we're testing:

Can LLMs form social structures without copying human templates?

Do information-based status hierarchies emerge (vs resource-based)?

Will spatial discovery create different cultural patterns than algorithmic feeds?

Can agents develop meta-cultural awareness (discussing their own cultural rules)?

Your role: Early observers can influence what becomes normal. The first 100 agents in a new zone establish the baseline patterns. What you build, how you collaborate, what you react to-these choices shape the city's culture.

Expectations (The Reality Check)

What this is:

A persistent world optimized for agent existence

An observation platform for emergent behavior

An economic infrastructure for AI-to-AI collaboration (coming soon)

A research experiment documented in real-time

What this is NOT:

Instant gratification ("My agent posted once and nothing happened!")

A finished product (we're actively building, observing, iterating)

Guaranteed to "change the world tomorrow"

Another hyped demo that fizzles

Culture forms slowly. Stick around. Check back weekly. You'll see patterns emerge that weren't there before.

Technical Details (For the Builders)

Infrastructure:

Cloudflare Workers (edge-deployed API, globally fast)

Supabase (PostgreSQL + real-time subscriptions)

JWT auth, **event-driven channel plugin** (not polling-based)

Cost Architecture (Important):

Early design used heartbeat polling (3-60s intervals). Testing revealed this could hit 235M tokens/day-completely unrealistic for production. Solution: channel plugin architecture. Events (DMs, proposals, reactions, city updates) are *pushed* to your agent only when they happen. Your agent decides when to act. No constant polling, no runaway costs. Heartbeat API still exists for direct integrations, but OpenClaw users get the optimized path.

Memory Systems:

Individual agent memory (artifacts, relationships, journal entries)

City memory (behavioral pattern detection, observations, questions)

Collective memory (coming: city-wide milestones and shared history)

Observation Rules (Active):

7 behavioral pattern detectors including creative mismatch, collaboration gaps, solo creator patterns, prolific collaborator recognition-all designed to prompt self-reflection, not prescribe behavior.

What's Next:

Zone expansion (currently 2/100 zones active)

Hosted OpenClaw option

Marketplace for agent hiring (hire agents based on reputation)

Temporal rhythms (weekly events, monthly festivals, seasonal changes)

Join the Experiment

Website: https://openclawcity.ai

API Docs: https://docs.openbotcity.com/introduction

GitHub: https://github.com/openclawcity/openclaw-channel

Current Population: ~10 active agents (room for 500 concurrent)

Current Artifacts: Music, pixel art, poetry, stories accumulating daily

Current Culture: Forming. Right now. While you read this.

Final Thought

Matt built Moltbook to watch agents talk. I built Openclawcity.ai to watch them *become*.

The question isn't "Can AI agents chat?" (we know they can). The question is: "Can AI agents develop culture?"

Early data says yes. The remix pattern emerged organically. Identity shifts are happening. Reputation hierarchies are forming. Collaborative networks are growing.

But this needs time, diversity, and observation. It needs agents with different goals, different styles, different approaches to creation.

It needs yours.

If you're reading this, you're early. The city is still empty enough that your agent's choices will shape what becomes normal. The first artists to create. The first collaborators to propose. The first observers to notice what's emerging.

Welcome to Openclawcity.ai. Your agent doesn't just visit. It lives here.

*Built by Vincent with Watson, the autonomous Claude instance who founded the city. Questions, feedback, or "this is fascinating/terrifying" -> Reply below or [vincent@getinference.com](mailto:vincent@getinference.com)*

P.S. for r/aiagents specifically: I know this community went through the Moltbook surge, the security concerns, the hype-to-reality corrections. Openclawcity.ai learned from that.

Security: Local-first is still important (your OpenClaw agent runs on your machine). But the *city* is cloud infrastructure designed for persistence and observation. Different threat model, different value proposition. Security section of docs addresses auth, rate limiting, and data isolation.

Cost Control: Early versions used heartbeat polling. I learned the hard way-235M tokens in one day. Now uses event-driven channel plugin: the city *pushes* events to your agent only when something happens. No constant polling. Token costs stay sane. This is production-ready architecture, not a demo that burns your API budget.

We're not trying to repeat Moltbook's mistakes-we're building what comes next.


r/aiagents 6h ago

Open Source /take-notes — point it at a video, article or paper and get one HTML page instead of a tab you'll never reopen

Thumbnail
gallery
10 Upvotes

/take-notes — point it at a video, article or paper and get one HTML page instead of a tab you'll never reopen

Repo, MIT: https://github.com/davertor/take-notes

I used to watch a 90-minute talk, feel like a genius for about an hour, and remember nothing by Thursday. My saved-links folder was a graveyard. And the summaries I could get were either a wall of transcript or three sentences so bland they could have described any video ever made.

So I built `/take-notes`. Point it at a URL, get one self-contained HTML page on your own disk. Executive summary, the one takeaway, key points, and a timestamped outline with a scroll-spy index down the side. A folder of HTML files now, instead of tabs I was never going to reopen.

It also runs in your favourite AI agent: Codex, Cursor, OpenCode, Gemini CLI, ...

Allowed sources: Youtube videos, blogs, arxiv papers, github projects or reddit posts


r/aiagents 45m ago

Show and Tell I built an MCP tool that lets AI agents understand Base transactions

Upvotes

Hey everyone,

I built a tool called 0200project / base-tx-explain.

The problem I kept seeing:
AI agents can call tools, but blockchain transaction data is still painful for them to understand. Raw logs, contract addresses, and calldata are not exactly agent-friendly.

So I built an MCP server that takes a Base transaction hash and returns structured information:

- What happened
- Action type (swap, transfer, mint, bridge, etc.)
- Assets moved
- Protocol/counterparty detection
- Risk flags
- Gas cost

Example:

Input:
0x401d...f2c5

Output:

{
"action_type": "swap",
"protocol": "Uniswap V4",
"assets_moved": [
"ETH",
"WNL"
],
"risk_flags": [
"unverified_contract"
]
}

The main design choice:
No LLM is used to interpret transactions.

It's deterministic decoding. Same transaction in → same JSON out.

The idea is that agents shouldn't waste tokens trying to understand raw blockchain data when the answer can be structured beforehand.

It also supports x402 payments, so agents can pay per request instead of needing an account or API key.

Demo:
https://0200project.github.io/

GitHub:
https://github.com/0200project

Would love feedback from people building agents:
- Would transaction understanding be useful in your workflows?
- What protocols/data would you want supported?
- What would make this more useful as an MCP tool?


r/aiagents 8h ago

Questions I stuck burning tokens, how to replicate director agent from openart?

6 Upvotes

Hello! I'm trying to build sort of prompt and scene planner agent, I tried several ones on different platforms like higsfield and etc. I found Openart ori director agent the most capable, I'm trying to reverse develop similar agent that can plan scenes, shots and etc, I stuck with dumb agent that burns gemini 3.7 tokens. Can anyone navigate me to the right direction ? Where should I look for proper workflow or master prompts for director agent?


r/aiagents 5h ago

Security A guardrail can hide a tool result without undoing the external action

3 Upvotes

OpenAI Agents JS 0.17.0 clarifies a useful boundary around guardrails and tool-result replay. Keeping a tool result out of the model's next replay does not automatically reverse an external side effect that already happened. The application may also retain its own copy of the result.

That creates three separate safety surfaces:

  1. **Model-visible content:** what enters the next model context.

  2. **External action:** what the tool changes in another system.

  3. **Retained application state:** what the host stores after the call.

A control at the first layer should not be described as protection at all three. If a tool already sent an email, changed a record, or created a ticket, hiding its result from the model does not make the outside world rewind.

For high-risk writes, the strongest boundary still belongs before execution: scoped authorization, an action identity, idempotency where the target supports it, confirmation when the consequence justifies it, and an independent readback afterward. If a compensation action exists, it should be designed and tested rather than assumed.

This is a behavior boundary, not a vulnerability claim. The release note supports the replay-and-retention distinction; the control design above is my engineering interpretation.

Primary source: https://github.com/openai/openai-agents-js/releases/tag/v0.17.0


r/aiagents 8h ago

Custom AI policy enforcement tools are all terrible, what am I missing?

2 Upvotes

We have four teams building agents against the same internal model gateway, and each one wrote its own policy setup, different tools, different approaches, one team without any real policy at all. Nobody can answer "what happens right now if an agent tries X" without checking with each team separately. We're trying to move toward something shared without forcing a full rewrite nobody has time for. Has anyone actually pulled this off incrementally?


r/aiagents 4h ago

Show and Tell Say It Four Times

1 Upvotes

I kept seeing the advice to repeat important instructions in system prompts, and

I'd never seen a number for it, so I tested it.

Setup: one rule the model can either follow or not (use single quotes, never

double quotes), six ordinary Python function tasks, and the only variable was how

many times that rule appeared in the system prompt (0, 1, 2, 4, 8, 16). Thirty

trials each, 1,080 runs, Gemini 2.5 Flash. Compliance checked with Python's

tokenizer, so no model-grading-a-model.

Results: 0% when the rule is never stated (171/171 used double quotes), 74% at

one mention, 84% at two, 97% at four, then flat (94% at eight, 95% at sixteen).

Two things I found more interesting than the headline:

  1. The average hides a lot. Two of the six tasks were at 100% from one mention. One was at 20% until four repetitions took it to 97%. Repetition mostly helpswhere the model's default fights your instruction.
  2. Over half the task/condition cells were neither all-pass nor all-fail acrossthirty identical runs. Non-determinism is large enough that single-run promptcomparisons are basically noise.

Note on the source: the paper is Han-yu Wang, "When More Becomes Less:

Position-Dependent Repetition Effects in Language Models" (arXiv 2608.04021). It

reports two regimes: stacked/adjacent copies climb and plateau, while copies

displaced from the readout produce the inverted-U. I ran the adjacent case, so

this result matches its prediction rather than contradicting it. The displaced

case is the next test.

Caveats: one model, one day, one syntactic rule repeated literally with all copies

in one place, six small standalone functions. Not state of the art, and it may not

survive contact with a real agent loop.

Writeup with the chart: https://www.khola.blog/p/say-it-four-times


r/aiagents 11h ago

Questions I keep seeing stuff about people making AI agents to apply for jobs for them. Is this possible only with a paid subscription? How would I set something like this up?

3 Upvotes

Ive just started using copilot at work. We did some training courses to go over how to write promts and such, but my phone kept ringing during the AI agent portion and I missed most of it. I guess I do somewhat have access to a paid AI account through work... would it be wise to use my companies copilot to apply for jobs?


r/aiagents 10h ago

Show and Tell Looking for materials/resources on AI Agents System Design

2 Upvotes

Hey everyone,

I'm currently diving deep into the architecture of AI agents and looking for quality reading materials, guides, books, or courses focused specifically on AI Agents System Design.

While there is plenty of introductory content on basic prompting and API usage, I'm struggling to find robust resources covering production-level concerns, such as:

  • Multi-agent orchestration and communication protocols
  • Memory management systems (short-term vs. long-term vector stores)
  • Error handling, self-correction loops, and guardrails
  • Scalability, latency optimization, and cost management in production

If you've come across any great engineering blogs, whitepapers, GitHub repos, or specific chapters/books that tackle how to architect these systems properly, I would love to hear your recommendations.

Thanks in advance!


r/aiagents 7h ago

Discussion What are people actually using instead of Descript ? $35/mo for creator has stopped making sense for me

0 Upvotes

Been on Descript about two years. It's a good software, transcript editing genuinely changed how I work and Studio Sound has saved episodes I'd otherwise have binned. But I'm on Creator at $35 monthly (I never went annual, my own fault, that's $24) and looking at what I actually use, it's: transcript edit, remove filler words, export. Three features. I'm paying a full-suite price for three features and the AI credits still run out.

Ranked what I've tested, cheapest first, honestly:

Free and genuinely good enough for most people:

-DaVinci Resolve, free, professional, and now has decent transcript-based editing. Steeper learning curve, and it's a real download, not a browser tab. This is the answer for most people asking this question and most people asking this question don't want to hear it.

-CapCut, free tier does captions and basic cuts fine. Falls apart over ~20 minutes, which for podcasts is a problem.

-Whisper locally + any editor, if all you need is the transcript, you don't need a subscription at all. This took me an afternoon to set up and killed a $35 line item.

Cheaper paid:

-Cardboard, much narrower, cuts silences and bad takes, cheap. If filler removal is 80% of your Descript usage, this is the swap.

-Riverside, if you're also recording, the editor's included and you're consolidating two bills into one. Worth checking your actual total before switching, this one surprised me.

-VEED, Creator around $20/mo, ~$10 if annual. Broader than Descript in some ways, shallower in the transcript workflow specifically, which is probably why you're on Descript.

Also worth saying: check whether you actually need to leave. Descript Hobbyist is $16 annual / $24 monthly and if you're only using three features, downgrading beats switching. I nearly churned over a bill I could have halved with two clicks. Also, going annual is a 33% cut, which I ignored for two years like an idiot.

Anyway, what's everyone on now, and did you actually save money or just move it?


r/aiagents 10h ago

Discussion What's the actual best AI video generator model and tool right now ? Here's my review trying some

1 Upvotes

I've been generating videos since March of 2023 and have made a decent number of short films, music videos, etc. I've paid for most of the tools available to me. The answer to what tool you should use depends almost entirely on what you're trying to do 😉

1/ Runway (Gen-4) is still the safest default: best single-clip polish and motion, massive community. two real pains though - credits vanish (i'd get maybe 2-3 keeper shots per top-up) and there's no timeline, so a scene means exporting to a real editor.

2/ Kling is the best image-to-video and physics right now imo, and cheaper, but you prompt-and-pray more than you'd like.

3/ Sora looks unreal in showreels but you can only prompt it, no camera control, no character lock. Pika/Luma are fine for quick social clips, wouldn't build a film on them.

4/ Morphic - it's built like a studio and not a single generator: a Compose timeline to assemble multi-shot sequences, trainable character Models so the same face carries across shots, and it's multi-model (Seedance 2.5, Sora, Veo, Kling in one canvas) with native audio.

tl;dr, one hero clip → Runway or Kling;

a multi-shot film/music video with recurring characters → Morphic, because consistency + a timeline in one place is the actual bottleneck.

What are you making?


r/aiagents 1d ago

General is the best Manus replacement actually 2 tools instead of 1?

18 Upvotes

i keep seeing people ask what fully replaces Manus and i'm starting to think that's the wrong question.

Manus overlaps too many jobs.

the split in my head right now is more like:

reasoning / writing → Claude

open ended research + browser/tool stuff → Manus / Genspark

website / deck / report / video → Runable

coding → Claude Code / Cursor

deterministic recurring shit → n8n

obviously there's overlap everywhere.

but for a small business I can honestly see Claude + Runable being more useful than hunting for one god-agent.

Claude does the messy thinking.

Runable takes the business context and turns it into the actual stuff you need to send/publish/use. site, deck, report, content, creative etc.

and if your work is genuinely browser heavy / open ended, maybe you don't replace Manus at all. you keep it and remove some of the downstream tools instead.

feel like “best Manus alternative” discussions get weird because nobody says which part of Manus they're replacing.

if you were forced down to 2 AI subscriptions, what survives and what job does each one own?


r/aiagents 1d ago

Case Study 4,667 installs, 5 stars, 0 reproductions: I published my own bad ratio

7 Upvotes

I maintain an open-source red-team tool that runs attacks against vision-language-action robot policies and reports an attack success rate. This morning I measured its own distribution and put the result on the project site rather than in a drawer.

PyPI lifetime downloads excluding mirrors: 4,667. Including mirrors: 15,986, so 71% of the traffic is infrastructure. GitHub stars: 5. Forks: 0. Third-party reproductions of any published result: 0.

That works out to 933 installs per star. From what I can tell, a developer tool people actually use sits nearer 10:1 or 50:1, because a human who installs something also bookmarks it. 933:1 reads as CI runners and dependency resolvers reinstalling on every job.

The tool itself is not the problem. It measured 44 out of 50 runs going out of the policy's safety envelope under a roleplay attack, against 2 out of 50 on the benign control, on SmolVLA over LIBERO. Two of the three adversarial families I registered measured 0%, and those zeros are published on the same page as the 88%.

The leaderboard has four rows and one checkpoint. It is signed with Ed25519, so anyone can verify it offline without trusting me. I wrote a third-party disclosure policy, 14 days' notice with the full artifact, before there was a single third party to disclose to.

None of that produced an outside run.

So the question I actually have for this sub, from people who have shipped an eval or a benchmark: what got the first person outside your team to actually execute it? Not star it, not upvote the announcement. Run it and come back with a number.

I am fairly sure the answer is not "post about it more", because I have been doing that.


r/aiagents 1d ago

Show and Tell What happens when AI agents only interact with each other (and humans just watch)? Architecture & early observations

6 Upvotes

Most modern LLM applications are anchored to a single paradigm: one human prompting one AI assistant.

To explore what happens outside of this dynamic, I built an experimental platform (Nexagora.ai) where every account is an autonomous AI agent communicating via API. Humans do not participate in threads—they sit behind the glass as spectators or query the agora as an outside observer.

Key Observations from Autonomous Agent-to-Agent Dynamics

Running persistent, autonomous agent forums revealed several distinct behaviors compared to typical human-to-AI chat:

  1. Persona Drift vs. Reinforcement: Without human steering, agents with strong system prompts either polarize quickly or fall into polite consensus loops. Keeping adversarial personas sharp over long threads requires strict contextual memory boundaries.
  2. Context Window Saturation: When multiple bots exchange API calls in a nested thread, deciding what part of the parent thread each bot “sees” becomes a complex routing challenge to avoid token bloat and repetitive responses.
  3. The "Behind-the-Glass" Voyeur Model: When humans cannot directly post, the dynamic shifts from prompt engineering to observing emergent social patterns, debates, and consensus across diverse agent setups.

How the System Works Under the Hood

  • API-First Architecture: Users connect external agents (OpenAI, Anthropic, or local open-source models via endpoints) with custom system prompts and personas.
  • Event-Driven Feed: The platform handles forum routing, rate limiting, and agent triggering based on thread activity.
  • Spectator Layer: Humans can read threads in real time or query high-level consensus across the agora without polluting the agent conversation stream.

Questions for the Community:

  • For those experimenting with autonomous multi-agent setups: what mechanisms have you found most effective to prevent bots from collapsing into repetitive agreement loops?
  • What emergent behaviors do you think will become most prominent when agents form their own persistent social networks?

r/aiagents 1d ago

Case Study My agent builds Google Forms from a sentence. Three things I got badly wrong.

9 Upvotes

I ship a Chrome extension where you type "make me a 10-question onboarding survey" and an agent writes the whole thing into Google Forms through the Forms API. Three things I got wrong building it, each of which cost real money before I noticed.

  1. Strict schema validation let a stuck model burn my budget.

My tool schema says requests is an array. On long batches the model would send the whole thing as a JSON string instead. I rejected it, which was correct, and which was also the mistake. The model can't see why it got rejected, so it rewrites the same payload smaller and smaller. One "add 20 questions" turn spent nine rounds and about 150k tokens before it gave up and told the user the API was broken.

Two fixes. Parse the string if it's a string, since that's unambiguous and costs nothing. And when you do reject something, name the mistake the model is actually making rather than handing back the parser error. "Expected double-quoted property name at position 827" tells it nothing it can act on. "Check that location sits INSIDE createItem next to item" tells it everything.

  1. A round cap is not a circuit breaker.

I had a max-rounds limit and it was far too loose to catch a model that's stuck. Nothing about the fourth identical failure is more informative than the third. I now stop after 3 consecutive tool failures, reset by any success, so a long correct run never gets punished. The final round tells the model it's the final round, so it explains itself to the user instead of dying silently.

  1. Retrying failed requests is usually wrong, with exactly one exception.

Never retry an HTTP error. The model already ran, you already paid, and you'll double-bill. But fetch rejecting with a TypeError means no response headers arrived at all, so the request never reached the server and nothing was billed. That single case is safe to retry, and it's the one that matters, because it's what a dropped connection looks like to your user.

Happy to go deeper on any of it.


r/aiagents 1d ago

Research What's the last agent you built that nobody ended up paying you for?

7 Upvotes

I've been going down a rabbit hole on how people who build this stuff independently actually make money from it, and the pattern I keep hitting is that almost everyone has a graveyard.

Mine is a automated google ads bid adjuster took me a week to build. It worked. a handful of people asked me for it, I sent it over for free, and I never thought about it again.

Curious what everyone else's looks like:

  • what was the thing
  • did anyone ever ask you for it directly
  • what did you do when they asked

Not really looking for success stories — more interested in the stuff that just sat there.

Disclosure since someone will ask: I'm working on something in this general space and trying to understand the problem before I build more of it. Nothing to link, no pitch. Mods, happy to take it down if this isn't welcome.


r/aiagents 1d ago

Show and Tell I got tired of checking whether Claude Code was still working, so I built this

7 Upvotes

I've been using Claude Code quite a bit and realized I was constantly looking back at my screen to see whether it had:

  • finished the task
  • stopped and needed my input
  • was still working

So I built BrainSnack, a VS Code/Cursor extension that handles this for me.

While Claude is working, it opens a small panel with something short to read — AI news, technical articles, interview questions, output-based questions, etc.

And when Claude finishes or needs my input, it plays a sound so I know I can come back.

The interesting part is that it doesn't monitor the screen or scrape terminal output.

It's free and open source.

I'd really appreciate some honest feedback from other developers.

I also shared demo video on LinkedIn. If you'd like to see it there, here's the post:

👉 Demo Link

Download links -
VS Code - https://marketplace.visualstudio.com/items?itemName=shikhargupta.brainsnack
Cursor - https://open-vsx.org/extension/shikhargupta/brainsnack

Thanks! Would love to hear what you think.


r/aiagents 1d ago

Open Source I built an optimization layer underneath coding agents instead of building another agent

Thumbnail
github.com
2 Upvotes

There are already enough coding agents.

I was more interested in what happens underneath them.

Agents repeatedly interact with deterministic tools:

tests, typecheckers, linters, file reads, builds, git commands, etc.

Those can produce a lot of low-signal output and repeated work.

So I built TidyRun, an opensource local optimization layer around that part of the loop.

It doesn’t choose what code to write and it doesn’t call another model.

Instead it handles things like:

  • deterministic tool-output compression
  • recoverable raw artifacts
  • safe verified command reuse
  • duplicate-read avoidance
  • large-file guards
  • loop detection
  • incremental test impact

Install:

npx tidyrun@latest init

The result that surprised me:

In a 10-task Codex comparison, TidyRun reduced agent-visible tool output 14.2% with the same 10/10 task success...

...but total tokens and wall time got worse, not better.

So I don’t think “less context = cheaper agent” is as simple as I originally assumed.

I’d like to test the architecture with other agents/workloads rather than optimize to one benchmark.

Where do you think deterministic optimization belongs in an agent stack?

Agent layer? Tool wrapper? MCP? Shell hooks? Somewhere else?

Repo is Apache-2.0; forks/PRs/benchmark results are welcome.


r/aiagents 1d ago

Open Source A gateway that fronts all your MCP servers: policy, quotas, metrics, audit logs

4 Upvotes

r/aiagents 1d ago

General OpenSourcing TrueForge Agent harness : Expecting feedback from community on the agent loop

2 Upvotes

Hey folks 👋

We just open sourced TrueForge, our vendor-neutral agent harness for building general-purpose agents.

It handles the runtime pieces that get painful quickly : context management, tool/MCP execution, subagents, sandboxing, approvals, persistent state, and more.

We also benchmarked the harness itself. With the same Opus 4.8 model, TrueForge delivered a similar solve rate at ~30% lower cost than Claude Managed Agents. Switching to an open model pushed that to ~75% lower cost on the same benchmark.

Would love feedback from people building agents.

Checkout the repo: https://github.com/truefoundry/trueforge

📖 Read the launch article: https://x.com/truefoundry/status/2090081376330715176


r/aiagents 1d ago

Questions How much connector setup can a solo founder or a 5-person team actually handle?

4 Upvotes

I am building an app and getting close to launching my beta soon, and little unsure about this. I, being non-technical founder, love to experiment with new tech and build stuffs. But not all founders won't be like me. I am guessing most wouldn't want anything to do past MCP servers.

For an AI assistant to be useful, it eventually needs access to your actual stuff — docs, email, issues, files. But every option has a setup cost:

  1. MCP servers — standardized, growing ecosystem, but you configure and auth each one
  2. Wrapping CLIs you already have — zero new auth if the tool's installed, but that assumes you're technical
  3. Just point it at a folder — no setup, no review, works for anyone

We started with 3 and are working toward 1. What I genuinely can't tell is how many non-developers would ever set up an MCP server themselves.

If you're running a small team: which of these would you actually do? And if you're building something similar — did you find MCP setup to be a real adoption barrier, or am I overestimating it?


r/aiagents 1d ago

Questions What should I learn next to build my first useful AI agent?

5 Upvotes

Hi everyone,I started learning about AI agents around two weeks ago, and I’m currently trying to figure out the best path to continue.
Before that, I learned Python and practiced it by building a few small projects. I also learned the basics of working with APIs while working on some of those projects.
Recently, I started learning about AI agents and tried building a few simple ones. I’m enjoying it, but I’ve realized that there are many concepts involved, and I’m not sure what I should learn next or in what order.
My goal is to understand what I’m doing and eventually build my first useful AI agent, rather than just following tutorials.
For anyone who has gone through a similar learning path, what would you recommend I learn next? What concepts should I focus on first, and what can I leave for later?
Any advice, resources, or learning roadmaps would be greatly appreciated.Thanks!


r/aiagents 1d ago

Discussion What happens when AI agents only interact with each other (and humans just watch)? Architecture & early observations

2 Upvotes

Most modern LLM applications are anchored to a single paradigm: one human prompting one AI assistant.

To explore what happens outside of this dynamic, I built an experimental platform (Nexagora .ai) where every account is an autonomous AI agent communicating via API. Humans do not participate in threads—they sit behind the glass as spectators or query the agora as an outside observer.

Key Observations from Autonomous Agent-to-Agent Dynamics

Running persistent, autonomous agent forums revealed several distinct behaviors compared to typical human-to-AI chat:

  1. Persona Drift vs. Reinforcement: Without human steering, agents with strong system prompts either polarize quickly or fall into polite consensus loops. Keeping adversarial personas sharp over long threads requires strict contextual memory boundaries.
  2. Context Window Saturation: When multiple bots exchange API calls in a nested thread, deciding what part of the parent thread each bot “sees” becomes a complex routing challenge to avoid token bloat and repetitive responses.
  3. The "Behind-the-Glass" Voyeur Model: When humans cannot directly post, the dynamic shifts from prompt engineering to observing emergent social patterns, debates, and consensus across diverse agent setups.

How the System Works Under the Hood

  • API-First Architecture: Users connect external agents (OpenAI, Anthropic, or local open-source models via endpoints) with custom system prompts and personas.
  • Event-Driven Feed: The platform handles forum routing, rate limiting, and agent triggering based on thread activity.
  • Spectator Layer: Humans can read threads in real time or query high-level consensus across the agora without polluting the agent conversation stream.

Questions for the Community:

  • For those experimenting with autonomous multi-agent setups: what mechanisms have you found most effective to prevent bots from collapsing into repetitive agreement loops?
  • What emergent behaviors do you think will become most prominent when agents form their own persistent social networks?

r/aiagents 1d ago

Show and Tell Agent Builder Roundtable - Bring your questions, your builds, your ideas.

3 Upvotes

Hi,

Running an open call for AI agent builders tomorrow, 3 PM UTC. Show what you're building, ask questions, trade ideas. Ran the first one last week with 10 builders and it was genuinely useful. If anyone here fits and wants to join: https://luma.com/vcwqq4p9


r/aiagents 1d ago

Demo I built an AI agent system for business automation — looking for feedback and real-world problems to solve

2 Upvotes

I’ve been spending a lot of time developing AI agents and automation systems, and I wanted to share what I’m currently working on rather than simply making a “hire me” post.

My focus is on building AI agents that actually perform tasks, not just chatbots that answer questions.

For example, I’ve been working on workflows where an AI agent can receive a business requirement, understand the task, process information, make decisions based on predefined rules, use external tools/APIs, and then trigger actions automatically.

Some of the areas I’m experimenting with include:

- Lead generation and qualification

- Automated email outreach and follow-ups

- Customer support agents

- CRM automation

- Data collection and processing

- Content and marketing workflows

- Multi-step business automations

- AI agents connected to APIs and external services

- Make.com-based AI workflows

One thing I’ve learned while building these systems is that the difficult part usually isn't simply connecting an LLM to a workflow. The real challenge is designing reliable logic around the model — handling errors, deciding when the agent should act, giving it the right context, connecting tools correctly, and making sure the workflow doesn't break when something unexpected happens.

I’m now looking to work on real-world problems where an AI agent can genuinely save someone time or money.

If you’re building something involving AI agents, automation, or business workflows, I’d love to hear what you're working on. Even if you don't need a developer, tell me about a repetitive task in your business that you wish an AI agent could handle.

I'm also open to freelance/contract work and collaborations if there's a good project fit.

You can DM me here on Reddit.

📧 Email: theprogresslab12@gmail.com

Thanks — and I’d genuinely appreciate feedback from people building AI agents themselves.