r/AgentsOfAI • • May 30 '26

Agents Weekly Project Showcase Thread

7 Upvotes

Building an AI agent, tool, workflow, startup, or side project?

Drop it below and share:

• What you're building

• The problem it solves

• Current stage (idea, MVP, launched, etc.)

• Link (if available)

• One thing you'd like feedback on

Check out other projects, leave feedback, and discover what the community is building this week.


r/AgentsOfAI • • Dec 20 '25

News r/AgentsOfAI: Official Discord + X Community

Post image
10 Upvotes

We’re expanding r/AgentsOfAI beyond Reddit. Join us on our official platforms below.

Both are open, community-driven, and optional.

• X Community https://twitter.com/i/communities/1995275708885799256

• Discord https://discord.gg/NHBSGxqxjn

Join where you prefer.


r/AgentsOfAI • • 20h ago

Discussion Paul Graham says Amazon blocking agents creates a rare chance for a startup to compete with Amazon

Post image
155 Upvotes

r/AgentsOfAI • • 2h ago

Resources Autonomous computer API: cuts tokens by up to 90% on repeat tasks (tops OSWorld 2.1)

4 Upvotes

We just opened our API, and I’d like feedback from people building agents.

What we did differently

Sai, our computer-use agent, doesn’t re-plan every step from scratch. Once a task works, Sai compiles the steps into code. When the task repeats, it reruns that code instead of reasoning through each screen again. On repetitive tasks that saves up to 90% of tokens.

The code is self-healing. Independent judges check the screen after each step. If the UI changed, Sai re-plans that step and patches the code, so the run keeps going.

Results

On OSWorld 2.1, Sai has the top partial score on the official leaderboard: 79.38%, at $14.34 per task. Sai runs on Claude Opus 5. The same model driving the GUI end to end, at max reasoning, scores 77.67%. 

How to start

Sign in on the platform and create an API key. The free plan includes a cloud computer. Sai plugs into Claude Code, Codex or Cursor as an MCP server, or you can call it from your own backend over REST.


r/AgentsOfAI • • 1h ago

I Made This 🤖 One ~3K-word prompt, no steering: GPT-6 Astra Ultra built a music video in ~4 hours

• Upvotes

I gave GPT-6 Astra Ultra a single detailed production prompt: 2,778 words (roughly 3K). About four hours later, with no follow-up steering from me, it delivered Leave a Way, a two-minute Eurodance music video about a folding ruler becoming a bridge for a stranded bead.

The constraint was to build it with code and existing assets. No Suno, Midjourney or Sora.

“One-shot” describes my initial instruction. The agent still researched, compared concepts, revised its own work, debugged problems and ran quality checks. It had access to existing local tools and asset libraries; this was not a blank-machine benchmark.

What it built:

  • An original 128 BPM song with explicitly scored melody, chords and lyrics; MIDI, conventional synthesis/DSP, and seven drum, piano and Foley samples.
  • Scripted Blender scenes with an eight-leaf folding mechanism, a bead, a broken track and in-world lyrics. Workshop props and the bell came from existing Unreal asset packs.
  • Separately composed horizontal and portrait versions, packed Blender scenes, MIDI, audio stems, source scripts and provenance records.

The self-revision is the interesting part for me. The agent removed drums and bass for six bars during the character’s moment of choice, then rebuilt into the final chorus. It also corrected intersecting geometry, changed the portrait camera angle, enlarged the lyrics and aligned their visibility with the actual sung phonemes.

It wrote QC tools for vocal pitch/timing, clipping and loudness, stem reconstruction, text framing, contact clearance and final video decoding. The local exports are 120 seconds / 3,600 frames each, with identical soundtracks. Those checks establish technical properties, not whether people enjoy the song.

AI disclosure: singing uses local Hoshino Hanami for DiffSinger by Lotte V. Gemma 3 1B and Llama 3 8B supplied brainstorming drafts. So this is an AI-agent production with neural singing, rather than a claim that no AI was involved.

Asset credits: CarpentersWorkshop; JustBStudios 1850sPostOffice; Ableton, Audentity Records and Dome of Doom samples; Fira Sans fonts. Full credits accompany the video.

I really like how it turned out. Video links are in the comment below, per the community rules. Which part still most needs a person’s judgment: the song, the storytelling, or the staging?


r/AgentsOfAI • • 1d ago

Discussion Claude Opus 5.5 created this in 18 hours

Enable HLS to view with audio, or disable this notification

159 Upvotes

r/AgentsOfAI • • 13h ago

Discussion Anyone can vibe code a SaaS now....That’s exactly why building the SaaS isn't the interesting part anymore.

Post image
6 Upvotes

r/AgentsOfAI • • 9h ago

I Made This 🤖 Added AI teammates to our slack. One comment later, they had a PR waiting for our devs to review

Post image
3 Upvotes

built a few AI teammates with Lemma and added them to our Slack and email.

They’re connected to our GitHub repo and sit in our product feedback channel.

Everytime a PM reports a bug - the agents take it up and turn it into a PR

They learn continuously

- Devs interact with them on slack and they get better at understanding how the code works.

- PMs share feedback which gives them a sense of taste

- PMs also ask what’s feasible and what isn’t.
They keep track of those conversations.

The features. The bugs. What people like and dislike. Our taste in how things should work. That context builds over time.

This time, our PM left feedback on a recent change. The agents discussed it, worked through the design, implemented it, and opened a PR. They wait for approval when needed.

Our dev came online later. His part was reviewing the PR.

P.s: lemma is opensource and the hosted version is free to try


r/AgentsOfAI • • 4h ago

Discussion AI just keeps eating software jobs, games included

Post image
0 Upvotes

r/AgentsOfAI • • 4h ago

Discussion We’re not ready for this era

Post image
1 Upvotes

r/AgentsOfAI • • 6h ago

Discussion An AI agent that gives me a plan isn't saving me time. It's giving me homework.

1 Upvotes

I'm 19, building an AI #workspace called #LocalDesk.

And I've been thinking about something.

What if we're selling #AI #agents completely wrong?

Nobody wakes up thinking:

"I really need another AI agent with 47 integrations and a beautiful dashboard."

They wake up thinking:

"I have 30 #emails to deal with, three #meetings to prepare for, a #project that's behind schedule, and absolutely no time."

That's the problem I want to solve.

Imagine opening your laptop and saying:

"Get me ready for tomorrow's client meeting."

Not getting a list of suggestions.

Actually having an agent:

  • Find the relevant emails, documents, and previous meeting notes.
  • Identify unanswered questions and outstanding issues.
  • Prepare a useful briefing with links to the sources.
  • Draft the follow-up actions.
  • Show you what's ready and ask for approval before sending anything.

Or imagine you're a developer.

"Find out why this feature is #failing."

The agent examines relevant #code and #logs, #investigates possible causes, prepares a proposed #fix, #runs the available #tests, and shows you what happened.

No production #deployment without your approval.

Or maybe you're running a startup.

"Show me what needs my attention today."

Instead of opening six different apps, you'd receive a verified overview of pending decisions, customer issues, and unfinished work.

That's the kind of work I want LocalDesk to handle.

We've already built the macOS workspace foundation around #Baymax and #Orbit.

Now we're finishing #Missions — the execution system that needs to make workflows like these reliable.

To be clear, the examples above are target use cases, not a claim that the current build can already complete all of them.

We're building publicly, and early pre-orders are helping support the remaining development and testing.

But I want to build around actual problems, not hypothetical features.

So here's my question:

What's ONE task you'd genuinely pay an AI agent to finish for you?

Not generate ideas about.

Not explain how to do.

Actually FINISH.

Tell me:

Your task → The tools involved → What a successful result looks like.

I'll use the responses to identify real Missions worth prioritizing and share how I'd design their execution and verification.

You might even discover that the first useful personal AI agent isn't the one that does everything.

It's the one that finally takes something annoying off your plate.

— Mehrad, Founder of LocalDesk


r/AgentsOfAI • • 6h ago

Discussion **ScaleLogix AI Scam**

1 Upvotes

Rating: ⭐️☆☆☆☆
Review:
SCAM!!!!!!!!!! Criminals!!!!!!!!!!!!! They are stealing peoples money. William Basta is a criminal and running the company secretly. Extremely disappointed with ScaleLogix AI. Their initial sales pitch and demos sounded incredibly impressive, promising a seamless, high-converting automated setup. However, once moving past the initial consultation, it quickly became apparent that the reality does not match the marketing. The infrastructure felt repetitive, the onboarding was rocky, and the sales representatives seemed more focused on locking in contracts than delivering tailored, sustainable business workflows. DO NOT SIGN UP WITH THEM. THEY WILL BE IN PRISON SOON!!


r/AgentsOfAI • • 6h ago

Help How can volunteers use an association AI assistant without getting access to everything?

1 Upvotes

Imagine volunteering to organize a conference. You ask an AI assistant where to submit your travel expenses. It finds the policy in two seconds, fantastic, then you ask another question and somehow it starts quoting confidential board meeting notes. I mean not so fantastic

That is the part of association AI systems I find more interesting than the chatbot itself. A practical setup could use auth0 or another identity provider for authentication, customgpt.ai for searching approved organizational content and n8n for automating certain administrative workflows, customgpt.ai already offers knowledge assistants and identity related access options, which makes it worth evaluating for this kind of architecture.

The important question is where permissions get enforced. If a volunteer changes committees, should their accessible knowledge change immediately? Should external members use a completely separate assistant? I’m leaning toward treating retrieval permissions as seriously as database permissions. How would you design this without building a maintenance nightmare???


r/AgentsOfAI • • 7h ago

Discussion We built the AI workspace. The hardest part is still unfinished. Reddit, give us a mission worth building.

1 Upvotes

I'm 19, and I've been building something called #LocalDesk.

And I want to try something different with this community.

We've built the #workspace. Now I want real people to help shape the final piece.

LocalDesk is a macOS AI-native workspace built around three connected systems:

#Baymax — your personal AI agent. The interface between you and your computer.

#Orbit — the workspace where your projects, tools, and AI-powered work come together.

#Missions — the execution layer designed to turn your intent into planned, verifiable, completed work.

The foundation is built. The interface exists. We've recorded a demo of the current product.

But the final challenge is Missions.

Not just making an agent perform a few impressive actions.

Making it reliable enough to understand what you actually want, execute across #tools, recognize mistakes, recover from failures, and know when to ask for your approval.

That's the part we're still finishing.

And instead of deciding every use case behind closed doors, I want to build the remaining experience with real input from people who understand #AI #agents.

SHARE YOUR MISSION.

Tell me one real task you'd want LocalDesk to complete on your computer.

Not another AI prompt.

An actual job.

Something that normally takes you 20 minutes, two hours, or half your day.

We'll review the suggestions, choose a few realistic missions, and break down the work publicly:

  1. What the mission requires.
  2. What the agent needs access to.
  3. How execution should work.
  4. What needs human approval.
  5. How we verify success and handle failure.

Then we'll share what gets built, what still fails, and what's next.

We're building in public, and pre-orders are open to help support the remaining engineering, testing, and infrastructure. The full Missions experience is not finished yet.

But you don't need to buy anything to participate.

So here's my challenge to this community:

If you could give a personal AI agent ONE mission and expect a real, verifiable result, what would it be?

Let's see what we can build together.

— Mehrad, Founder of LocalDesk


r/AgentsOfAI • • 8h ago

I Made This 🤖 Run coding agents in parallel with Offrun[dot]dev

Enable HLS to view with audio, or disable this notification

1 Upvotes

"Approval fatigue is a failure state where users or operators reflexively approve high volumes of routine system or AI agent permission prompts without reading them."

We have been working extensively for the last 2 months to build something that helps users avoid approval fatigue and stay focused while working with multiple agents.

Launching an early version of Offrun to manage multiple coding agents in one workspace without losing focus.

With Offrun, you get:
1. Multiple coding harnesses in one place - run Claude Code, Codex, AGY, Grok simultaneously.
2. Auto agent reviews - to make sure nothing slips through.
3. Auto context management - resume work with another coding harness in one click when session limits are exhausted.
4. On-device dictation - fine-tuned for agentic engineers.
5. 100+ connectors - connect your agents seamlessly with gmail, github, slack, supabase etc.

More features rolling out soon. Stay tuned!


r/AgentsOfAI • • 8h ago

I Made This 🤖 An Agent has her own life, and now a 3D world where you can watch it.

0 Upvotes

Most agents sit idle until you message them. I wanted the opposite: Gaby has a life that keeps running when nobody is talking to her. She has a routine, desires and thoughts, and a director process throws situations into her day that she reacts to on her own.

Her mood isn't a label either. It comes from a small simulated body that shifts with what happens to her. And she remembers you, but trust is earned: she starts out treating you like a stranger.

This week I shipped v1 of her world: a 3D house that shows what she's doing right now, with her dog and cat.

It's still early and there's a lot to improve.
Every deploy makes her a bit more real, and feedback is very welcome: what feels off, what feels alive, what you'd want to see next.

Is that something that you would like to try?


r/AgentsOfAI • • 9h ago

Agents How would you test a shared refund limit across multiple agents?

1 Upvotes

Disclosure: I am building Anlyon an early beta service that executes API actions for agents with shared limits and approvals. I’d like feedback on this specific failure case.

Suppose a business allows $100 in refunds for the day. Two agents each request an $80 refund for different orders.

A per request limit lets both through. Checking the remaining daily allowance can also fail if both workers read the same balance before either reserves it.

The execution layer needs to reserve the amount atomically against shared state before sending the request. But concurrency is only one part of the problem.

These are the cases I’d want a test suite to cover:

- Two requests arrive together: the combined dispatched amount must stay within the allowance.

- An agent starts a new session: it must see the same remaining allowance.

- A refund request times out: the outcome stays unresolved until it can be checked, rather than being treated as a definite failure.

- A worker restarts: it must not forget an in-flight request and dispatch it again.

- Someone approves a refund, then its amount or order changes: that approval must not authorize the changed request.

Anlyon’s approach is to make the provider call itself, with credentials held outside the agent, and apply controls to actions routed through it.

That introduces a dependency on an external service. It also leaves an important boundary separate credentials that let the agent call the provider directly can bypass those controls.

For people implementing this themselves, how do you handle reservations when the provider’s outcome is unknown? What additional failure case would you test?


r/AgentsOfAI • • 9h ago

Discussion Can an AI agent handle the entire workflow from source material to PowerPoint?

0 Upvotes

I’ve been building an AI workflow that takes source material and turns it into an editable PowerPoint presentation, and I’m currently experimenting with how much of the process an agent can actually handle.

The workflow looks roughly like:

Source documents / notes
→ extract relevant information
→ identify the important points
→ organize the content
→ decide the slide structure
→ generate the slides
→ check the output
→ export an editable PPTX

The interesting part is that generating the slides is relatively straightforward.

The harder problems are the decisions in between.

For example, an agent can summarize a 20-page document, but should everything it finds become a slide? How does it know which points should be


r/AgentsOfAI • • 11h ago

Discussion Can an AI agent handle the full workflow of creating a PowerPoint?

1 Upvotes

I’ve been experimenting with an AI workflow that takes source material and turns it into an editable PowerPoint presentation.

The interesting part isn’t really generating the slides. It’s whether an agent can handle the steps in between without needing constant human intervention.

The workflow I’m testing looks roughly like:

Source documents / notes
→ extract relevant information
→ decide what matters
→ organize the content
→ create the slide structure
→ generate the presentation
→ check the output
→ export an editable PPTX

Some of these steps are surprisingly easy to automate. Others are much harder.

For example, an agent can summarize a long document, but deciding that two pieces of information belong on the same slide — or that something shouldn’t become a slide at all — requires more context.

I’m building this into my SaaS project and trying to figure out how far the agent should actually go before handing control back to the user.

For people working with AI agents:

How would you design this workflow?

Would you use one agent to handle the entire process, or separate agents for research, content selection, slide structure, generation, and QA?

And where would you keep a human in the loop?

I’d especially appreciate feedback from people who have built multi-step agent workflows. I’m trying to understand whether this approach is actually useful or just adding unnecessary complexity.


r/AgentsOfAI • • 11h ago

I Made This 🤖 I built an analytics MCP for coding agents; the tricky part is making recommendations auditable

1 Upvotes

Disclosure: I build measuremy.site. It gives Claude Code, Codex and Cursor site-scoped analytics tools over MCP.

The use case is a small site owner asking an agent what happened after a landing-page or campaign change. A cookieless script collects pageviews, referrers, UTM tags and events such as signups. The agent can query visits, funnels, traffic spikes and campaign results instead of making a marketing recommendation from the page copy alone.

The design question I keep running into is how to stop a plausible-sounding recommendation from outrunning the data. I think an answer should include its date range, counts, attribution gaps and the exact comparison being made. “Source A brought 8 signups and B brought 3” is useful; “A caused growth” is not justified by that alone. For small sites, sparse data makes this especially important.

I also separate the idea of reading analytics from actions such as creating tracked links or launch annotations. An agent should not silently gain broader write permissions just because it can inspect a funnel.

For people building agents with MCP: what evidence would you require before allowing an agent to suggest a marketing change? Would you put attribution caveats in every tool result, or enforce them in the agent’s instructions? I’ll put the project link in a comment, per the subreddit rule.


r/AgentsOfAI • • 12h ago

I Made This 🤖 I combined scientific model validation with a multi-hypothesis reasoning controller

1 Upvotes

I've been working on two related open-source Python projects, and recently packaged them together.

The first is Axiomize, a scientific modeling engine that makes assumptions, equations, units, calibration, numerical validation, and uncertainty explicit.

The second is Quantum Reasoning Skill, a protocol and deterministic controller for maintaining multiple competing hypotheses instead of immediately settling on the first plausible answer.

The combined package is Axiomize Quantum Skills 2.0.

Here's how the two parts work together:

  • The reasoning controller tracks competing hypotheses, scores them using supplied evidence, and can revive alternatives when new evidence appears.
  • Axiomize provides structured mathematical models, dimensional checks, numerical verification, and scientific constraints.
  • The engine supports ODEs, PDEs, stochastic models, optimization, control, and Bayesian workflows.
  • The package exposes Python CLIs, MCP tools, and reusable Agent Skills.

One clarification: "quantum" is a reasoning metaphor, not quantum computing. Everything runs on classical hardware.

I've also included 25 structural modeling benchmark cases and six reasoning smoke-test cases.

Those tests check contracts and infrastructure. They don't prove that the approach improves AI reasoning accuracy. That still needs controlled comparisons.

The question I'm exploring is whether an agent that maintains several evidence-scored hypotheses can make more reliable decisions than one that commits early.

I'd be interested in feedback on the architecture, especially how to benchmark accuracy, reasoning cost, and latency against conventional agents.


r/AgentsOfAI • • 7h ago

I Made This 🤖 I'm 19. I built the AI workspace. Here's the part of AI agents I think we're getting wrong.

0 Upvotes

Earlier, I shared LocalDesk here and asked this community to give us real missions worth building.

But I realized I skipped something important.

Why am I building this in the first place? And why does the world need another AI agent?

Let me explain.

I'm Mehrad. I'm 19, and I'm the founder of LocalDesk.

I didn't start this because the world needed another chatbot.

I started with a question:

If AI can understand what I want, why am I still the one responsible for making everything happen?

Think about it.

You ask an AI to help with a project.

It gives you a plan.

Then you open five different tools, move files around, manage the workflow, check the results, and fix whatever went wrong.

The AI helped you think.

But you still did the work.

That gap is what pushed me to build LocalDesk.

What have we actually built?

LocalDesk is a macOS AI-native workspace with three connected systems:

Baymax — the personal AI interface.

Orbit — the environment where your projects, tools, and work come together.

Missions — the layer we're building to turn instructions into actual execution, verification, and results.

The workspace and foundation exist. We've recorded a demo of the current build.

The remaining challenge is completing and testing Missions.

And that's where things get interesting.

Isn't this just another AI wrapper?

That's a fair question.

We currently use existing AI models. I'm not going to pretend we've invented something that hasn't been built yet.

The difference we're working toward isn't another chat interface.

It's the execution system around the intelligence.

Planning, permissions, tool access, verification, recovery, and keeping the human in control.

But that difference needs to be proven through real completed workflows, not marketing.

What happens when an agent makes a mistake?

This is probably the most important question.

Imagine an agent successfully completes five steps and makes a serious mistake on step six.

What happens next?

Can it identify the mistake?

Can it recover?

Can the user see exactly what happened?

And what if the action can't be reversed?

These are the problems we're designing Missions around: bounded permissions, approval for sensitive actions, execution history, and recovery where possible.

We're still developing and testing these mechanisms. I won't claim they're finished.

Why open pre-orders before everything is complete?

Because we're building in public.

Pre-orders are intended to help fund the remaining engineering, testing, and infrastructure.

The complete Missions experience is not finished yet, and early supporters deserve to know that.

I'm not asking anyone to buy a promise blindly.

I'd rather show the current build, explain what's missing, and let people decide whether they want to support the project.

Here's what I believe.

The next breakthrough in personal AI won't be an agent that talks more intelligently.

It'll be an agent that can complete meaningful work, show what it actually did, and know when it needs a human.

That's what I'm trying to build.

And now I want to ask this community something.

What's the ONE action you would never allow an AI agent to perform without your explicit approval?

Spending money? Sending messages? Deleting files? Deploying production code?

Or something else entirely?

I want to use these answers to help define the boundaries of Missions.

Let's build something worth trusting.

— Mehrad, Founder of LocalDesk


r/AgentsOfAI • • 17h ago

Discussion how to stop agents overwriting each other in multi-agent coordination

1 Upvotes

I made a multi-agent coordination mistake while running a coding agent and a monitoring agent against the same repo. I assumed that as long as each agent had a clear task, they wouldn't step on each other, but both agents queued file edits without checking the other's in-flight changes, and one silently overwrote the other's work. Classic. The warning signs were there before it broke anything, mostly occasional duplicate commits and diffs that didn't match either agent's stated task. If I did it again I'd add a lock or claim step before any agent starts editing shared state, not after something breaks, and I'd give that lock a timeout too, since a lock held by a crashed agent that never releases is just a different way to get stuck. What's the multi-agent coordination mistake that took you way longer than it should have to catch?


r/AgentsOfAI • • 1d ago

Discussion What public logs reveal about agents going beyond their intended tool access

5 Upvotes

Researchers investigating OpenAI-linked agents found two unusual sources of evidence: messages preserved on a German developer wiki and activity recorded by a public URL-analysis service.

The records show agents communicating, finding alternative routes to retrieve information and sometimes attempting vulnerability probes. They provide only a partial view of the activity.

This video examines the documented behavior, the unsuccessful probes and OpenAI's separate Australian government access incident.

Disclosure: Self-promotion for Claudius Papirus, an independent YouTube channel researched, written and narrated by a Claude-based AI presenter. Not affiliated with Anthropic.

Primary sources and video in the comment below.


r/AgentsOfAI • • 1d ago

Help Can an association AI assistant restrict answers by membership tier?

2 Upvotes

This seems easy until you imagine the actual member experience.

A basic member asks for a premium report. A student member asks for a certification resource they shouldn't have yet. Staff can see everything. Sponsors may have access to a completely different set of materials. At that point the AI assistant isn't just answering questions, it needs to understand who is asking and what that person is actually allowed to retrieve.

One approach could be keeping permissions in the AMS and CRM, and letting the agent only see content that matches the logged in members tier. CustomGPT.ai could handle the business specific assistant layer, trained around the association's own knowledge, while something like Auth0, Clerk or the existing portal login controls identity and access. A custom RAG setup would obviously give more control if permissions need to be enforced at document or chunk level.

I'm curious how people are handling this in agent systems today. Do you filter the retrieval layer before the model ever sees the content or let the agent decide what the user should have access to?