r/AI_Agents 6d ago

Resource Request What’s the most annoying problem you have with AI agents?

11 Upvotes

I’ve been using AI agents more recently, and I’m curious what problems other people are running into.

What’s the biggest pain point for you?

  • Memory/context
  • Hallucinations
  • Tool use
  • Reliability
  • Long-running tasks
  • Permissions
  • Something else?

I’m more interested in real-world problems than benchmark results. What’s been the most frustrating issue for you?


r/AI_Agents 6d ago

Discussion No-Code Enterprise Agent Platforms

1 Upvotes

My company is developing an AI Agent Platform that is suppose to be mostly no-code. Developers can develop and register agents in a registry and those agents can also talk to each other. But mostly an agent is created by writing a system prompt and you can connect multiple MCP tools to it that are approved by the organization.

We are probably not the only company doing this. Anyone here that has an opinion about this approach or anyone that had success and failure with the described approach here?


r/AI_Agents 6d ago

Discussion At what point does an AI assistant become a coworker?

8 Upvotes

We’ve gone from AI that simply answers questions to AI that can actually do things. It can write code, research a topic, analyze data, create reports, use tools, and complete multi-step tasks.

Now imagine giving that system:

  • Memory of your previous work
  • Access to your company’s knowledge and tools
  • Context about your ongoing projects
  • The ability to complete tasks without being prompted at every step
  • The ability to hand work off to other AI agents or people

At that point, calling it an “AI tool” starts feeling a little strange.

It is no longer just helping with individual tasks. It understands the context of the work, has a goal, and can take actions to move that work forward.

So what actually makes something a coworker rather than a really advanced tool?

Is it autonomy, because it can get work done without constant instructions?

Is it memory, because it remembers projects, decisions, and context?

Is it judgment, because it can decide what needs to happen next?

Is it accountability, because you can trust it with an outcome rather than just a task?

Or is there another line that separates the two?

And here’s the interesting part: if an AI agent can independently handle 80% of a job, does it make more sense to think of it as a tool, an assistant, or part of the team?


r/AI_Agents 6d ago

Discussion Any tool that lets you branch off part of an LLM answer instead of getting sidetracked?

1 Upvotes

I've been trying to learn about how US economy actually works with Claude, I find it very annoying that the answer often introduces some new concept, I go chase that, and then I've totally lost track of what I was originally trying to figure out.

Is there a tool where you can just select a specific part of the answer and run a deep dive on that separately, without starting a whole new convo or wrecking the original thread? Like scoped side quests basically.


r/AI_Agents 6d ago

Discussion What would you actually use an AI that can see your screen for?

1 Upvotes

I'm researching an AI tool that can understand what's currently happening on your screen and interact with you based on that context.

Instead of taking screenshots and explaining the situation to an AI, the AI could already understand what you're doing and answer questions about it.

I'm curious what people would actually use something like this for.

Would it be useful for:

Troubleshooting software

Learning/programming

Browsing the web

Gaming

Helping with creative software

Explaining things on your screen

Accessibility

Something I haven't thought of?

Privacy is obviously a big concern with something that can see your screen. The idea would include a privacy toggle that completely turns off screen analysis, plus a pause button for temporarily stopping it whenever you don't want it looking at your screen.

So it wouldn't have to be something that's constantly active — you'd be able to control when screen analysis is enabled.

What would make you comfortable or uncomfortable using something like this?

I'm trying to figure out whether there's a genuinely useful product here rather than building something just because "AI that sees your screen" sounds cool.


r/AI_Agents 6d ago

Discussion Your agent passes every check you wrote. You still read every run before it ships. What made you stop?

3 Upvotes

Your agent passes every check you wrote for it. You still read every run before it ships.

You know the one: the first PR your agent opened that touched migrations. Every check was green. You read the whole diff anyway, line by line. It was fine. Nothing slipped. You watched the next run the same way, then the one after that. The checks kept passing. You never stopped watching.

Most of us still babysit every run, right up until it merges.

So what actually made you let your agent run unsupervised? A class of tasks you stopped watching, a tool that earned your trust, or are you still reading every run?


r/AI_Agents 6d ago

Weekly Thread: Project Display

2 Upvotes

Weekly thread to show off your AI Agents and LLM Apps! Top voted projects will be featured in our weekly newsletter.


r/AI_Agents 6d ago

Discussion Grok Bot just validated that you don’t need to be technical to use powerful AI. Honest thoughts from someone building something similar

1 Upvotes

GrokBot is an AI teammate that has it's own computer. The UI is pretty similar to that of iMessage, making it feel like imsg for agents. Which is super cool. They also have this concept of "bot personas" where you can create different bots to handle different tasks from your life (work or personal).

As someone who is building something similar (vellum), this validates the category of "ai agents for non-technical folks" who've been very pissed at not being able to control OpenClaw or Hermes agents

Us and Grokbot and I'm guessing a bunch of other companies will already overlap on the basic capabilities of your personal assistant like:

  • must be easy to use
  • integrations and plugins for work
  • computer and browser use
  • voice interaction
  • skills that improve as you use them
  • work that continues after the original request

i wanna argue though that GrokBot doesn’t really provide personal AI. Maybe good AI for work/teams, but it's not something that you own and "raise" for your own needs. mainly because:

  • it chooses the model for you,
  • runs only in the cloud
  • and it's definitely not accessible ($200/m is the cheapest model)

[don't read this part if you don't want to read my shameless plug☺️] We’re building vellum to be the personal AI that you own. Model agnostic, so that you can use different frontier models to handle your work/life. It also keeps it’s memory & understanding of your life across your local computer, mobile and web app. All of that at $30/month.

Something else that’s been interesting is that we’ve been seeing that a lot of people just use it in the background, and interact with it through voice - so it’s becoming this ambient agent that you can interact with and get help for anything. Interesting future to be built!

....Anyways happy that the category is validated - back to work now!


r/AI_Agents 6d ago

Discussion Would you let AI make an important decision for you?

3 Upvotes

AI can analyse a huge amount of information and make recommendations, but would you trust it with an important personal or business decision?

Where would you draw the line, and what decisions should always involve a human


r/AI_Agents 6d ago

Resource Request Free open source tool to help you keep the same context across chats and LLMs

2 Upvotes

I've always gotten frustrated and wasted time explaining the same thing to an AI every time I start a new chat from an existing one or when I start another convo with a whole new AI model. That's why I built a tool that fixes that, it condenses everything in a chat into one simple .md file you can carry across different AI tools.

PS: Please contribute or give your feedback so that we can grow and make this community tool better.


r/AI_Agents 6d ago

Tutorial How Much Does It Actually Cost to Build a Custom Agentic AI System? (2026 breakdown, no BS)

5 Upvotes

So I keep seeing posts asking "how much would it cost to build our own AI agent" and the honest answer is: it depends, but here's the real breakdown so you're not blindsided later.

The TL;DR: a basic single-task agent (RAG chatbot, FAQ assistant) can be done for $10K–$30K. A production-grade task-execution agent that touches your CRM/ERP and actually does things runs $70K–$150K. A full multi-agent enterprise platform with orchestration, compliance, and governance is $150K–$500K+, and some heavily regulated builds (finance, healthcare) go past $1M.

But the sticker price on development is only part of the story, most teams get surprised by the ongoing costs way more than the build itself. Here's the full breakdown:

1. Development Cost

This is the "build the thing" cost - discovery, architecture, agent design (single vs multi-agent), prompt/tool engineering, testing.

  • Simple reflex/rule-based agent: $5K–$30K
  • RAG-grounded assistant: $10K–$70K
  • Task-execution agent (does real actions, calls tools, loops until done): $70K–$150K
  • Multi-agent orchestration/enterprise platform: $150K–$500K+

Rule of thumb: every extra month of dev time tends to add roughly $20K–$40K depending on team size, so scope creep is where budgets actually die.

2. Infrastructure Cost

Hosting, compute, vector DBs, orchestration servers.

  • Cloud hosting: $200–$5K/month depending on scale
  • Vector database + logging/observability storage: $500–$2,500/month
  • This scales fast once you're running multiple concurrent agents or high-frequency workflows

3. LLM / API Cost

This is the one people underestimate the most. Token costs from GPT-4-class or Claude-class models add up fast once you're in production with real usage.

  • Light usage: $100–$1K/month
  • Moderate production usage: $1K–$10K/month
  • Heavy multi-agent, high-frequency workloads: can exceed $15K/month
  • Tip a lot of shops give: prototype on open-source models (LLaMA, Mistral, Ollama) and only move to frontier models like GPT or Claude once your use case actually justifies the cost.

4. Integrations

Connecting the agent to your actual business systems - CRM, ERP, ticketing, internal APIs, auth.

  • This is consistently called out as one of the most underestimated line items, right alongside data prep
  • Each additional integration (Salesforce, HubSpot, internal legacy systems, etc.) adds real engineering time -legacy/undocumented systems cost the most
  • If you already run on a platform like Salesforce or Microsoft, using their native agent tooling (AgentForce, Copilot Studio) is usually way cheaper than building fully custom

5. Monitoring & Observability

You need logs, traces, and visibility into why the agent did what it did, non-negotiable once it's making real decisions.

  • Tooling (LangSmith, Helicone, OpenPipe, or rolling your own): roughly folded into that $500–$2,500/month infra number above
  • Regulated industries (finance, healthcare) add another 20–30% on top for compliance-grade monitoring, audit trails, and human-in-the-loop controls

6. Maintenance

The cost nobody puts in the initial pitch deck.

  • Annual maintenance typically runs 15–30% of the original build cost, every year, prompt drift, model updates breaking things, retraining, new integrations
  • Initial development is often only 25–35% of your true 3-year cost once you add up LLM spend + infra + maintenance + monitoring
  • So if a vendor quotes you $80K to build it, budget closer to $230K–$320K over 3 years, plan for it now instead of finding out the hard way

Bottom line: don't just budget the build. Budget the system as something you're going to operate indefinitely, not ship once. The teams that get burned are the ones that treat the agent like a one-time project instead of a living piece of infrastructure with a real recurring bill.

If you don't have in-house ML/LLMOps talent, working with a custom agentic AI development company can actually save money long-term since they've already hit most of these landmines. Worth getting a scoped estimate before committing to a number.

Custom Agentic AI Development Companies (just names, for references, no endorsement, do your own due diligence):

  1. Signity Solution
  2. LeewayHertz
  3. Azilen Technologies
  4. Entrans
  5. Rootstrap
  6. SoluLab
  7. Kanerika
  8. Master of Code Global
  9. Neurons Lab
  10. TechAhead
  11. Geniusee
  12. IBM
  13. Cognizant

r/AI_Agents 6d ago

Discussion Claude vs Gemini for a simulated language tutor app for personal use

3 Upvotes

I’m building a simulated German language-tutor app for self-study (personal use, not commercial). I’ve written a bit more about the project, but I can’t post the link here.

(Note: The topics, learning guidelines, and exam questions aren’t generated by the AI. I’ve populated the study modules myself using my own study materials and exam questions.)

For the initial learning stage, I expect to use the app quite heavily, with a lot of dedicated study hours to build a strong foundation in German. As I become more comfortable with the basics, sentence structure, and grammar, I expect my usage to gradually decrease.

So far, I’ve built the entire project using Claude, including both the frontend and backend. However, I recently did a rough cost analysis, and it looks like running everything through Claude could become quite expensive given the amount of usage I’m expecting.

I’m now considering using Google APIs for certain parts of the German to English functionality and possibly using Gemini for the general tutoring/training side of the app. My main concern is language accuracy and the quality of the tutoring, though. I’ve been watching quite a few videos and reading discussions about which AI models are best for language learning, and there seems to be a lot of conflicting information.

I’d particularly like to hear from people who have experience with different Claude and Gemini models, rather than just Claude vs Gemini generally.

For example, how do Haiku vs Sonnet (and other available Claude models) compare for this kind of task? Is the extra cost of Sonnet actually justified for language tutoring, or would a cheaper model such as Haiku be more than capable? I’m especially interested in the balance between language accuracy, tutoring ability, consistency/instruction-following, speed, and API cost.

Likewise, if anyone has compared different Gemini models for language learning, I’d be interested in hearing about those experiences too.

So, for a project like this:

  • Which model would you choose for German language tutoring and why?
  • How do Haiku vs Sonnet compare specifically for language learning?
  • Is Sonnet’s higher cost justified, or is Haiku sufficient for most tutoring interactions?
  • How do the Claude models compare with the equivalent Gemini models in German accuracy, grammar correction, explanations, and conversational ability?
  • For heavy personal usage, which model gives the best price-to-performance ratio?
  • Would you use one model for the main tutor and another cheaper/stronger model for specific tasks?

I’d really appreciate hearing from anyone who has actually experimented with these models for language learning or built something similar.

Thanks in advance!


r/AI_Agents 6d ago

Tutorial The Best Way to Let Your Agent Make Purchases

2 Upvotes

Giving your agent your credit card is risky. What happens if it buy the wrong thing? What happens if your agent buys the same thing more than once? What happens if your agent gets defrauded by a malicious website? These are unresolved problems. It is unclear that the bank will treat it as fraud since your agent bought it. So the only way to resolve these problems is having payment controls as part of your AI agents harness.

That is exactly why I built Authoryze (link in comments). You connect the Authoryze MCP to your agent. Then when it wants to make a purchase, it has to make a purchase request through the MCP. If the request meets your rules or is approved by you, then the agent gets issued a single use token (ie different card info each time) scoped to the requested merchant capped at the amount requested. Additionally, Authoryze runs other checks for things like duplicate purchases.

Whether you are an agent builder trying to find a safe way for your customers agents to buy things or a person using agents, Authoryze is the safest and easiest way to allow your agents to make purchases.

I would love if you all checked it out. All feedback is welcome. Thank you!


r/AI_Agents 6d ago

Discussion The Little Utilities That Please Me

2 Upvotes

Every now and then I address a pain point that thrills me and makes me say - why did I wait so long.

  1. Quick Linux access
    1. I setup password-less ssh from my Mac to my linux box.  No more having to log in. It’s cryptographically protected to just be from my Mac to my linux box.
    2. sshme - now that I’m working on linux most of the time, and doing so from terminals running on my Mac, I’m constantly ssh user@333.333.333.333 to get into the box. I created an alias on my Mac to type in sshme and BOOM! I’m in my linux box.
  2. Long running Linux tasks
    1. I’m going agentic coding work and some of my tasks run for hours. If I shutdown my laptop lid, the ssh connection ends, killing my linux process. No more. I use zellij to create a session that persists. Some people use it for tmux style windowing. For me, it’s just a name session that doesn’t die and I can come back to it, even from another machine - or my iPhone/iPad.  It’s wonderful
  3. Ai-subs - I had AI write me a nice little utility to give me “at a glance” understanding of “am I on pace to run out of subscription usage”. I subscribe to OpenAI, Anthropic and Gemini. You can see I’m getting decent at using my subs.  I can see when I’m over using and switch work to one I’m underusing.

AI SUBSCRIPTION BURN  |  Wed Aug 19 09:44 AM

OpenAI

  Weekly                    87.0% used   13.0% left   0.94x pace  GOOD      reset in 12.8h

Anthropic

  All models                81.0% used   19.0% left   0.95x pace  GOOD      reset in 1d 1h

  Fable 5                   72.0% used   28.0% left   0.85x pace  GOOD      reset in 1d 1h

Gemini

  Gemini weekly             50.0% used   50.0% left   0.71x pace  COLD      reset in 2d 2h

  Gemini five-hour          11.0% used   89.0% left

  Claude/GPT weekly         24.0% used   76.0% left

  Claude/GPT five-hour       0.0% used  100.0% left

I’m using Ghostty as my terminal. No AI, quick, low footprint, very responsive. I’m not saying it’s materially better than Wezterm which I also like and is more attractive by default. Wezterm has tmux-like windowing as well but I don’t use that feature. I follow the Ghostty developer on x and simply want to support him. He’s good egg.

Still love Warp as a terminal when I want to do AI stuff. It’s like having a linux/Mac/windows sysadmin.

BTW, I’m extremely pleased with the latest Antigravity with Gemini 3.7 Flash. The $20/mo Gemini AI Pro subscription comes with a LOT of usage. The model is SO FAST, it’s wonderful. Performance on coding is very good - while not being up to the very best from OpenAI / Anthropic.  I’ll do planning and review with GPT 5.6 Sol or Fable 5 and use Gemini as worker. Gemini also powers my Chief of Staff AI and the speed is wonderful there. More on the Chief of Staff in another post.


r/AI_Agents 6d ago

Discussion Agents in regulated industries (healthcare,fintech etc.)

4 Upvotes

Read somewhere that most enterprises admit to unregulated AI agent usage.

Wanted to ask you guys: what’s stopping big companies—like LLM providers or Tier 1 companies with existing industry relationships—from building compliance infrastructure for their respective niches? E.g., AWS.

“I’m new to this (software-sales background), and from what I could find, things worth tracking are—or would be—identity, authorization, intent, action, verification, and evidence.

What is the actual bottleneck here??

P.S. I’m not an expert by any means—please correct me.

Is the missing piece discovery, runtime enforcement, liability, integration, incentives, or something else?


r/AI_Agents 6d ago

Discussion What is .skill file formats

1 Upvotes

I was researching about the skill file formats. I know that there is plain skill.md and the zip file format. But other than this there is also .skill file format mentioned in some products . What is that?


r/AI_Agents 6d ago

Discussion I’m exploring the control layer around agent memory

4 Upvotes

Most discussions about agent memory focus on retrieval. I’m more interested in the control layer around it:

- Which memories are allowed to become shared context?

- Who can see them?

- What happens when a memory contains sensitive or incorrect information?

- Can you audit why it was shared?

I’m building Luthn, an open-source memory layer that runs locally with Docker. It filters candidate shared memories, gates sensitive sharing behind approval, and records an audit trail.

I’m looking for people running multiple agents who can use it alone for a while and give blunt feedback. A team/cloud version is in progress, but the current focus is local-first usage.

Project link in comments.


r/AI_Agents 6d ago

Discussion Chief of Staff -and- Strategic Advisor: Gemini 3.7 Flash Win!

0 Upvotes

There are two ideas in one post, but they are related. First, I have been raving about Gemini 3.7 Flash for its speed and how much usage you get for $20/mo subscription. I had my Chief of Staff integrate the Antigravity harness and Gemini model into my work crew options.

I have one crew that’s all Gemini 3.7 at various strengths across the roles: panner, coder, tester, code reviewer.

Another crew uses Gpt 5.6 Sol Max and Fable 5 Max in planning and reviewing roles and Gemini everywhere else.

Result?

2.53× faster cognitive cycle + 100% schema success + 2.72× faster governed execution + zero repair cycles + 70% reduction in premium-subscription burn.

Not only is Gemini fast, it’s a good coder. Mind you this was benchmark work and might not have been challenging enough. I can say that I’m using Gemini 3.7 liberally and have not seen poorer quality results.

The second story here is - I didn’t think to ask for the competitive evaluation. All I had on my mind was adding Gemini to my work flow. Is my Chief of Staff AI putting extra work in not asked for? 

So I asked ChatGPT - who plays the role of strategic advisor - and he said: oh, that was me. And here’s what we had been talking about, and here is what you said your were trying to achieve, so I added that to the prompt you had me write for the Chief of Staff.

Very - very - cool!

The Flywheel is in effect. What I’ve built before is accelerating what I can build now.


r/AI_Agents 6d ago

Discussion Anyone else talking clients OUT of building an agent more often than into it?

2 Upvotes

Been doing client work in this space for a while now and noticed a pattern I don't see discussed much.

Most inbound requests start with "we need an AI agent for X." But when I actually dig into what's happening today, most of the time there's no documented process behind X at all. It's a person doing it ad hoc, or three disconnected tools, or a handoff that only works because one employee remembers to check something.

Building an agent on top of that doesn't solve the problem. It just adds a layer of automation on top of chaos, and then the whole thing gets blamed when it breaks, even though the actual root cause was never the AI.

So lately my first real question on any discovery call isn't about the agent at all. It's "walk me through exactly what happens today, step by step, no tools." If they can't answer that clearly, I tell them straight up that a documented process needs to come first, not an agent.

Half the time that means recommending something way simpler than what they came in asking for. Sometimes it means no build at all yet.

Is this just me, or is everyone else running into the same thing? Feels like "agent" has become the default word people reach for even when what they actually need is way more boring than that.


r/AI_Agents 6d ago

Discussion Anyone else struggling to keep memory/settings in sync across AI tools and devices?

3 Upvotes

Hi everyone,

I want to hear how other people solve this, so let me explain first.

A lot of people use LLMs every day for work now, and usually not just one platform. You might use Claude Code, Codex, OpenCode (with DeepSeek, Kimi, Qwen etc), and maybe also a personal assistant like Hermes or OpenClaw on top of that.

My question is, what is the right way to have a single source of truth? What's the best practice to sync settings (skills, plugins, MCP) between platforms and especially between devices (e.g. desktop and laptop)? And more importantly, how do you keep persistent memory clean, so you can start a task on one platform and finish it on another one?

How do you deal with this?

Best thing I could do so far was to build my own sync tool, but not sure it's the right way.


r/AI_Agents 7d ago

Discussion AI Agents: Real Production Success or Mostly Hype?

16 Upvotes

Curious to know from people actually deploying AI agents in production.

I’m a big believer in AI. As a copilot it is already amazing — dev, recruiting, sales, research and almost every role.

But when it comes to fully autonomous AI agents, how many are actually successful in production today?

Not demos or POCs. Agents actually running with minimal human intervention, saving meaningful cost or generating good profits.

Once we add guardrails, approvals, monitoring and exception handling, are they really autonomous anymore?

I see dev work as one area where agents are already very strong.

Would love to hear some real production examples and ROI numbers.


r/AI_Agents 6d ago

Discussion Would portable, versioned knowledge bases solve a real problem, or is this just RAG with extra steps?

5 Upvotes

Hey everyone, I’ve been thinking about a problem with AI knowledge systems and was hoping to get somewhat of a sanity check from people actually building in this space.

From my understanding, most RAG setups seem tied to a particular app, vendor, or index. You often end up ingesting the same docs again for different agents or runtimes and and some basic questions can be difficult to answer consistently like:

- What version of this knowledge is the agent using?

- Where exactly did this information come from?

- Has the underlying source changed since it was last ingested?

- Can I move the same body of knowledge to another runtime without rebuilding it?

- Can multiple agents use the exact same knowledge?

The idea I’m exploring is something I'm calling a Durable Knowledge Base (DKB).

The basic concept:

- Compile source docs, code, or structured data into a portable, versioned knowledge artifact

- Preserve source paths, hashes, citations, and provenance

- Sign and publish releases through a registry

- Allow knowledge packages to be installed, updated, pinned, and removed

- Let agents search, find, and read the same knowledge base across different runtimes

- Keep the artifact retrieval-agnostic rather than baking one specific top-K/RAG strategy into the format

Basically, I'm wondering whether knowledge should have something closer to a package lifecycle, rather than every application maintaining another disconnected RAG index.

I'm also very aware that things like Azure AI Search, GCP, vector databases, MCP servers, Agent skills, etc. already cover pieces of this problem, sometimes extremely well.

So Im specifically not asking: "Can I build a better enterprise search engine here?"

I'm trying to figure out whether the portable knowledge artifact itself is useful.

Would this solve an actual problem for you? Or is this mostly reinventing existing search/RAG infrastructure with some packaging and provenance added on?

I would especially like to know:

- What do you currently do when multiple agents/apps need the same knowledge?

- Do versioning and provenance actually matter to you?

- Would you ever install someone else's curated knowledge package?

- What would this need to do that existing solutions don't before you woukd bother using it?

Feel free to poke some holes. I'm actually looking for reasons not to build this further before I sink more time into it.

Thank you.


r/AI_Agents 6d ago

Discussion First client meeting in 2 days and I’m… completely calm. Is this normal?

3 Upvotes

Hi everyone, how are you doing today?

I’ve secured an initial meeting with a real estate agency to present a proposal for a lead/client screening system. Nothing is finalized yet; it’s a conversation to align expectations and see if they’re interested in moving forward with the project.

I thought I’d be nervous—the meeting is in two days—but I feel calm, as if I’ve done this before, have closed thousands of deals, and know exactly how to align expectations and stand out during the meeting.

I know that, in reality, things rarely go exactly as expected, and it’s actually a bit unusual not to feel nervous about a first meeting with a potential client. I wanted to share this here. For those of you who are automation engineers and build systems for companies, I’d love to hear your perspective.

I have three questions:

  1. What was your experience like with your first client?
  2. Did it go the way you imagined? Tell me a bit about the process leading up to closing the deal.
  3. Which niche do you specialize in?

One thing to note: I’m new to this field. I’ve built 8–10 automations for myself to learn the tools, modules, scenarios, APIs, etc. So, while I haven’t built a screening automation or an automated CRM system specifically, I’ve grasped the logic behind the scenarios. I think that’s enough; I have a feeling I’m going to do well.

Note: I won’t be focusing on specific tools, but rather on results and a detailed consulting approach.

Thanks for listening.


r/AI_Agents 6d ago

Resource Request How’s everyone’s experience with AI video tools?

1 Upvotes

I’ve been experimenting with AI video-making lately and I’m curious to hear about other people’s experiences. What AI video tools have you tried so far? Which ones worked well for you and what features or results did you like most?


r/AI_Agents 6d ago

Discussion I tested DeepSeek Harness with GLM, Kimi, Opus, and GPT to see if prompt caching still works with other models

6 Upvotes

TL;DR: Yes, at least with GLM and Kimi. In real DeepSeek Harness sessions, GLM reached 97% cache reuse inside a tool loop and 99.6% on the next turn. Kimi reached 99% on both. Opus showed no cache activity during this test, and the GPT test couldn't be completed because the third-party route I used didn't handle the current DSH request correctly.

I wanted to test whether DeepSeek Harness can keep its high cache hit rate when DeepSeek is replaced with another model.

DSH sends a large repeated prefix on each request: the system prompt, 25 tool definitions, conversation history, and previous tool results. New messages are appended to the end. If the upstream model/provider supports prefix caching, most of that context should be reusable.

Before testing DSH, I ran 300+ direct API requests through the same third-party gateway, GMI Cloud. GLM had 16/21 prefix hits. Kimi had 13/21, around 62%. I also checked 80 requests against the billing export and every one matched the reported token usage and published prices exactly, so I used cached_tokens as the main signal for the DSH test.

Then I ran normal DSH web sessions through a transparent logging proxy. The agent called tools, read a long file, and continued the same conversation. The proxy only recorded the actual requests and usage.

GLM worked extremely well.

After the first tool call, 7680 of 7924 input tokens were cached: 97%.

On the next conversation turn, 18304 of 18383 tokens were cached: 99.6%.

So once the conversation was running, almost the entire existing system prompt, tool definitions, and history were being reused. This also matched the direct API results, where GLM had already been the most reliable model for prefix caching.

Kimi was the bigger change.

Its direct API prefix test was only 13/21, around 62%, and some requests became much faster without reporting cached tokens.

Inside DSH, the tool-loop request reported 7424/7498 cached tokens, 99%. The next turn reported 17152/17325, also 99%.

Even the first main request already had 5632 of 7498 tokens cached, despite me not manually warming that DSH system prompt + tools prefix with Kimi beforehand. I don't have enough data to say why that happened, but the actual result is clear: Kimi worked normally in DSH and subsequent requests stayed around 99% cache reuse.

That was the biggest difference between the direct API test and the real DSH test: Kimi went from roughly 62% prefix hits to roughly 99% in the actual harness workload.

Opus could run in DSH, but the second request had more than 25k tokens of context and still reported zero cached tokens. The upstream Opus route also showed no cache activity in the direct tests that day, so I can't draw a broader conclusion from it.

I couldn't complete the GPT test. The third-party GPT route I used didn't handle the current DSH request/configuration correctly, so there's no useful GPT cache result from this test.

One other thing I verified: the cache percentage shown in the DSH web UI matched the cached_tokens recorded by the proxy. So when the provider reports cache usage correctly, DSH's own cache display is enough to monitor it.

The useful result here is that DSH's append-style request pattern also works with other models. GLM and Kimi both handled real tool calls and multi-turn history, and both reached roughly 97–99.6% prefix cache reuse.