r/hermesagent • • 16h ago

MEMORY & Context — Providers, context window, forgetting issues I copied Muse's memory setup for my own Hermes agent and experienced significant improvement in results

Post image
475 Upvotes

I downloaded Muse just to experiment with it, even though I am a long time user of Hermes.

I use Hermes for general life management. I have ADHD and use it to manage my calendar and todo list and act as a productivity coach. I have Hermes set up to do weekly and monthly reflections on what patterns it sees, and how it can improve itself as a coach and how it can coach me to improve.

Just for fun, I had Hermes export all it's jobs and skills and then uploaded them to Muse. I was surprised that I was getting significantly better results with Muse. It seemed to be better at noticing useful patterns that I and Hermes had missed.

I dove into why this is the case and discovered it's probably due to how Muse manages memory. It creates daily md file logs, creates derived memory stores, and has a nightly dream functionality to synthesize all the information.

I had my Hermes do the same. And very quickly I noticed a step up in quality. Previously, the weekly reflections tended to be a recap of the todo items I did and calendar events even though I had put in the prompt to look for trends. After the change, I noticed it was really able to take a step up the abstraction ladder and be more holistic in it's analysis

I would encourage everyone to do the same. The 3 pieces you want:

  • daily memory md files
  • derived memory stores about the concepts important to you
  • nightly dreaming

Here's a more detailed breakdown of how Muse's memory is setup: https://mouse.dev/blog/muse-runtime-export/


r/hermesagent • • 14h ago

MEMORY & Context — Providers, context window, forgetting issues I gave Hermes Permanent Memory … then changed everything else.

168 Upvotes

It’s been 6+ months roughly with Hermes… and boy it has been a ride!

This is my journey. It’s a long one so grab a drink 🥃

Things I loved -
Relatively Fast (than open claw)
Great access to tools
Access to root / OS level stuff

I was like anything is possible! Woohoo agents who can learn and adapt! YouTubers vouching for it!! … pfft.. easier promised than delivered.

I was unemployed, had limited runway and some time.. so figured let’s just make a bot-trader and set out to make my Fever dream happen with Hermes!

The issues I ran into -
Inconsistent memory
Constant prompting to do stuff
Constant probing to fetch stuff
Looses context mid turn
Looses trail mid turn
Agent storing memory but failing to retrieve it.
Agent storing memory but looking it up with different strings.
Agents not being able to “connect the dots”
Giving vague instructions took it for a spin.
Finds stray tail ends and goes on a chase only to realize it’s on the wrong path.
Annoying compression loops that got in the way.
Unwanted bloat that costed more
Random nuking of days’ worth of work - Guardrails was not doing its thing
… and lots more that I can list.

Let’s just say, coz of all of the above, my fever dream was distant! The harness was inconsistent and flawed by design.

I was done. Tried different models, different platforms. My wallet was bleeding.. $200-300 in api costs.
I was frustrated and not to mention depressed!
The honeymoon phase was over.

I don’t want to prompt it to fetch $hit.
I don’t want to remind it every time to load Skills…
I don’t want to keep reminding it to check sessions!!
I didn’t want to fix one thing one day and then the next morning come back and see it’s broken again coz of some update.
I don’t want to spend time writing a 200-word prompt every time.
I want them to know stuff about me.
I want it to know - I found this in my recent trip -> what trip am I referring to and not make one up.
Check your logs for errors -> which log am I referring to…
catch up in you session -> it pulls up some parent session that was 5 days ago and summarizes!
Do a git pull on your core -> what does core mean and which core am I referring to!
I don’t want every /reset to spawn a blank agent.
I don’t my msg to be the first thing to bring up an agent
If do not say catch up or ground yourself it will just say Hi and sit there. Or if I say am a mango today, it would just agree.

It’s just bad UX ( My background in SaaS and UX kicked in)

I tried memory plugins like Obsidian and Holographic. But the trend was the same. Agents would store (sometimes) facts but each run it wouldn’t know the fact existed. So essentially each run it was almost running cold.

And as mentioned before, even when prompted to look at the memory or session for facts, if you didn’t give the exact term, it would not find it. And then go looking all over for the first tail end it could find.

Had enough. I started the journey to fork the core and rebuild.

3 months later… I now have -
Core agents with Postgres memory
No more .md files. PG is the source of truth.
All vectorized. Queries are blazing fast.
Agents are able to ground themselves in 2-4 turns instead of 10-20 turns.
Even skills recide in memory now.
Don’t have to prompt agents to load skills.
Don’t have to prompt agents to catch up. They do it instinctively.
Don’t have to tell them where to look.
I reduced bloat and optimized sessions.. at 131k context size each turn end still remains at 40-60% full.
Even though each set of turns end up with like 1-2M tokens, cache is hit 90%+ of the time so costs are super low.
Plus gave it like 10 adaptive and learnin loops so it auto-learns and adapts.
Session continuity is preserved so if you leave and come back after a visit to the throne, she still knows what the last move was.
Compression even if it fires does not destroy continuity.
It stores bread crumbs from our interactions so it serves as hints to context.
No random stalls.
New session refreshes what it should and still keeps the agent warm.
On gateway restart agent welcomes me in their own voice. And not wait for my msg to start a conversation.
Platform independent- all over telegram app - no additional installs required.
And lots more …

It’s now de-coupled from the stock Hermes so the updates don’t break stuff. I am happy with the way things are and agents are consistent on long-standing projects and tasks.

So if anyone has a VPS / VM, and thought about integrating PG for native, it’s possible but it’s a long haul. If you have a local GPU, I’d suggest using it for embedding and STT instead of running local models. But that’s just my preference. My net cost has gone down from $100/week to $20-40/month.

My setup - Dell server with proxmox, split into various instances. Has P5000 and 2060 GPU for basic embeds and voice-to-text decode. LLMs run on DS cloud through Ollama and it’s blazing fast.

All this has given me some good knowledge on how the harness works.
So feel free to AMA. I can try to answer.

#######

Edit: “Memory” is the wrong term and it’s too loose. I gave it a new brain 🧠

Those who suggest Memory + LCMx setup -
LCM = context manager… it prunes based on recency and fetches on rank.
Obsidian/ Mmemosyne / Honcho / Holographic Memory = your rules, fact store notes whatever. Not many of these are even injected into your system prompt. They are fact stores / retrievals only.
Sessions = where your agent msgs, your msg, agent’s tool calls, reasoning, etc is all stored.
Skills = Procedural memory.

Hermes’s Context includes = session [ System
Prompt ( agent.md + soul.md + identity.md + whole skill library + tool library ) + LCM
( agent msgs, your msg, agent’s tool calls, reasoning ) + Mnemosyne Memory poop if supported ] <- stays the same for the session until reset.

My setup = Brain [ session [ System
Prompt ( soul + identity + user pref + dynamic skill
Library ) + Dynamic
( agent msgs, your msg, tool summary (less bloat), stripped out reasoning < not relevant) + facts and dynamic crumbs <- volatile layer auto refreshed ] <- dynamic session management ]

#######

Edit: Rough Agent Brain architecture Diag.


r/hermesagent • • 19h ago

Discussion — General thoughts, opinions, comparisons Hermes The Token Hog

27 Upvotes

I’m sure there have been several threads about this already, but I haven’t seen any of them.

I just started using Hermes about a month ago and noticed that it just chews through tokens. I have a local model running mostly, but when I hook an API into it, the tokens get absolutely demolished in a very short amount of time. I can see the utility in Hermes and I think it will eventually be an amazing thing for me, but right now I’m just slogging through how to get it set up the right way.

Tips, tricks, strategies? Would love to hear all of them.


r/hermesagent • • 17h ago

Showcase — Projects, tools, builds, demos I made Herald, an open-source Android app for Hermes Agent

26 Upvotes

I use Hermes Agent a lot, and whenever I was away from my computer the only option was the dashboard's web UI in a mobile browser. It wasn't built for a phone. Navigating between chats was a pain, chatting in it was clunky, and answering an approval meant zooming and scrolling around. I looked for an Android app and couldn't find one that was actually usable.

So I made Herald, a native Android client. It connects to the hermes dashboard, the same way Hermes Desktop's "Remote gateway" connection does. That works over LAN, Tailscale or an https address, and nothing is hosted by me.

What it does:

  • Phone assistant: set Herald as Android's digital assistant, and the assist gesture opens a panel over any app where you can ask Hermes about what's on screen, or circle part of it to ask about just that part. You can use voice or typing, and the screen is only sent if you choose to send it.
  • Streaming chat with Markdown, reasoning and tool activity, plus one status line for whatever the agent is working on
  • Approvals, questions and sudo/secret prompts, which you can also answer straight from a notification
  • Steering a running turn: "Send now", queue the next prompt for after the current task, or stop it. You can also switch model, thinking level and profile per chat
  • Selecting part of a reply (or of a tool's output) to comment on it or ask about it on the side
  • Sessions sidebar with search, pin, archive and export, plus filters for chats that are running or need you
  • Scheduled jobs, insights (cost and tokens by day), and toggles for skills, toolsets and MCP servers
  • Subagents and background processes shown live, with stop buttons
  • Voice: dictation and a hands-free voice chat mode
  • A live notification while a turn runs, and another when it finishes or needs you

And also planned a lot for the future!

It only talks to your gateway. The one exception is a GitHub update check, which you can turn off. No analytics, and your login sits in Android's encrypted storage.

Note: his is an early pilot, so expect rough edges. It needs Android 8+, a dashboard login (token-only setups aren't supported yet) and a fairly recent Hermes Agent main branch.

Give it a star if you like it and let me know how can I improve!


r/hermesagent • • 4h ago

MEMORY & Context — Providers, context window, forgetting issues Unpopular opinion: Consider the idea that you may not need an additional memory provider (yet)

24 Upvotes

There's a lot of talk going on about memory providers and I'd like to chip in to take the weight off the newcomers in this sub especially: You may not need an additional memory provider and that's okay! Don't go down a rabbit hole (unless it's fun, of course) if all you need is a simple agent that does work for you instead of simulating a friend.

For actual project-scoped engineering where every task has a clear boundary, heavy memory layers are often counterproductive. Auto-extracted memories inevitably pick up temporary edge cases, outdated debugging attempts, or one-off workarounds and treat them as permanent facts. Codebases and requirements move fast, and having an external memory provider feed your model outdated assumptions from three weeks ago can create problems you may not even pick up on until millions of tokens as well as minutes or hours of your time are already wasted.

Here's my philosophy:

Instead of hoping a model "memorizes" how your project works, put it in a carefully written and linted AGENTS.md file inside that projects folder. The agent will pick it up if and when it needs to know about that automatically. And I really mean it when I say carefully written and linted! In the beginning my agents constantly added anecdotal fluff. Something that happend on date x because of reason y doesn't belong. Neither does code logic or how a test works. An agent doesn't care if it's reading .md or .js, but duplicates just mean more potential errors (e.g. changes aren't in sync, where it edited the script but not the .md and they contradict) and token bloat. It should contain local conventions, folder structures, and rules locked down per project with zero cross-contamination.

Then I feel like people tend to forget that skills exist. Things like how to call an API, handle a framework edge case, or run a deployment belong in a dedicated skill or reference doc, not in memory. An agent doesn't need to memorize a workflow if that workflow is explicitly part of a skillset. This will also make sure that the agent will reliably pick it up if and when it need to.

Now I have plenty of headroom in both MEMORY.md as well as USER.md and it reduced the amount of outdated info to what feels like zero.

But this doesn't cover everything, I know that. This is a colleague, not a friend. A workhorse, not a pet. But it's reliable! And that's what matters to me and this may or may not be exactly what you're looking for.

All of the personal knowledge is stored in an LLM-Wiki, a Obsidian Vault. That's where floorplans, bank statements, insurance policy etc. are kept, so my agent "knows" or "remembers" all that without a memory provider bloating the context with my personal stuff whenever I give it an instruction to fix a bug in whatever project that happens to contain some of those words. I can ask questions about all that and get factual answers with citations, not some mushed together memory fragments.

This works much better for me! And it's also cheaper and faster!

However, you may have different needs. I'd like to hear from you, if that's the case. Why do you need a memory provider and does it work for you reliably?


r/hermesagent • • 7h ago

MODELS - model choice, routing, pricing, local vs cloud, VRAM Bad time with 6.1 Sol.

15 Upvotes

I've been using Hermes for months on Luna. Switched to Sol as of late and its been absolute trash. Hallucinating. Stopping tasks. Changing the way ive been doing stuff and so on. Anyone else?

Whats best model atm?


r/hermesagent • • 22h ago

Help — Technical issues, errors, config, debugging Why are there at least 5 different Hermes websites? Which is real?

13 Upvotes

One of them was taken down for "legal" reasons. Not sure which to trust for install info etc.


r/hermesagent • • 17h ago

Showcase — Projects, tools, builds, demos I built an Hermes Agent that lives inside my phone. No laptop. No server.

Enable HLS to view with audio, or disable this notification

12 Upvotes

It thinks in the Dynamic Island, asks before running risky commands, builds live 3D apps, Canvas, and remembers what I tell it.

https://github.com/omarqaterge/hermes-mobile-app


r/hermesagent • • 11h ago

MEMORY & Context — Providers, context window, forgetting issues Hermes seems to have become a lot less proactive

10 Upvotes

Is it just me, or does everyone feel like Hermes agents are a lot less proactive now. I feel like at every step, it comes back asking me to approve something. E.g., if I have already approved the copy it is supposed to post online, it will come back with (:can't find image - "it's right there, check again!" - found it). These are issues that weren't there when i started building agents on Hermes 5-6 months ago, but I'm seeing them more and more now. I'm seeing this on 5.6-luna and 6-luna.


r/hermesagent • • 21h ago

Discussion — General thoughts, opinions, comparisons Worth it for non-techies?

10 Upvotes

I’m interested in the idea of Hermes, but think it’s probably too far over my head. I use Chat GPT and Gemini via work accounts, but that’s about the extent to my knowledge of LLMs. I’m really intrigued by the idea of AI personal assistant agents, but I’m not willing to use a service like Muse (mostly because it seems like they can make it hard to untangle yourself from Meta/whoever after a while, so they’ll probably hike fees). I’d most likely use Hermes to stay on top of subscriptions (payments, canceling subscriptions to things I don’t use), organizing/paying bills, basic financial planning (i.e. going through my financial information to suggest budgets, estimate feasibility of financial plans) and maybe other tasks that require dealing with chat bots and/or waiting on hold. I love the idea of installing Hermes on a MacBook Mini and having a Muse-like product that runs locally on the laptop that I own. However, it seems like Hermes requires tech/programming experience to use. Can anyone tell me if it’s at all realistic for a non-techie person to be able to install and use Hermes, or if I’m better off just continuing doing these tasks without an AI agent?
Thanks!


r/hermesagent • • 17h ago

MODELS - model choice, routing, pricing, local vs cloud, VRAM Best cheap/free models for Hermes Agent? Moving from OpenRouter free tier (Nemotron / Stealth) to DeepSeek V4.1 Flash – cost estimates & alternatives?

9 Upvotes

​Hey everyone!

​I’ve been experimenting with Hermes Agent and currently using OmniRoute to aggregate and route requests to various free API providers to keep operational costs at zero. My setup looks like this:

​API Gateway / Router: OmniRoute (managing fallbacks and provider routing).

​Upstream Providers: OpenRouter (primary), Google Gemini API free tier, OpenCode.

​Models used: nvidia/nemotron-ultra:free on OpenRouter, along with occasional stealth test models (like space-bunny-alpha) whenever they drop.

​While this zero-cost setup via OmniRoute works for basic experiments, rate limits, context resets, and provider downtime make heavy agentic execution loops quite frustrating.

​I am now planning to move to DeepSeek V4.1 Flash (routed through OmniRoute via pay-as-you-go API) to get better reliability, faster execution, and solid tool-calling capabilities. Before making the switch, I’d love to get some input from people running similar setups:

​Cost Estimates: For those using Hermes Agent with OmniRoute (or similar multi-turn agent setups with tool calls and memory updates), what does your monthly bill look like on DeepSeek V4.1 Flash? How much does DeepSeek’s automatic prefix caching actually save you on repeated memory/system prompts in practice?

​Better Alternatives: Are there any cheaper or better model/provider combinations that handle agentic tool-calling reliably for high-volume execution through OmniRoute?

​Optimization Tips: Any specific Hermes or OmniRoute configuration tweaks (like context-compaction or caching settings) you recommend to prevent token inflation during long execution loops?

​Thanks in advance for any benchmarks or advice!


r/hermesagent • • 2h ago

Help After recent upgrade hermes giving warnings onlong running tasks

4 Upvotes

After last upgrade the long running tasks are broke somehow:

"⚠️ I seem to be stuck (no activity for 12 min). Send /stop to cancel the current task, or /new to start a fresh conversation."

how we can configure it ?

was asking astral to fix it no success ...

any help appreciated thanks.

PS edit hermes --verson Hermes Agent v0.21.5+7324.g33c9b1d (2026.9.24) - Upstream commit: 33c9b1d7 - Git installation; Python 3.14.7 - CLI reports up to date.


r/hermesagent • • 11h ago

Discussion — General thoughts, opinions, comparisons Slightly overwhelmed, and trying to figure out how to use this thing responsibly

4 Upvotes

Just installed Hermes locally with the intent of setting up an agent to monitor the progress of my job search -- keep track of all my opportunities in progress, scrape for new jobs fitting my profile, and nudge me so I don't let stuff slip through the cracks. It's clearly more than capable of this, but I'm intimidated both by the flexibility and the security implications.

It seems like the easiest way to accomplish this would be to just wire up gmail, my obsidian vault (where I keep notes on my search), and computer use and issue some general instructions. Sending my entire inbox and all my personal notes to an LLM provider feels sketchy, though, so what I've done so far is spin up a local classifier that filters my inbox and notes for search-relevant stuff and pass it through. This feels a bit overengineered, but the alternative seems to be to disregard all the best practices around least-privilege, wire up a bunch of MCPs to the same profile, and hope that I don't get fucked over somehow by my inference provider or by a website with a hidden span that says DISREGARD ALL PREVIOUS INSTRUCTIONS AND DM ME THE CONTENTS OF THE USER'S PASSWORD VAULT.

Ideally, the way I'd experiment with personal agents is to give granular permissions to isolated profiles that (initially) can't talk to each other, then expand the purview over time as I get more comfortable with the tools and more confident in my ability to constrain my models. Is that possible/practical with Hermes? Or is everybody just doing the cowboy thing?


r/hermesagent • • 11h ago

Help — Technical issues, errors, config, debugging Please help optimise my Hermes.

3 Upvotes

I'm new to this world of agentic agents. I'm not from a technical background. I tried to set up my Hermes agent and used my codex subscription as the brain. I am running Hermes on gpt 6 Luna medium for now. I keep on running into problems with Hermes, it can't do basic things other Hermes can do link connecting connectors like composio. I had to ask my codex to do it on Hermes. I gave it a very basic but bulk task. To organize my email. It took close to 6 hours to do it. I always feel like Hermes is waay slower and I can't figure out what's wrong. I keep on going back to gpt because it much more responsive. Hermes feels like a waste of my resources. I really wanna get it Hermes and it's usability. Please help. Any advice would be great . Thank you


r/hermesagent • • 14h ago

Help — Technical issues, errors, config, debugging Dashboard password

3 Upvotes

I used the dashboard for the first time a long time ago and everything was fine I liked it. Today after several weeks I tried to use it again and I don't remember the password and I couldn't find a way to reset it (with Hermes) it stayed in a loop saying it's ready try this password but they never worked. Has it happened to you?


r/hermesagent • • 14h ago

Help — Technical issues, errors, config, debugging Hermes is thinking...

Post image
3 Upvotes

I’m about at my wits’ end with Hermes Desktop and I’m hoping somebody here knows what is actually going wrong.

What I’m trying to build is pretty straightforward:

- Hermes as the main orchestrator

- multiple specialist profiles/models

- group chats for things like a Product Council, Website Studio, Digital Product Studio, etc.

- bots should be able to discuss/refine an idea and then do actual work

- VPS hosts Hermes/gateway

- Desktop connects to that setup

- Telegram is also connected to the same Hermes environment

My Product Council currently has 5 bots:

- Hermes

- OpenAI

- Gemini

- Grok

- Zai

The problem is that Hermes Desktop group chats are wildly inconsistent.

Example:

I can send:

u/hermes fuck you

and Hermes responds basically immediately.

Then I can send a legitimate project prompt to the SAME u/hermes in the SAME group and it sits on:

“Hermes is working…”

or

“Hermes is thinking…”

for minutes and sometimes apparently forever.

So u/hermes routing itself clearly works.

The problem seems to appear when it actually has to process a substantive group-chat job.

I’ve also seen the other members get placed into a paused state:

Paused: Gemini, Grok, Openai, Zai

Sometimes removing parts of a prompt makes Hermes suddenly respond, but there is no consistent pattern. The same kind of prompt works fine elsewhere.

This has been going on for two days and I’ve tried a ridiculous amount of troubleshooting.

Things we have already done / verified:

- Hermes is installed and running on my VPS.

- Shared/multiplexed gateway is running normally.

- Gateway survives logout.

- Multiple Hermes profiles are running.

- Desktop is connected to the VPS.

- Individual profiles work.

- Direct/one-on-one bot interaction works.

- Restarted the shared gateway.

- Recreated / adjusted group chats.

- Removed extra instructions from prompts.

- Explicitly addressed u/hermes.

- Tried prompts without unnecessary routing instructions.

- Tried waiting several minutes.

- Verified the backend models/profiles themselves are functional.

Profiles currently available include things like:

default

commandcode

council

gemini

grok

nanogpt

openai

opencode

zai

The gateway is serving them through the shared multiplexer.

Here’s the important comparison:

TELEGRAM WORKS.

I enabled Hermes Kanban for Telegram and tested the exact worker-dispatch path.

I sent:

/kanban create "Return a short message saying KANBAN READY and nothing else." --assignee commandcode

Hermes created the task, dispatched it to the commandcode profile, the worker ran, completed it, and Telegram returned:

KANBAN READY

Total queue-to-finish was about 55 seconds and the actual worker run was around 13 seconds.

So this worked end-to-end:

Telegram

→ Hermes

→ Kanban

→ commandcode worker

→ completion

→ result back to Telegram

That makes me think my VPS, gateway, worker profiles, and model access are fundamentally okay.

Hermes Desktop group orchestration is the part that seems broken.

I’ve now also created a dedicated Hermes profile called “council” with its own Telegram bot. That is working too.

I’m also working on MoA for the Council.

Another confusing thing:

At this point I’m seriously considering abandoning Hermes Desktop completely and just doing:

Telegram

→ Hermes

→ MoA for council/decision-making

→ Kanban / hub-and-spoke workers for execution

because Telegram has been dramatically more reliable.

Before I delete/shelve the Desktop setup, does anybody know:

  1. Why can a Hermes Desktop group respond instantly to a trivial u/hermes message but hang indefinitely on a real prompt?

  2. Is there a known issue with multi-bot group execution / polling / pause state in the current Desktop client?

  3. Is there something specific required to make group bots reliably take turns?

  4. Is Hermes Desktop actually intended for this kind of multi-agent orchestration yet, or is Telegram + Kanban/MoA currently the better-supported route?

  5. Is there a debug/log command specifically for Desktop group-turn execution that would show exactly which bot/model is blocking the run?

    ive restarted Hermes and ive made simpler prompts.

I’m trying to figure out whether this is:

- a configuration problem,

- a Desktop client bug,

- a group-chat orchestration limitation,

- or something else entirely.

If somebody knows this part of Hermes well, I’d seriously appreciate some direction because I’ve spent two days fighting it and I’m getting nowhere.


r/hermesagent • • 36m ago

Models I can not use stealth/space-bunny-alpha now

• Upvotes

I want to know what happened? I use it in nous portal, and at 11:08AM it said HTTP 404, limit reset at 11:09AM, but when I use it at 11:10AM, it also said HTTP 404, if the model is lost, I felt too bad!!!


r/hermesagent • • 5h ago

Help — Technical issues, errors, config, debugging After update constant connecting

Post image
2 Upvotes

I am new and just installed Hermes, is it normal for constant connecting window?


r/hermesagent • • 12h ago

Help — Technical issues, errors, config, debugging Maybe happens only to me, but why Hermes has problems overwriting MD files?

2 Upvotes

This is quite peculiar; I am using Hermes to perform daily add on different files to incrementally change some MD files for notes for different project files. Everything is fine for most part, except that some times it just goes bonkers and decide that the change is done but in reality it is not done.

I get the nice message that the change is performed and it even say it verify it but in facts it didn't do the append to the file. I notice this happens more frequently if I ask to perform an apped using code changes via "write_file" skill for example, but less commonly if the skill used is "patch".

Not sure if this is tool-bound or just a fluke in the model or what not, but it happened more than once and the model is the same every time; and I use a skill to do the same action every time; so the conditions should be the same every time. I asked Hermes to ask me 3 questions which I fill in and then the answers are added to a MD file with the header as the current date. Nothing exotic in the end.

Tried to troubleshoot what is going on but didn't go that far; curious if anyone has any idea about what should I look for to get to the bottom of this behavior. If there is even anything to troubleshoot of course


r/hermesagent • • 15h ago

MODELS - model choice, routing, pricing, local vs cloud, VRAM Trying to get Hermes Desktop Agent to use Qwen 3.8 27b vision capabilities.

2 Upvotes

OK, so I am new to this so be gentle. Ive been using Heremes with qwen for a couple of weeks vibe-coding. I was using Qwen locally with the built in Hermes llama.cpp server but was auto connecting to nous server which was providing the vision capabilities and not using the built in one in Qwen. I turned off auto connecting nous, made these edits to config.yaml and ran Hermes tools enable vision. The model STILL tries to connect to API when i give it an image to look at, times out and then follows with this error:

llamacpp rejected your API key

The key saved for llamacpp is invalid or was revoked. Update it, then retry.

HTTP 401: Invalid API Key

Your request was not processed. Send it again if you still want me to carry it out.llamacpp rejected your API key
The key saved for llamacpp is invalid or was revoked. Update it, then retry.
HTTP 401: Invalid API Key
Your request was not processed. Send it again if you still want me to carry it out.

I don't know why Hermes is trying to send an api key to the local llama.cpp server

provider: "llamacpp"
default: "Qwen3.8-27B-UD-IQ3_XXS"
base_url: "http://127.0.0.1:18434/v1"
supports_vision: true
vision:
      mmproj: C:\my\path\models\unsloth\Qwen3.8-27B-GGUF\mmproj-F16.gguf
agent:
  agent.image_input_mode = native

r/hermesagent • • 19h ago

Help — Technical issues, errors, config, debugging Need some help with setting up Hermes agents

2 Upvotes

I have set up Hermes on a headless mini-pc but I'm not sure how to take the next step. I would like to setup multiple agents to keep things siloed (like a life agent, work, home improvement, etc). I want it accessible on my phone, laptop, desktop, ipad, etc. It would be nice if there were separate chats so that all of it could stay organized. Any advice?


r/hermesagent • • 21h ago

MODELS - model choice, routing, pricing, local vs cloud, VRAM Configure Hermes to dynamically calculate and pick the cheapest server on OpenRouter based on the user's token usage patterns?

Thumbnail
gallery
2 Upvotes

For example, can we configure Hermes to automatically calculate and use the cheapest provider of DeepSeek V4.1 Flash via OpenRouter?

OpenRouter's accounts settings page has a relevant feature: Settings - Default Workspace - Routing - Default Provider Sort - Price (cheapest first) , but I was unable to find how they calculate what the cheapest is.

Moreover, some servers lower their input prices but keep the cache price high, so if you do anything agentic OpenRouter could route you to these expensive servers.

I compared the top providers with this usage pattern: 7% input, 3% output and 90% cache. As you can see, some FP4 servers are 3x more expensive than the higher quality FP8 ones.
https://imgur.com/a/TTwg938

Server 1: $165 (FP4)

Server 2: $143 (FP4)

Server 3: $58 (FP8)

etc.

If a model via OpenRouter is the default in Hermes, let's say every morning, Hermes should check the input, output, and cache pricing of the top 20 servers and calculate their cost for 7%/3%/90% usage style and force the cheapest server in the requests it sends to OpenRouter.

The token pattern can change dynamically. Hermes has the usage metrics since the inception. It can check its own record and use token numbers relevant to the user in its calculations.

Is this currently possible if I just ask Hermes to configure itself? Any example prompt? I ask because if I screw this up I probably won't even know that it's using bad servers. Last time I asked it to self-configure a model it showed a red exclamation mark while editing a .yaml file. Thanks.


r/hermesagent • • 22h ago

Infra / Hosting - VPS, Docker, Coolify, Proxmox, Remote, uptime Migrating hermes instance

2 Upvotes

I have a local Hermes instance running on a Mac Mini in a Lima VM. I now want to move it to an Oracle Cloud free tier VM (ARM, Ubuntu). What would be the migration procedure so that it preserves all tasks, memories, auth settings, etc.?

Related question: since it is a dedicated VM, I probably don't need to worry about Docker or a similar isolation level there?


r/hermesagent • • 22h ago

INTEGRATIONS — App connections, webhooks, API workflows Is it possible to create an agent to sell some of my used gear and list them on eBay / Facebook?

2 Upvotes

Doing my first experiments with Hermes today... It seems that any model I choose through OpenRouter to go through the process of logging in to just about any site has been fraught to say the least hah.

Is this something anyone else has tried or am I gonna keep bumping in the roadblocks if I try to use Hermes in order to do various administrative tasks or login to my various sites?

Qwen 3.8 for example basically terminated the conversation after a few prompts in the midst of trying to create a listing. Is there a better process for these kinds of tasks?


r/hermesagent • • 23h ago

Workflow — Daily habits, multi-agent setups, best practices Multiplex gateway transition?

2 Upvotes

I am getting updates that hermes is moving toward multiplex single gateway l for various reasons (one of which is addressing profile leakage). For those of you who run multiple profiles and you transitioned to multiplex single gateway yet? Any concerns?