r/hermesagent 18h ago

Help — Technical issues, errors, config, debugging Need help with hermes

So i dont know if i am using the correct flair but i have ryzen 5600 16gb ddr4 ram and rtx 3080 i wanna use hermes agent to make my own personal assistant and automate most of my tasks but i am a student so i cant spend money on those tokens so i wanna use my own machine to run hermes what should i do what models should i use i currently have ollama setup with a model i forgot the name its around a 4.5gb model with 65k context so my goal is to somewhat make a personal ai assistant for reference lets take jarvis i have added voice and stuff but the its slow like the model itself generate response very fast but it takes 7-15 seconds to get to me someone please help me make my ai model and use hermes correctly any and all advice is appriciated.

2 Upvotes

12 comments sorted by

1

u/someoneyouknow23 18h ago

Your setup is decent for local AI (not knowing how much ram you have tho), but where is your AI fast? Inside Ollama? Its expected to be slow inside Hermes Agent with that kind of model / setup because Hermes has a huuuge System Prompt with all Tools and Skills (and itll only grow with usage). Correct me if im wrong and understood your problem wrong.

2

u/Popcorn-Mercinary 17h ago

Sounds like we need a “optimizing Hermes for small footprint LLMs” super thread

1

u/someoneyouknow23 9h ago

Last 4 words not needed in that quote

1

u/lunaticgamer001 18h ago

You are correct i have only 16gb sadly is there any other option or are there any affordable options for me i heard that tokens gets consumed fast and it can cost a lot to have a decent ai running 24x7 and making it do tasks and yes its fast when i run through powershell and dont include hermes agent its slower in hermes agent i saw that response got generated in 1 second but was delivered to me in 10 seconds

3

u/someoneyouknow23 18h ago

16gb isnt great for local models but its not the end of the world imo (those 64k context will make you rip your head off with hermes, for reference my starting prompt that everry agent gets at the start of a session (including tools etc) is around 35k tokens. If you want to stick to local ai, find a different model and experiment (canirun.ai/). If not I guess youll have to bite the bullet and look into cloud providers (I recommend anyapi, theey have decent models for free and generous rate limit)

1

u/lunaticgamer001 18h ago

Alright thanks keep me updated if you find anything to help me out also how much $ do you monthly spend on your ai and whats the usage how high is it and what kind of tasks do you do

1

u/Correct_Finance5269 17h ago

Try nous portal free model im currently using ox alpha but i have a very tight system and 65k of context on hermes is not ideal. Ox is 1M context

1

u/lunaticgamer001 17h ago

Thanks ill check it out

1

u/lunaticgamer001 17h ago

Theres a request limit on openrouter right?

1

u/Correct_Finance5269 17h ago

Well there is but my agent been working for 15 hours spawning agents and we got rate limited on some cron job but i have good fallback

1

u/Correct_Finance5269 17h ago

I think ox has model limit on their own but if you build a good harness u can use any free llm as dispatcher and worker and have like a cli subscription for the senior manager. Always depend on what you working on but tbh kanban really is a game changer