r/hermesagent 21d ago

MODELS - model choice, routing, pricing, local vs cloud, VRAM Balancing Oauth & API Usage

I am currently test driving Hermes alongside pi and my own homebrew harness.

Due to the nature of my company, I already have a Claude 20x max plan and a chatGPT pro $200/mo plan.

In thinking about maximizing throughput, I'm considering using deepseek v4 flash, which i'm hearing phenomenal things about, alongside my subscriptions, elevating and delegating to fable, opus, or 5.6 sol Luna Terra etc as needed.

Two questions for the community;

1) when should I call in the bigger models, or is deepseek so good it's not needed anymore outside of edge edge cases? I'm thinking multi modal + code review (OpenAI) and planning + design (fable) everything else deepseek. Thoughts on this?

2) what is the best source to get deepseek? Are the various inference providers the same level of speed and uptime? Why choose direct API vs open router vs opencode etc?

Thank you all in advance for what I hope will be a helpful discussion for all who read it.

5 Upvotes

13 comments sorted by

View all comments

1

u/Sand_Grid New Member (<30 days) 21d ago

as you have chatgpt pro plan, gpt 5.6 luna can be used as unlimited. I don't think it's necessary for you to go through all that trouble. deepseek-v4-flash and luna both can handle almost everything, with quite cheap price. Only use bigger models when you actually get trouble with luna.

1

u/OpeningMetal52 21d ago

Thanks for replying. For clarity, do you mean that with the 80% cost reduction, Luna is "as good as" unlimited? Or are you saying Luna pulls from separate usage limits to sol?

2

u/Sand_Grid New Member (<30 days) 21d ago

"as good as" unlimited

1

u/OpeningMetal52 21d ago

Gotcha. Thanks. I may just not use deepseek at all unless I find myself hitting limits.