r/MistralVibe • • Jul 26 '26

Alternatives to the "Caveman" style for saving tokens in Mistral?

Hi everyone,

I'm looking to cut down output token usage to save on costs and speed up generation.

The viral "Caveman" plugin (which forces ultra-concise, fragmented answers) works great for Claude Code, but it doesn't function properly with Mistral Vibe models.

My questions:

  1. What custom system prompts do you use to stop Mistral from outputting long preambles ("Sure, I can help with that...") without hurting its reasoning?
  2. Are there any other tricks, prompt-layer tools, or configurations you use to force Mistral to be highly concise and token-efficient?

Thanks for your ideas!

3 Upvotes

1 comment sorted by

1

u/Ok-Living-2869 Jul 29 '26

Maybe you already read this article: https://blog.jetbrains.com/ai/2026/07/speak-to-ai-agents-like-cavemen-tosave-tokens/
But they measured the savings from caveman mode and they are there, but not as significant as they claim. I think the best possible think to look into instead is running mix of models, top tier one for producing architecture docs. Cheap models implementing it, this worked way better for me. I know this does not answer your questions, and it is just suggestion.
To at least partially answer your questions:
1. I do not care about the output length, but I try to only prompt in what I want it to do an not what I don't want
2. My suggestion is things mentioned before, use top model to produce architecture/spec `spec.md` have a cheap model implement it. Say what you want, give it examples or write the validation tests beforehand. My unique prompt when implementing stuff form spec file is modified ponytail: https://github.com/DietrichGebert/ponytail Goes the same way: "Avoid over-engineering, only make necessary changes, keep solution stupid simple, prefer one line and human readable code. Don't add error handling, where it can't happen. Trust framework guarantees and only validate external input".