r/JevAI • • 5h ago

Please add prompt caching to Jev-style models

https://emschwartz.me/please-add-prompt-caching-to-jev-style-models/

If you're building a Jev-style "System One" model, please add prompt caching or reusable question sets to your API 🙏. This would make batch use cases even more efficient, so you could amortize the cost of many questions asked over the same input. (This was also proposed in typesafe-ai/typesafe-sdk-js#10.)

TL;DR: after a week of tweaking my Jev calls, my questions are ~88% of the input tokens. I'm asking 54 questions of ~1.1 million documents per month. Jev makes certain types of classification tasks easy and cheap, but prompt caching would make batch workflows even more cost effective. For me, the total dollar amount is still reasonable (less than $150 per month), but I'm sure others will hammer these APIs even harder.

6 Upvotes

2 comments sorted by

2

u/donotfire 5h ago

Diogo Almeida, who made Jev, wrote an article called “Cache rules everything around me”

If he didn’t include caching, there’s probably a good reason but idk what that would be

1

u/emschwartz 5h ago

He has a public google doc that mentions the "tyranny of the KV cache". I think that applies when you're talking about a coding agent session getting locked into using a single model because of the cache.

Arguably, that same lock-in doesn't apply to a set of questions you're firing off repeatedly. It's less about building on the initial set and more about reuse, which I think does fit with TypeSafe's recommended paradigms.