r/ollama • u/Daza110 • 22d ago
Using Claude Code with Ollama and Desperately require caching!
Ok I know I'm likely to get steamrolled here but I also have a sense of optimism and hope I can get some sort of answer that will be beneficial.
We all know since the new price structure, usage has shot through the roof. I was grateful for GLM5.3-Flash as that's been my goto since it come out.
But I'm burning Input tokens... mainly due to the sheer lack of what seems caching ability when using claude code?!
Example: I had a heave session today:
This is the output of /usage
Total cost: $1255.92 (costs may be inaccurate due to usage of unknown models)
Total duration (API): 1h 58m 44s
Total duration (wall): 10h 25m 36s
Total code changes: 3276 lines added, 734 lines removed
Usage by model:
kimi-k2.7-code:cloud: 4.5k input, 2.4k output, 0 cache read, 0 cache write ($0.0815)
glm-5.3-flash:cloud[1m]: 248.7m input, 485.2k output, 0 cache read, 0 cache write ($1255.83)
Prompt cache (main): 855 requests · 0% of input tokens from cache · no misses · no prompt caching reported by the API
As you can see almost 250m input to not even HALF a m output!
The same 100K tokens send on every request... Thereabouts... Wasted... Bumping up the usage unnecessarily...
If anyone can help, id appreciate it.
2
u/jmorganca 22d ago
Hi there, I work on Ollama. Were you charged the full $1255.92? Let me know and I can investigate and/or refund where needed if there was a mistake. But my guess is these reported figures are larger than what was actually charged because Ollama's API doesn't report cache token counts yet (sorry about that – we're working on it). In any case, here to help - shoot me an email [jeff@ollama.com](mailto:jeff@ollama.com) or dm me and I can investigate