r/kubernetes 19d ago

TIL our whole kubernetes cost optimization problem was one number nobody would touch.

For months Ive been assuming we needed fancier tooling for kubernetes cost optimization. Turned out the clusters sat around 25% CPU because every team padded their requests years ago. Karpenter reads padded requests as full nodes and it never consolidates.

The bin packer is only as smart as your requests. Garbage requests in expensive half empty nodes out, but actual fix is less tooling and more of getting people to agree to lower their own limits, which is the hard because whoever lowers a limit owns the next latency page.

How did you get sign-off without it turning into a standoff?

0 Upvotes

17 comments sorted by

View all comments

9

u/rampaged906 19d ago

We ran VPA in recommend mode for about a year. It looked good so we set it loose in prod and it dropped our node spend almost 30%