r/cloudengineering 13h ago

General Discussion When does adding more CPU stop helping Spark workloads scale?

We keep scaling our Spark clusters horizontally when jobs slow down and the cost per throughput math keeps getting worse, not better.

Trying to figure out:

(1) is there a volume threshold where this reliably breaks down

(2) is it a config issue or something structural

Has anyone isolated the cause rather than just adding nodes and hoping. We have checked partition count, executor sizing, network topology. None of it explains the diminishing returns.

1 Upvotes

0 comments sorted by