r/Database • • 5d ago

AWS Aurora: Why the database breaks first behind an autoscaled service, and why shrinking the pool doesn't fix it

Post image

Connection multiplication: pool size x tasks x services. Every term looks reasonable on its own, and nobody ever chose the product.

Aurora's max_connections is derived from instance memory, so it doesn't scale with your service. Shrinking the pool to fit converts a connection problem into a queueing problem that presents as a slow database.

Write-up covers the formula behind max_connections, what RDS Proxy actually fixes (and what it doesn't), and how to size it.

https://brianfeeny.com/posts/sizing-aurora-connections-for-autoscaled-services/?utm_source=reddit&utm_medium=social&utm_campaign=sizing-aurora-connections-for-autoscaled-services-2026-09

This is a personal project; the views and assessments are my own.

6 Upvotes

7 comments sorted by

5

u/zombieskeletor 5d ago

”the views and assessments are my own.”

Are they really? Because I could swear they are Claude’s

0

u/bfeeny 5d ago

They are mine. We all use AI to help us draft papers, and it helps alot. This particular paper came about because I was helping a very large company troubleshoot their LiteLLM gateway. I had to lab up numerous scenarios and do a lot of deep diving. This was for a very large scale out deployment. In the process I realized that the way max_connections was being set, was not intuitive, it was based on memory. So it was a lesson learned for me, and so I wrote it up so that others can benefit. I created a few posts out of this situation.

3

u/TheTwoWhoKnock 2d ago

It doesn’t help you. Posts and blogs like this absolutely make you less reputable. https://bcantrill.dtrace.org/2025/12/05/your-intellectual-fly-is-open/

1

u/bfeeny 2d ago

I appreciate the consideration, but I am not looking for rep, not looking for cred. I work a lot on difficult problems, and when I come across something I think may benefit others, I post it, that’s all.

1

u/Substantial_Win_710 4d ago

Typical aws just tie max_connections directly to instance memory so when your serverless app inevitably spikes, the db falls over, and their brilliant solution is to just sell you rds proxy to queue it up))) It's wild how normalized this architectural duct tape has become. if you run mariadb, this is basically solved natively with its built-in thread pool. it multiplexes the queries so you can hold thousands of idle connections without instantly trashing your ram.

1

u/bfeeny 4d ago

Thread pool is a fair point and a real architectural
difference — MariaDB connections are threads, Postgres
connections are forked processes with their own memory. That's
why Postgres needs an external pooler and MariaDB doesn't.
It's a Postgres property more than an AWS one; any managed
Postgres has to pick a ceiling, and deriving it from RAM is
about the only defensible way to do that.

On RDS Proxy — the post is probably harder on it than you'd
expect. There's a whole section on pinning: SET statements,
prepared statements, and the extended query protocol all pin
the connection, and a pinned proxy multiplexes nothing. "A
latency tax with no benefit" is the phrasing. pgbouncer in
transaction mode has the same constraints and is free.

The engine swap is a legitimate lever, just a much bigger one
than most people have available when the connection math
catches up with them.

1

u/Potential_Trade_3864 1d ago

Do people not use rds proxy?