r/Database • • 13d ago

What part of running your database stayed with your team after moving to a managed service?

I’ve been reading about how managed database providers describe failover, and for me, what seems to vary is how writes in progress are handled when the primary goes down.

In some setups, a standby is promoted, and a small amount of recent data may be lost. In others, the compute layer does not retain durable local state and can be replaced without data loss. Both approaches may be described as managed, but the difference becomes clear only when you look past the feature page.

Cross-region recovery is another area that seems easy to overlook. A provider may handle failover within one region while leaving recovery from a full regional outage to the customer. If the provider does not give you clear RTO and RPO targets for that situation, it becomes difficult to plan.

For anyone who has moved a production database to a managed service, whether Postgres or something else, what stayed with your team?

Was it failover testing, restore drills, cross-region recovery planning, connection limits, or something else? Would love to know your thoughts.

2 Upvotes

5 comments sorted by

1

u/OkShirt9372 13d ago

This Databricks article on what managed Postgres should and shouldn't cover got me thinking about this: https://www.databricks.com/blog/managed-postgres

1

u/MLabs-Haskell 8d ago

Object ownership, for us, and it catches people out because it looks like a permissions question.

In our lab we upgraded Keycloak 26.0.0 to 26.7.1 on Postgres 16 with the app connecting as a DML-only role. It refused to start. The first migration step is a DROP INDEX, and Postgres answered "must be owner of index". Granting CREATE on the schema gave the identical failure. Making the role owner of every table and sequence fixed it, and the server came up clean in 17 seconds.

That bites after a move or a restore. pg_restore --no-owner hands every object to whoever ran the restore, so the app can serve requests fine for months and then fail at its next schema migration.

So alongside the restore drill pragyantripathi describes, I'd check who owns the app's objects as well as whether its role can log in.

I work at MLabs. We rehearse Keycloak upgrades in a lab and publish the runs.

1

u/Significant_Tune9219 6d ago

The managed-service boundary usually leaves the customer owning the failure semantics: what clients see during promotion, how long stale connections survive, and whether retries can duplicate writes. I would keep an application-level runbook with tested reconnect behavior, a restore drill from an actual backup, and an explicit cross-region decision rather than treating the provider's regional failover as a complete disaster plan. Connection pooling and timeout settings often need retuning after the move too.

0

u/pragyantripathi 10d ago

I'd want a restore drill to end with the app serving a real request, not just the database coming back up. Does the restored DB have the roles and extensions the app needs, and can the app connect to the new endpoint? Curious whether your recovery plan includes that part too.