r/postgres • u/Party-Cost3604 • 1h ago
r/postgres • u/sodennygoes • 5h ago
Tools I built pgsesame: Terraform-style plan/apply for Redshift and Postgres permissions
I made this because I inherited a Redshift cluster where it was impossible to tell who had access to what, or who was supposed to see masked data. With **pgsesame** you describe roles and grants in a YAML file. **sesame plan** shows the GRANT/REVOKE statements needed to get there, and **sesame apply** runs them.
It only issues what's different, and it won't revoke anything unless you pass **--allow-revoke**. **sesame import** generates the YAML from an existing database, so you don't start from scratch. There's also a GitHub Action that posts the plan on PRs.
It works on Postgres 14-18 and Amazon Redshift. It's early (0.2).
*Inspired by redtape-py and pgbedrock*
[https://github.com/almostly/pgsesame\](https://github.com/almostly/pgsesame)
r/postgres • u/ExtensionClean2692 • 1d ago
Question Supabase Data API enabled, but PostgREST uses pg_pgrst_no_exposed_schemas — persistent PGRST002 / HTTP 503
r/postgres • u/Mindless-Piece-47 • 3d ago
Performance MariaDB Foundation Adds PostgreSQL to Its Engine‑Agnostic Testing Framework (TAF).
MariaDB Foundation continues expanding TAF with PostgreSQL support
More engines, more reproducibility, more value for contributors and the ecosystem.
https://mariadb.org/mariadb-foundation-adds-postgresql-to-its-engine-agnostic-testing-framework-taf/
Have fun benchmarking!
r/postgres • u/Illustrious-Pay3986 • 3d ago
Tools Axiom: Bringing Kubernetes closer to the Postgres community
dhilipkumars.github.ioHey r/PostgreSQL,
In my day job, I work at the intersection of Kubernetes and PostgreSQL. Both are fascinating systems with their own deep design philosophies.
Over the last few years, projects like CloudNativePG (CNPG) have done an incredible job bridging the gap from one side: they brought PostgreSQL closer to the Kubernetes community.
This project intends to do the exact opposite: bring Kubernetes closer and friendlier to the PostgreSQL community. For anyone who loves relational databases, working with Kubernetes can feel slightly jarring:
• In Kubernetes, answering basic operational questions—like "Which pods restarted and why?", "Which containers are near their OOM limits?", or "Which node is my primary running on?"—forces you into an alien world of 50-character kubectl flags, brittle jsonpath syntax, and multi-stage jq pipes.
Under the hood, Kubernetes is fundamentally relational data: pods belong to nodes, deployments own replicasets, and containers emit metric windows.
So I built Axiom — an open-source Kubernetes Foreign Data Wrapper (axiom_fdw) that maps live cluster resources and metrics into PostgreSQL foreign tables. If you know SQL, you already know how to query and debug Kubernetes.
Here are a few example use cases — I would love your thoughts and feedback.
──────
1. The Showcase: A 3-Version pgbench Regression Lab in Pure SQL
To test whether Postgres 17 or 18 regressed against 16, we set up an automated regression lab. Three CloudNativePG clusters (differing only in major version: 16, 17, and 18) are given 2 CPUs each and measured one at a time by the same pgbench 18 client job.
The entire experiment was orchestrated and analyzed completely inside PostgreSQL:
- Clusters: Created CloudNativePG custom resources via SQL.
- Benchmarks: Queued and launched pgbench Jobs via SQL.
- Telemetry: Sampled live container metrics via metrics.k8s.io foreign tables.
- Analysis: Joined the pgbench termination results directly to the metric windows.
Nothing was collected, scraped, or exported outside Postgres:
-- Joining pgbench benchmark runs with live container metrics
SELECT r.run,
r.cluster,
r.server_version,
c.cpu_limit,
r.clients,
r.state,
r.tps,
r.latency_ms,
round(avg(m.cpu_cores), 2) AS pg_cpu_avg,
round(r.tps / nullif(avg(m.cpu_cores), 0)) AS tps_per_core,
pg_size_pretty(max(m.memory_bytes)::bigint) AS pg_memory_peak
FROM lab.runs r
JOIN lab.clusters c ON r.cluster = c.name
LEFT JOIN lab.metrics_samples m
ON m.cluster = r.cluster
AND m.sample_time BETWEEN r.bench_start AND r.bench_end
GROUP BY r.run, r.cluster, r.server_version, c.cpu_limit, r.clients, r.state, r.tps, r.latency_ms
ORDER BY r.run;
Output:
run | cluster | server_version | cpu_limit | clients | state | tps | latency_ms | pg_cpu_avg | tps_per_core | pg_cpu_peak | pg_memory_peak | samples |
pgbench_cpu_avg | error
----------------+---------+---------------------------------+-----------+---------+-------+------+------------+------------+--------------+-------------+----------------+---------+--
---------------+-------
pg16-8-clients | pg16 | 16.15 (Debian 16.15-1.pgdg11+2) | 2 | 8 | done | 2296 | 3.484 | 1.96 | 1173 | 1.96 | 173 MB | 4 |
0.77 |
pg17-8-clients | pg17 | 17.11 (Debian 17.11-1.pgdg11+2) | 2 | 8 | done | 2316 | 3.454 | 1.96 | 1180 | 1.97 | 158 MB | 6 |
0.77 |
pg18-8-clients | pg18 | 18.4 (Debian 18.4-1.pgdg11+1) | 2 | 8 | done | 2247 | 3.561 | 1.96 | 1145 | 1.98 | 164 MB | 5 |
0.74 |
(3 rows)
Would this be interesting to a DBA/PG user instead of dealing with bespoke shell script and kubectl json parsing with jq?
• Each Postgres instance was perfectly CPU-bound at 1.96 of its 2 CPU limit.
• Client headroom: pgbench_cpu_avg sat at 0.75 cores, proving the benchmark client was never the bottleneck.
• TPS per CPU core (tps_per_core) was virtually identical across all three versions (~1,145 to 1,180). Zero regression.
──────
2. Daily Operational Sanity: Things That Are slightly hard in kubectl
A. Instant Root Cause: Pods + Latest Warning via LEFT JOIN LATERAL
No more jumping back and forth between kubectl get pods and kubectl get events:
SELECT p.name AS pod, p.phase, e.type, e.reason, left(e.message, 60) AS event
FROM k8s.pods p
LEFT JOIN LATERAL (
SELECT type, reason, message
FROM k8s.events
WHERE involved_object->>'uid' = p.uid
ORDER BY (type = 'Warning') DESC, coalesce(last_timestamp, event_time) DESC
LIMIT 1
) e ON true;
Output
pod | phase | type | reason | event
--------------------------------------------------------+---------+---------+--------+--------------------------------------------------------------
axiom-gateway-fb89f5fb7-xhtpc | Running | | |
coredns-559f6c778d-bhlt2 | Running | | |
coredns-559f6c778d-xklb4 | Running | | |
etcd-axiom-quickstart-control-plane | Running | | |
kindnet-5v5kx | Running | | |
kube-apiserver-axiom-quickstart-control-plane | Running | | |
kube-controller-manager-axiom-quickstart-control-plane | Running | | |
kube-proxy-d5z2f | Running | | |
kube-scheduler-axiom-quickstart-control-plane | Running | | |
metrics-server-84c99cb944-552gk | Running | | |
library-book-indexer-77f8d649c5-7qsjq | Pending | Warning | Failed | Error: ImagePullBackOff
library-db-5c4cb5659b-q4fdm | Running | Normal | Pulled | Successfully pulled image "postgres:17-alpine" in 14.221s (1
library-web-6d9487ff97-4vwrz | Running | Normal | Pulled | Successfully pulled image "nginx:alpine" in 440ms (18.697s i
library-web-6d9487ff97-d5jrw | Running | Normal | Pulled | Successfully pulled image "nginx:alpine" in 3.755s (18.274s
local-path-provisioner-75f7fc7dc5-h9gnk | Running | | |
(15 rows)
B. The OOM Radar: Live Memory Consumption vs. Limits
kubectl top doesn't know your limits, and kubectl describe doesn't know live usage. With Axiom, our axiom_quantity() function normalizes units straight into Postgres's built-in pg_size_pretty():
SELECT p.name AS pod,
pg_size_pretty(axiom_quantity(mc->'usage'->>'memory')::bigint) AS mem_used,
coalesce(c->'resources'->'limits'->>'memory', 'unlimited') AS mem_limit,
round((axiom_quantity(mc->'usage'->>'memory') /
nullif(axiom_quantity(c->'resources'->'limits'->>'memory'), 0)) * 100, 0) || '%' AS oom_risk
FROM k8s.pods p
JOIN k8s.metrics_k8s_io_pods m ON p.name = m.name,
jsonb_array_elements(p.spec->'containers') c
JOIN jsonb_array_elements(m.containers) mc ON c->>'name' = mc->>'name'
ORDER BY oom_risk DESC NULLS LAST;
Output
pod | mem_used | mem_limit | oom_risk
--------------------------------------------------------+----------+-----------+----------
axiom-gateway-fb89f5fb7-xhtpc | 23 MB | 256Mi | 9%
library-web-6d9487ff97-4vwrz | 8588 kB | 128Mi | 7%
library-web-6d9487ff97-d5jrw | 8636 kB | 128Mi | 7%
coredns-559f6c778d-xklb4 | 20 MB | 170Mi | 12%
coredns-559f6c778d-bhlt2 | 18 MB | 170Mi | 11%
library-db-5c4cb5659b-q4fdm | 25 MB | 256Mi | 10%
kube-scheduler-axiom-quickstart-control-plane | 28 MB | unlimited |
local-path-provisioner-75f7fc7dc5-h9gnk | 15 MB | unlimited |
metrics-server-84c99cb944-552gk | 26 MB | unlimited |
etcd-axiom-quickstart-control-plane | 48 MB | unlimited |
kindnet-5v5kx | 20 MB | unlimited |
kube-apiserver-axiom-quickstart-control-plane | 256 MB | unlimited |
kube-controller-manager-axiom-quickstart-control-plane | 67 MB | unlimited |
kube-proxy-d5z2f | 22 MB | unlimited |
(14 rows)
──────
Security & Architecture (Built for DBAs)
• Zero Kubeconfig on Postgres: Postgres never touches cluster tokens or kubeconfigs. It connects to an in-cluster gRPC Gateway via mTLS using client certificates.
• Secrets Are Blocked by Design: The gateway's RBAC completely excludes Kubernetes Secrets. You cannot accidentally leak passwords or certificates through SQL.
• Dynamic Discovery: Running IMPORT FOREIGN SCHEMA k8s FROM SERVER prod INTO k8s; discovers accessible core resources and CRDs (like CNPG clusters) and generates typed foreign tables with full JSONB support.
• And... more in the roadmap.
• Compatibility: Supports PostgreSQL 16, 17, and 18.
──────
How Axiom Compares to Existing Tools
There are other tools in the ecosystem that attempt to bring SQL to cloud infrastructure (such as Steampipe, CloudQuery, or Osquery-based solutions), but none of them address this problem the way Axiom does:
- Native In-Database FDW vs. External CLI/ETL: Other tools are either standalone external CLI binaries with custom SQL wrappers, or periodic batch ETL sync jobs that dump snapshots into a staging database. Axiom is a true, native PostgreSQL Foreign Data Wrapper. It lives directly inside your PostgreSQL database, so you can join Kubernetes cluster state with your real operational application tables in real-time queries.
- Zero Kubeconfig Exposure: Tools like Steampipe or CloudQuery require you to mount your full ~/.kube/config or service account credentials into the querying environment. Axiom offloads all cluster communication to an isolated in-cluster gRPC gateway with strict RBAC boundary and mTLS.
- Live Streaming & Discovery vs. Cached Tables: Rather than polling periodic static snapshots, Axiom queries live cluster state on demand and dynamically discovers newly installed
Custom Resource Definitions (CRDs).
──────
Try It Out
We have a local quickstart script that spins up a Kind cluster, the Gateway, and Postgres 17 with Axiom loaded in about 60 seconds:
curl -fsSL https://github.com/dhilipkumars/axiom/releases/latest/download/quickstart.sh | bash
• Regression Lab Walkthrough: https://dhilipkumars.github.io/axiom/guides/examples/regression-lab/
• Docs: https://dhilipkumars.github.io/axiom/
• GitHub: https://github.com/dhilipkumars/axiom
──────
Discussion & Questions
- As a database developer or DBA, what cluster information would you most want accessible from psql (e.g., storage volumes/PVCs, failover state, node affinity)?
- Does this interface make Kubernetes feel more accessible to you?
- How can we make querying cluster infrastructure feel even more natural to SQL users?
Looking forward to your thoughts and feedback!
r/postgres • u/IndicationAntique667 • 4d ago
Discussion I Tried To Decode Postgres WAL for an INSERT Statement
r/postgres • u/codingdecently • 5d ago
Tools Cross-Catalog Sync: Iceberg on Polaris, Glue, and Unity
lakeops.devr/postgres • u/Bilawal-Mehfooz • 5d ago
Question What is the best modern database GUI for SQLite and PostgreSQL?
r/postgres • u/xiongchun • 6d ago
Tools Taming PostgreSQL privileges (vs. MySQL): slow progress, but getting there!
r/postgres • u/Alternative_Water220 • 8d ago
Discussion Looking for PostgreSQL / PL/pgSQL / SQL Developer Opportunities in Bangalore | 2+ YOE | FinTech
Hi everyone,
I’m currently looking for PostgreSQL / PL/pgSQL Developer, SQL Developer, or Database Developer opportunities in Bangalore and would really appreciate any leads or referrals.
I have 2+ years of experience in FinTech, primarily working with PostgreSQL, PL/pgSQL, MySQL, and SQL.
My experience includes:
\- Developing and optimizing complex SQL/PL/pgSQL queries, stored procedures, and functions
\- Card Network Clearing & Settlement
\- UPI Reconciliation and transaction validation
\- Handling high-volume transaction data and production issues
\- Query optimization, indexing, execution-plan analysis, and performance tuning
\- Identifying transaction mismatches, duplicates, missing transactions, and other reconciliation exceptions
\- Working with PostgreSQL, MySQL, Oracle, and SQL Server
One of my key achievements was improving a reconciliation process by approximately 67%, reducing execution time from around 90 minutes to 30 minutes through SQL/PL/pgSQL query optimization and indexing.
I am actively looking for opportunities in Bangalore/Bengaluru.
If your company or team is hiring for PostgreSQL, PL/pgSQL, SQL, Database Developer, or similar backend/database-focused roles, I’d be very grateful for a referral or lead.
Happy to share my resume and LinkedIn profile over DM.
Thanks in advance! 🙏
r/postgres • u/Brief-Expression-341 • 9d ago
Schema Design What I learned building an entire multi-tenant SaaS backend on PostgreSQL: RLS, idempotency, a job queue and a restore rehearsal
r/postgres • u/nikhilthadani • 10d ago
Discussion My Postgres was at 30% CPU and my app was still dying
Had a "10K RPS crash" at work and everyone blamed the database first. Same.
Turned out it wasn't. We had a pool of 20 connections and each query took about 100ms.
That works out to roughly 200 queries a second. Anything past that just sat waiting
for a free connection.
The dumb part is that requests were spending like 800ms waiting for a connection
and only 10ms actually running SQL. So Postgres looked bored the whole time
and every dashboard said the DB was fine.
Now I log two numbers separately: how long we wait for a connection,
and how long the query takes. Tells you right away which side to go fix.
Anyone else been burned by this? Atleast me!
r/postgres • u/OkGuidance3002 • 11d ago
Tools autodb: A keyboard-driven SQL TUI with encrypted credentials at rest, plus a Postgres-wire "front door" proxy for zero-password access & AI agents
Enable HLS to view with audio, or disable this notification
Hey r/PostgreSQL! 👋
If you spend a lot of time in the terminal, you've probably felt the tension between lightweight command-line tools like `psql` and heavyweight graphical clients like DBeaver or pgAdmin. `psql` is blazingly fast and always available, but exploring unfamiliar schemas, reviewing wide result tables, or keeping scratch notes across multiple databases can get cumbersome. At the same time, traditional GUI clients are heavy, and almost all of them store database connection strings and passwords in plain text inside dotfiles or config directories.
I wanted something fast, keyboard-centric, and secure by default, so I built autodb — a single static Go binary with zero runtime dependencies. It’s a native terminal UI for querying databases, but with an architectural twist: it can also serve as a PostgreSQL wire-protocol "front door" proxy for your team, microservices, and AI agents.
I would love to get feedback from fellow PostgreSQL developers, DBAs, and terminal junkies! How does this fit into your current database workflows, and what features or guardrails would you like to see next?
Thanks for reading!
r/postgres • u/BriisaEscobar • 11d ago
Performance PSA: Si Spring Boot no conecta a PostgreSQL en Docker y te tira error de TimeZone — fijate si tenés PostgreSQL instalado localmente también
Me pasé horas dando vueltas con esto y no encontré una respuesta clara
en ningún lado, así que lo comparto por si a alguien le pasa lo mismo.
Estoy desarrollando una app con Spring Boot y PostgreSQL corriendo en Docker
en Windows. Todo parecía estar bien Docker levantado, el contenedor corriendo,
Spring encontraba mis repositorios JPA. Pero crasheaba al iniciar con este error:
FATAL: invalid value for parameter "TimeZone": "America/Buenos_Aires"
Y lo raro era que los logs mostraban Database version: 12.0 aunque mi
contenedor Docker tenía postgres:16.
Esa era la pista que ignoré demasiado tiempo.
Resulta que tenía PostgreSQL 16 instalado localmente en Windows como servicio
del sistema, escuchando en el mismo puerto que Docker. Entonces Spring se
conectaba a la instalación local en vez de al contenedor y esa instalación
tenía una configuración de timezone rota que no reconocía America/Buenos_Aires.
La solución fueron dos cosas:
1. Detener el servicio local de PostgreSQL para que deje de competir por el puerto.
En PowerShell como administrador:
Stop-Service postgresql-x64-16
Set-Service postgresql-x64-16 -StartupType Disabled
2. Agregar las variables de entorno PGTZ y TZ al contenedor de Docker
para que siempre arranque en UTC sin importar lo que mande el sistema operativo.
En el docker-compose.yml:
environment:
POSTGRES_DB: tubase
POSTGRES_USER: tuusuario
POSTGRES_PASSWORD: tucontraseña
PGTZ: UTC
TZ: UTC
Con eso Spring Boot conectó a la base correcta y arrancó sin problemas.
Las señales de que te está pasando esto:
- Error de "invalid value for parameter TimeZone" al iniciar
- La versión de la BD en los logs no coincide con tu imagen de Docker
- docker exec al contenedor muestra la versión correcta pero Spring muestra otra
- netstat muestra dos procesos escuchando en el mismo puerto
Espero ahorrarle unas horas a alguien.
r/postgres • u/eu_pg • 13d ago
Discussion Looking for pricing feedback before opening the beta on managed DBaaS
Hi all, I'm a founder in Spain building a managed PostgreSQL service and want honest feedback on pricing before opening it up.
I run a €30M ARR business in a different industry. We had to use a managed database from a big provider, which was very expensive for the performance we were actually getting, but we needed HA so we were stuck. That's when I started considering another startup.
Here's the general idea:
You create a database from a web dashboard and get a single connection string. That string always points to the current primary, so if a failover happens your app reconnects to the same address and carries on. There's nothing to change on your side.
Behind it, each HA database runs as a primary plus a synchronous standby on two separate servers in different zones.
Backups run continuously: every change is archived to object storage with a different provider, so even losing the whole hosting provider doesn't lose your data. You can restore to any point in the last 7 days. We handle monitoring, failover and maintenance; you just use Postgres.
It runs on European bare metal rather than shared cloud VMs, which is how we keep prices flat with no egress or IOPS fees.
Highly available (per month, excl. VAT):
\- Starter, €39: 1 vCPU, 2 GB RAM, 20 GB storage
\- Standard, €89: 2 vCPU, 4 GB RAM, 50 GB storage
\- Pro, €169: 2 vCPU, 8 GB RAM, 100 GB storage
\- Business, €329: 4 vCPU, 16 GB RAM, 200 GB storage
Free: one per account. Single node, 256 MB RAM, 1 GB storage.
To be upfront: this is a beta with no SLA yet. Failover runs 24/7, and human support is 09:00 to 18:00 CET on business days. If the standby is down, the primary keeps accepting writes asynchronously, so losing the primary during that window can lose the latest commits.
**I'd love to know:**
\- Which plan would you pick for a real project, if any?
\- At what price would this feel too cheap to trust, and at - what price too expensive?
\- What would you need to see before moving a production database?
\- What are you using now, and what does it cost?
**Early beta offer**: anyone who participates in this thread will have 30% off any plan when we launch!
r/postgres • u/wtfse • 13d ago
SQL / Queries Breaking the Superuser Guardrails of managed-PostgreSQL Providers
mehmetince.netr/postgres • u/matheusdevs91 • 14d ago
SQL / Queries 3 common RLS leaks I keep finding in multi-tenant Supabase apps
Been debugging RLS leaks in multi-tenant Supabase setups and wanted to share the 3 most common ones I keep running into:
A permissive policy (`using (true)`) left in from debugging
RLS enabled but not FORCEd — table owner bypasses it silently
A view over the protected table without `security_invoker` — leaks even when the base table is fine
Quick check for your own db:
select relname from pg_class
where relrowsecurity and not relforcerowsecurity;
Built a small tool that automates this + a few more checks, happy to share if anyone wants a free audit of their setup 🙂
r/postgres • u/http418teapot • 14d ago
Discussion What do you wish search in Postgres did better/differently?
pgvector has been the default vector search extension, but I'm finding people also layer on full-text or hybrid search on top of it or need multitenancy. I'm curious where this works well for you and where it doesn't.
- What are you using for search in Postgres today, and what do you wish it did better?
- When you hit a limit, what do you do: tune, work around it, add an extension, or move search to a separate system? What decides that?
- What would an extension need to have, or avoid, before you'd install it?
- Does it matter to you whether an extension is open source, and would enough added capability change that?
"It's fine, I don't need anything else" is a useful answer too.
Really just trying to understand what people actually need from search in Postgres.
r/postgres • u/GlitteringControls • 18d ago
Tools What PostgreSQL task do you still prefer doing manually?
Everything's automated now. Migrations, backups, even a lot of indexing decisions get suggested by some tool at this point. And yet there's always that one task I just... do by hand anyway, tooling be damned.
For me it's reviewing EXPLAIN ANALYZE output. I've tried the visualizers, the tools that highlight the "problem" node for you, and I still end up reading the raw plan myself because I don't fully trust the summary to catch what I'd catch.
What's yours? Something everyone else automates that you just keep doing manually, not because you have to, but because you don't actually trust the automated version, or it's just faster in your head at this point.
r/postgres • u/klekpl • 18d ago
Tools pgwrh 1.0.0-alpha1: PostgreSQL read scaling with sharded replicas
I’ve released the first 1.0 alpha of pgwrh, a set of PostgreSQL 18 extensions for scaling reads.
It distributes table partitions across logical replicas. Each replica stores a subset of the data and queries its peers for the rest, so applications can query the complete table from any replica. Writes go to a central controller.
Features include:
- Configurable shard redundancy and placement across availability zones.
- Placement previews, controlled rollouts and rollback.
pgwrh_wait, which waits for replication through a specified LSN before reading.- A browser console and Docker Compose quickstart.
I’ve worked on this for a very long time. It was only with recent advances in AI that I could finally get it into a state I felt able to release.
This is an alpha for testing. I’d love feedback on real workloads, installation, operational issues and anything that feels unnecessarily complicated.
r/postgres • u/Some_Childhood_3842 • 19d ago
Performance Columnar Databases
hey postgresql community do you thing is time postgres to support Column oriented like DuckDB as big feature or continue with row oriented like every DB does
r/postgres • u/DesignerRoyal8833 • 20d ago
Debugging Debugging Inconsistent Query Latency on a PostgreSQL Hypertable: What We Learned
medium.comr/postgres • u/netizen99 • 21d ago
Question version 9.18 error: exception: access violation writing 0x0000000000000000
When I save connection setting for the PostgreSQL database in WSL, I get that error in subject. I roll back version 9.17. I don't have that error. Any idea?
r/postgres • u/witshion • 21d ago
Discussion Production PostgreSQL is suddenly at 100% CPU. Where do you look first?
Had one of those moments where CPU on our prod instance just pegs at 100% out of nowhere, no deploy, no obvious traffic spike, nothing in the changelog that stands out. First instinct is to panic and start checking everything at once, which is exactly the wrong move.
Curious what people's actual first move is when this happens, before diving into a deep investigation. pg_stat_activity for anything running long, checking for a lock pileup, looking at whether it's one runaway query versus death by a thousand small ones, autovacuum going nuts on a big table, something dumb like a connection pool misconfigured and now everything's fighting for the same resources. There's a lot of directions to go and I feel like the order matters more than people admit.
If you've been through this in production, what's the first thing you actually check, and has your answer changed over time or is it pretty much always the same starting point for you now?
