Notes from running PostgreSQL in production. I'm the architect of SupraScan, a blockchain indexer running 500–1,000 TPS on a database that grew to 10 TB before I cut it to 2.7 TB.
Content Type
Our PostgreSQL blockchain indexer's TOAST table grew past 5TB because unbounded JSONB arrays kept getting bigger. Here's what TOAST was costing us, how to measure it, and how we moved hot JSONB fields into normal columns without taking the indexer offline.
An UPDATE returns in two milliseconds, but underneath, Postgres just ran WAL logging, checkpoint bookkeeping, wait-event tracking, vacuum, and ANALYZE to make that write durable, recoverable, and still fast to plan. Here's what each of those five processes actually does, and which one is usually behind the latency spike that has no slow query in sight.
ORMs like Prisma and TypeORM don't lie to you about your database, they just don't tell you everything. Here's where that gap shows up in production: N+1 queries, migrations that silently drop columns, connection pool exhaustion in serverless, and the SQL your ORM actually generates versus what you think it's running.
An order gets saved to Postgres. The event that's supposed to tell payments, shipping, and notifications never goes out, because the process crashed one line later. Here's why that gap exists, and the two real patterns (Outbox, Sagas) that close it.
Search gets slow, someone says 'we need Elasticsearch,' and two weeks later there's a new cluster, a sync pipeline, and a class of bugs that didn't exist before. PostgreSQL's tsvector, GIN indexes, and pg_trgm cover the full-text and fuzzy-matching workload most teams actually have, without a second database to keep in sync. This is the case for starting there and adding Elasticsearch only when the workload actually demands it.
A deploy that used to take ninety seconds now takes six minutes, and nobody changed the code. The real story is in how Docker layers, build caching, and registries accumulate weight that never comes back off on its own, plus how to actually measure it before you guess.
A write-heavy ingestion table on Postgres starts choking under load: autovacuum can't keep up, WAL grows fast, and every insert costs more than it should. A Cassandra table doing the same job barely notices. The difference isn't tuning. It's a decision made before either database wrote a single line of code: B-Tree or LSM-Tree.
Your habits as an engineer aren't shaped by courses and effort alone; they're shaped by the PRs, teammates, and incidents you sit next to every day.
For twenty years, every PostgreSQL process read one disk page at a time and waited. PostgreSQL 18 finally lets it ask for many pages at once. Here's what changed under the hood, and how to verify the gain on your own workload.
New posts and case studies on PostgreSQL internals and production incidents, plus a short digest when new course modules go live. No spam, unsubscribe in one click.
Prefer a feed reader? Follow via RSS