How savepoints quietly throttled our Postgres queue
Blog post from Incident.io
Incident.io improved the performance of its Postgres-backed queue for on-call escalation processing, which handles roughly 15 million state checks daily, after load tests exposed a ceiling where acquisition queries became slower than the work itself and database CPU could not be fully utilized. The team found that accumulated dead tuples increased scans of the queue table, prompting them to move queue data from a large escalations table to a much smaller escalation_jobs table. A more significant bottleneck came from using Postgres savepoints for individual jobs within a batch transaction: the resulting MultiXact metadata caused contention on an internal LWLock as workers used SKIP LOCKED to scan rows held by others. They eliminated these subtransactions by leasing jobs through a claimed_until field and processing each job in its own transaction, accepting a rare potential delay of up to 10 seconds after a process failure. To reduce write amplification and cleanup overhead, claimed_until was intentionally left unindexed so updates could use heap-only tuple updates, with page fillfactor adjusted to preserve room for them. Finally, each application pod adopted a single dispatcher that claims jobs and supplies workers through a buffered channel, reducing competing queue scans and providing natural backpressure. Together, these changes increased peak escalation throughput by 3.5 times and shifted the remaining limit toward conventional database CPU scaling rather than internal lock contention.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.