Postgres locks do not scale
On a normal day our Postgres database is at a cpu load of 20-40%. March 24th was not a normal day.
Suddenly, for no discernible reason, our database was slammed at 100% CPU. All of it in system time. Our bots couldn’t join calls, customers couldn’t query our API, dashboards were unresponsive, and we couldn’t even get a psql into the database.
The usual suspects were considered first: there were no new deploys that morning, no current AWS incidents, swap usage was nominal, and all metrics before ...
Read more at recall.ai