$RodHat_
An empty office chair under a desk lamp at night

Rod's Tales

War stories. His own fuckups included, especially his own.

The listen queue had been five since 2014

The listen queue had been five since 2014

Nine years of incrementing connection drops, masked by retry logic added in 2017 and promptly forgotten. The counter was in netstat -s. The backlog was in ss -tlnp. Both said the same thing for nine years.

Editorial card: The ephemeral port range had eight hundred slots left

The ephemeral port range had eight hundred slots left

The inventory service was throwing EADDRNOTAVAIL on database connections. CPU fine. Memory fine. The connection pool was configured exactly as intended. That was the problem.

The conntrack table had twenty-four slots left

The conntrack table had twenty-four slots left

The load balancer was healthy. CPU fine. Memory fine. One in thirty connections was timing out and nobody could explain why. The answer was in dmesg, six lines down, logged and completely ignored.

The transaction counter hit two billion

The transaction counter hit two billion

Production PostgreSQL stopped accepting writes on December 21st. Disk fine. CPU fine. Connections fine. The error said something about 'wraparound data loss.' I had seen it once before and I knew immediately how bad the next six hours were going to be.

The load was forty and the box was idle

The load was forty and the box was idle

Load average alert fired at 47.8 on an eight-core box. SSH'd in, ran top, saw 95% idle. Every metric looked healthy except the number that triggered the page. Thirty-two processes in D state, waiting on an NFS mount that had stopped answering three hours earlier.

The machine that forgot what year it was

The machine that forgot what year it was

A dead CMOS battery on a physical box meant the hardware clock reset to January 1, 2000 on every reboot. NTP would correct it in about ninety seconds. Everything that happened during those ninety seconds was wrong in ways that took hours to untangle.

The disk was forty percent free

The disk was forty percent free

Production throwing ENOSPC. df shows 40% free. You run df again because you don't believe it. Same numbers. The filesystem is completely full and completely not full at the same time, and once you understand why, you will never run df without -i again.

The process that wouldn't die

The process that wouldn't die

kill -9 is supposed to be final. You send it, the process is gone. Then one morning you send it and nothing happens, and you send it again, and the process is still sitting there in ps, and you have to go rebuild an assumption you've had for twenty years.

The conntrack table was full

The conntrack table was full

Connections to services behind our firewall started failing intermittently. The iptables rules were correct, the NICs were clean, the routing was fine. The problem was a kernel table we had never configured, silently dropping packets when it ran out of room.

The TCP window that ate a gigabit

The TCP window that ate a gigabit

The cross-datacenter link was provisioned for 1Gbps. iperf3 consistently showed 180Mbps. The hardware was clean, the routing was clean, the fiber was clean. The problem was a 208KB number that nobody had changed in years, and the unforgiving math of physics.

The disk that du couldn't find

The disk that du couldn't find

df said 97% full. I walked every directory with du and came up 40GB short. The filesystem was not corrupted. The math was just wrong in a way that took embarrassingly long to recognize for what it was.

The file that root couldn't touch

The file that root couldn't touch

A log directory was filling up. Deletion failed with EPERM. Root was confirmed root. Permissions were correct. The mount was read-write. It took an embarrassingly long time to remember that root is not omnipotent, and that two years earlier someone had applied a hardening script.

The pool was full of dead connections

The pool was full of dead connections

An intermittent burst of broken-pipe database errors had been logged as "flaky" for two months before anyone looked closely enough to notice they were always the first query on a connection. The connection pool was handing out corpses.

The OOM killer was doing its job

The OOM killer was doing its job

A slow memory leak ran undetected for five weeks because the kernel's out-of-memory killer, executing its heuristic correctly, kept choosing the monitoring agent over the leaking service. The pager never fired. The monitoring gaps were there in the data the whole time.

The number 1024 was not a coincidence

The number 1024 was not a coincidence

A file descriptor limit from 2003 survived three infrastructure migrations and one complete platform rewrite because nobody ever questioned a number that had always been there. It bit us at 2am on a Tuesday.

The pipeline succeeded. We found out from a customer.

The pipeline succeeded. We found out from a customer.

A deployment that broke production ran undetected for forty minutes because the notification step was written to never fail, and the monitoring alert that should have fired went to the same broken endpoint.

The postrotate script worked. We had no logs.

The postrotate script worked. We had no logs.

Three months of production logs, all gone. logrotate ran clean on schedule every week. The daemon was healthy and writing the whole time. The postrotate signal was firing correctly, just to a PID that had been dead since deployment day.

a close up of a rack of computer equipment

The storage was fine. We were swapping.

Three weeks, two engineers, one open storage vendor ticket, and roughly forty collective hours chasing read latency on a ZFS pool. The pool was fine the whole time. There was a swapfile nobody remembered adding seven months earlier.

The change was correct. Deploying it to every region simultaneously was not.

The change was correct. Deploying it to every region simultaneously was not.

A one-line config change, reviewed, tested, and correct in every environment we tried it in. It went to all six regions in the same rollout because config isn't code, and config doesn't need a staged rollout. It does now.

The alert had been firing for eight months and it was right the whole time

The alert had been firing for eight months and it was right the whole time

A rule that pages nightly and gets acknowledged nightly isn't monitoring, it's a ritual. We had trained an entire team to dismiss a specific alert without reading it, and then it started telling us about something new.

The observability platform was down during the outage. Good. Now do it with your hands.

The observability platform was down during the outage. Good. Now do it with your hands.

A junior engineer froze during a production fire because the SaaS dashboard that watches production was part of the fire. RodHat on the tools that were on every Unix box before the kid was born, the racket that sold competence back to us as a monthly invoice, and why the fire isn't out.

The health check was green because it was checking the wrong damn thing

The health check was green because it was checking the wrong damn thing

A service can answer HTTP 200 while its queue is wedged, its database writes are failing, and every useful request is dying. RodHat on health checks that prove process existence instead of service capability.

Staging had its own everything, except the one thing that mattered

Staging had its own everything, except the one thing that mattered

Separate servers, separate config, separate deploy pipeline, a big banner saying STAGING. And a database connection string that pointed, through two layers of indirection, at production. We found out during a load test.

We wrote the retry logic to survive a blip. It turned a blip into four hours.

We wrote the retry logic to survive a blip. It turned a blip into four hours.

One backend got slow for ninety seconds. Every client retried, in lockstep, three times, with no jitter. The retries were larger than the original traffic, the backend never recovered, and every fix we tried made it worse until we did the thing nobody wanted to do.

The temporary NFS mount that ran production for six years

The temporary NFS mount that ran production for six years

Somebody stood it up in an afternoon to unblock a launch. It had no monitoring, no owner, no backup, and no entry in any diagram. It survived two datacentre moves, a company acquisition, and every engineer who knew it existed.

The billing job ran twice for six weeks and the totals still balanced

The billing job ran twice for six weeks and the totals still balanced

No errors. No alerts. The reconciliation report matched every single day. And a few hundred customers were being double-charged, because the job that generated the charges and the job that verified them were the same code with the same bug.

The migration was flawless. The TTL was 86400.

The migration was flawless. The TTL was 86400.

Six weeks of planning, a rehearsed cutover, and a maintenance window we finished forty minutes early. Then a fifth of our traffic kept arriving at a datacentre we'd already started decommissioning, for a full day, and there was nothing whatsoever we could do about it.

a close-up of a server room

The microservice that should have been a function call, and the eighteen months it took anyone to say so

An architecture rant, grounded in a real bad design RodHat watched calcify for a year and a half before anyone had the standing to kill it.

The restore worked perfectly and we still lost eleven days

The restore worked perfectly and we still lost eleven days

We tested our backups. Monthly, documented, signed off. The restore ran clean, the checksums matched, the database came up on the first try. And the data in it was eleven days old, because for eleven days the job had been backing up a directory nobody was writing to any more.

a rack of servers in a server room

The Friday deploy that ate my weekend, and whose fault it actually was (mine)

A sysadmin war story about a Friday-afternoon deploy, a silent DNS TTL assumption, and the two-day outage it caused. RodHat owns every part of it.