Linux infrastructure engineer. Production estates and platform migrations for banking, government and telecoms.
Thirty years in IT, more than twenty-five on Unix and Linux. Red Hat and Ubuntu at production scale, containerised platforms, high availability and disaster recovery, Ansible, and the migrations that move an estate from one to the other.
RHEL · Ubuntu · Docker · Kubernetes · Ansible · PostgreSQL · AWS · Azure · VMware · Nginx · Caddy · HAProxy · Prometheus · Grafana · ELK · TLS and PKI · SPF, DKIM and DMARC
Post-mortem
534,686 delivery attempts. Zero messages delivered. Every dashboard reported the system healthy.
A self-hosted Postal mail server, running in Docker on a VPS, had been accepting mail and queueing it normally for months. Messages went in. The queue looked healthy. Nothing came out the other side, and nobody noticed, because the layer that was broken was not the layer being watched.
Every attempt in the database carried the same failure:
Resolv::ResolvError
The obvious check is to resolve a domain from inside the container. I did, in Ruby, the same library Postal uses. It returned in 0.03 seconds.
That result was a false negative, and it cost real time. It had already produced a plausible and completely wrong theory, that name resolution was failing under concurrency and succeeding when tested by hand.
The manual lookup read /etc/resolv.conf. The worker read the resolver configured in
Postal's own config file. Two different files. The diagnostic and the failure were never
looking at the same thing.
Postal's configured resolver was 10.0.1.1. That address is the gateway of a Docker
bridge network, br-e538c46466eb. It is a routable address, so nothing errors
immediately, and it answers on no port at all. Nothing listens on 53.
Every MX lookup the mail server had ever made timed out against a gateway that was never a resolver. The config file was dated 8 May. The failures ran from then until September.
127.0.0.11, Docker's embedded DNS, which also resolves
container names, with 1.1.1.1 and 8.8.8.8 behind it.timeout:2 attempts:2, so a resolver failure fails fast instead of holding a
worker open.The nightly job ran every night. It reported success every night. Nothing had been uploaded since 15 August. I found out on 10 September.
The cause was not technical. A free-tier cloud account lapsed and the provider suspended it. The backup script carried on running to schedule, the uploads failed, and the failure was not loud enough to reach anyone.
Every alert in the chain was watching the wrong thing. They were all asking did the job run? The job ran. It ran perfectly, twenty-seven times, to a destination that was no longer accepting anything.
Nothing was asking the only question that matters: how old is the newest backup?
Checking the restore path afterwards, one leg was invisible on the first pass because its cron
runs as root rather than as the service user. Reading one crontab is not reading the
backups. Every schedule on the host is part of the estate, whoever owns it.
Nothing was lost. The historical snapshots were still there because the sixty-day prune had not yet reached them. That is timing, not design. Four more weeks and the outcome is a different article.
Two ways in. The price is on both, because you should not have to ask.
Start here
You describe what it is doing and when it started. I read the logs and the config, reproduce what I can, and send you a written diagnosis within one working day: the actual cause, the actual fix, and what the fix will cost.
You can hand that to your own team and never speak to me again. That is a legitimate outcome and the price assumes it.
Paid when you have the diagnosis, not before.
The work itself
Production infrastructure: Red Hat and Ubuntu, Docker, Postgres, high availability, disaster recovery, TLS and certificate automation, and the migrations that move an estate from one place to another. Thirty years of it, across banking, government, healthcare and telecoms.
Direct engagement, invoiced through my limited company or an umbrella. Fully remote, four weeks' notice.
A monitoring retainer with a four-hour response window. I keep my availability capped, so I cannot honestly promise to be at a keyboard at 2am, and an incident SLA I cannot meet is worth less than no SLA at all.
If what you need is someone on a rota, you need a team, and I will say so.