Byron Coke

Linux infrastructure engineer. Production estates and platform migrations for banking, government and telecoms.

contracts@paxmentis.com ·Remote, UK ·Available at four weeks' notice

Thirty years in IT, more than twenty-five on Unix and Linux. Red Hat and Ubuntu at production scale, containerised platforms, high availability and disaster recovery, Ansible, and the migrations that move an estate from one to the other.

RHEL · Ubuntu · Docker · Kubernetes · Ansible · PostgreSQL · AWS · Azure · VMware · Nginx · Caddy · HAProxy · Prometheus · Grafana · ELK · TLS and PKI · SPF, DKIM and DMARC

A recent diagnosis, written up in full

Post-mortem

Half a million delivery attempts, nothing ever sent, and a test that passed

534,686 delivery attempts. Zero messages delivered. Every dashboard reported the system healthy.

A self-hosted Postal mail server, running in Docker on a VPS, had been accepting mail and queueing it normally for months. Messages went in. The queue looked healthy. Nothing came out the other side, and nobody noticed, because the layer that was broken was not the layer being watched.

Every attempt in the database carried the same failure:

Resolv::ResolvError

The test that made it worse

The obvious check is to resolve a domain from inside the container. I did, in Ruby, the same library Postal uses. It returned in 0.03 seconds.

That result was a false negative, and it cost real time. It had already produced a plausible and completely wrong theory, that name resolution was failing under concurrency and succeeding when tested by hand.

The manual lookup read /etc/resolv.conf. The worker read the resolver configured in Postal's own config file. Two different files. The diagnostic and the failure were never looking at the same thing.

The cause

Postal's configured resolver was 10.0.1.1. That address is the gateway of a Docker bridge network, br-e538c46466eb. It is a routable address, so nothing errors immediately, and it answers on no port at all. Nothing listens on 53.

Every MX lookup the mail server had ever made timed out against a gateway that was never a resolver. The config file was dated 8 May. The failures ran from then until September.

The fix

What I take from it

Selected work

Twenty-seven days with no backup, and nothing said a word

The nightly job ran every night. It reported success every night. Nothing had been uploaded since 15 August. I found out on 10 September.

The cause was not technical. A free-tier cloud account lapsed and the provider suspended it. The backup script carried on running to schedule, the uploads failed, and the failure was not loud enough to reach anyone.

Why it ran for four weeks

Every alert in the chain was watching the wrong thing. They were all asking did the job run? The job ran. It ran perfectly, twenty-seven times, to a destination that was no longer accepting anything.

Nothing was asking the only question that matters: how old is the newest backup?

The part I nearly missed

Checking the restore path afterwards, one leg was invisible on the first pass because its cron runs as root rather than as the service user. Reading one crontab is not reading the backups. Every schedule on the host is part of the estate, whoever owns it.

Honest postscript

Nothing was lost. The historical snapshots were still there because the sixty-day prune had not yet reached them. That is timing, not design. Four more weeks and the outcome is a different article.

What I take from it

Working with me

Two ways in. The price is on both, because you should not have to ask.

Start here

A written diagnosis · £250, fixed

You describe what it is doing and when it started. I read the logs and the config, reproduce what I can, and send you a written diagnosis within one working day: the actual cause, the actual fix, and what the fix will cost.

You can hand that to your own team and never speak to me again. That is a legitimate outcome and the price assumes it.

Paid when you have the diagnosis, not before.

The work itself

Engineering time · £600 per day

Production infrastructure: Red Hat and Ubuntu, Docker, Postgres, high availability, disaster recovery, TLS and certificate automation, and the migrations that move an estate from one place to another. Thirty years of it, across banking, government, healthcare and telecoms.

Direct engagement, invoiced through my limited company or an umbrella. Fully remote, four weeks' notice.

What I will not sell you

A monitoring retainer with a four-hour response window. I keep my availability capped, so I cannot honestly promise to be at a keyboard at 2am, and an incident SLA I cannot meet is worth less than no SLA at all.

If what you need is someone on a rota, you need a team, and I will say so.