building vodou.

rsync --delete from a stale build/ rolls production back

A local build/ folder three policy versions old nearly overwrote live files and deleted server-only ones. Why dry runs miss it, and a five-minute check.

Chad Priest / / 11 min read

On 2026-07-30 I was publishing version 1.4 of our privacy policy. The file on my laptop, frontend/build/privacy.html, said v1.1. The server was already serving v1.3.

The deploy command I was about to run pushed the whole build/ directory with rsync --delete. It would have put v1.1 over v1.3, and it would have deleted every file on the server that my laptop didn’t have. rsync wouldn’t have printed an error. The rolled-back privacy.html would have kept returning 200, just with older legal text on it, and that’s the kind of wrong nobody reports.

I can’t give you the number of files it would have deleted. I never ran the dry run that prints it, and build/ has been rebuilt since, so that count is gone. That’s my mistake, and it’s part of the point of this post. I do know some of the files that were in the set, because the deploys the day before left them there on purpose: privacy.html.bak-v1.1 and privacy.html.bak-v1.2, the backups taken before each policy went live; index.html.bak-20260729-155530; and the previous JS bundle, main.44d2bb50.js, which was left on disk so a bad frontend release could be rolled back by swapping one file. None of those existed on my laptop. Those deletions would have 404’d and shown up in nginx’s access log. Nobody requests a rollback file until the day they need it, though, and by then it would have been gone.

It didn’t ship. This post covers why it nearly did, why the usual advice (“dry-run first”) wouldn’t have caught it on its own, and a check you can run against your own deploy in about five minutes.

build/privacy.html said v1.1. The server was serving v1.3.

The app is a Create React App frontend. The layout matters here, because it’s common and it’s where the bug lives:

  • A generator writes the policy to frontend/public/privacy.html.
  • CRA copies public/ into build/ when you run npm run build.
  • nginx serves build/.

So the file the generator updates isn’t the file that gets served. To get a new policy live you either rebuild the whole app or copy the one file across. Nobody rebuilds a React app to publish an HTML page.

Here’s how v1.2 and v1.3 got to production. Both went live on 2026-07-29, the day before, from other working sessions on the same machine. Neither one touched my build/. Each uploaded the new page, backed up the live file on the server as privacy.html.bak-v<old version>, and put the new one in place as root with install -m 644. As a single-file deploy that’s sound. It also changed production twice through a path that left no trace in the directory my mirror deploy was about to read from.

So there are two halves to the bug. The artifact was stale, and the reason it was stale is that production had a second write path the artifact knew nothing about. If you’re wondering which one is your risk: it’s both. Every side-channel write to the server widens the gap between what’s live and what’s in your build output.

Nothing about that gap is visible from a terminal. ls build/ looks the same whether the contents are from yesterday or from April.

I thought build/ was a copy of production

I’ll admit the wrong belief first, because it’s the whole bug. I treated build/ as “what’s live”. It’s a folder named after an action, sitting next to the source and full of production-looking files. Mentally, deploying it just meant syncing prod with itself plus my change.

That belief is exactly what a mirror deploy encodes. rsync -a --delete src/ dest/ means “make dest identical to src”. rsync doesn’t read that as a request to upload your changes. It reads it as “src is the truth”. If src is old, rsync makes production old, and it’s very good at doing that.

The deletions bother me more than the rollback. A rolled-back privacy.html is at least a file I was looking at. The deletions hit files I wasn’t thinking about at all: the backups other deploys had made, the rollback bundle, and anything else written straight to the server that never passed through my laptop. A fresh build from the newest commit wouldn’t have saved those either. They were never in any commit.

Why rsync -n would not have told me v1.1 was older than v1.3

The standard advice is correct and I follow it now. The Margrop post on rsync dry runs puts it well: split every risky sync into a plan and an execution, use --itemize-changes so each line tells you what kind of change it is, and treat deletion as a separate, explicit permission. The sysax write-up on rsync mirroring adds the other half: “the failure is not knowing which one you are running”, mirror or accumulating copy.

Both are right, and neither covers the case where you know perfectly well you’re running a mirror and the source is simply old.

Here’s what the itemized dry run shows for a file that changed:

>f.st...... privacy.html

That’s the line for v1.3 → v1.4, and it’s also the exact same line for v1.3 → v1.1. s means the size differs and t means the mtime differs. Neither one tells you which side is newer. A dry run tells you that something will change. It doesn’t tell you whether the change goes forward or backward.

The same holds for every tool in this family. Hugo’s deploy docs describe it accurately: compare names, then sizes and MD5s; “any difference triggers a re-upload, and remote files not present locally are deleted.” Checksums catch difference, not direction. aws s3 sync works the same way.

The obvious guard doesn’t save you either. rsync --update skips files that are newer on the receiver, but it judges that by mtime. The fix for this setup is cp public/privacy.html build/privacy.html, and a plain cp stamps the copy with the current time. Copy an old file today and it’s newer than anything on the server. --update doesn’t stop deletions in any case.

Prateek Sharma’s post about an AI-generated aws s3 sync --delete in his GitLab CI is the usual framing of this failure: someone didn’t notice the flag. That’s a real failure, but it isn’t this one. I knew about the flag. What I didn’t know was how old the thing on the left side of the command was, or what was on the right side that the left had never seen.

A mirror deploy trusts the age of a directory nobody dates

Take the React app and my laptop out of it and here’s the general failure:

A deploy with mirror semantics treats its source directory as the truth, and the source is a build artifact that nothing guarantees is current, going to a destination that other paths also write to.

It shows up wherever the input to a delete-capable sync is a folder that gets regenerated only sometimes:

  • aws s3 sync dist/ s3://bucket --delete run by hand from a Vite or CRA project, as a quick fix between CI releases.
  • firebase deploy or netlify deploy --prod from a laptop whose public/ or dist/ came from a branch three weeks ago.
  • A static-site rsync --delete from a staging box that someone deployed to from somewhere else in the meantime.

CI mostly hides the problem because CI builds from a clean checkout on every run, so the artifact is never older than the commit. The bug comes back the first time a human runs the deploy step on its own, or the first time anyone writes to production without going through CI.

Here’s the property that was violated, written so it’s either true or false of your deploy:

Every directory that feeds a delete-capable sync is produced from a recorded, clean source commit in the same run that pushes it, and that commit is at least as new as the one production was built from. The sync’s delete set is empty, or every entry on it is on an explicit allowlist.

My setup broke all three parts. build/ wasn’t produced in the run that would have pushed it. Nothing recorded which commit it came from, so nothing could compare it to production. And nobody had ever looked at the delete set, let alone decided what was allowed on it.

Run this against your own deploy

You need your build directory, your deploy target, and a shell. Nothing here comes from our stack.

1. Is production ahead of your build? This is the check that matches my failure. Pick one file that carries a version number and compare the two sides. Adjust the pattern to however your page writes its version.

pat='[Vv]ersion[^0-9]{0,40}[0-9]+(\.[0-9]+)+'
live=$(curl -fsS https://your-site.example/privacy.html | grep -oE "$pat" | grep -oE '[0-9]+(\.[0-9]+)+' | head -1)
here=$(grep -oE "$pat" build/privacy.html | grep -oE '[0-9]+(\.[0-9]+)+' | head -1)
[ -n "$live" ] && [ -n "$here" ] || { echo "no version found on one side"; exit 1; }
newest=$(printf '%s\n%s\n' "$live" "$here" | sort -V | tail -1)
[ "$newest" = "$here" ] || { echo "build is BEHIND production: $here < $live"; exit 1; }

Passing: exit 0. Failing: exit 1 with the two versions printed. sort -V matters here, because it orders 1.10 after 1.9, which a plain string compare won’t. On 2026-07-30 this exits 1: build 1.1, production 1.3.

If no file you serve carries a version, skip this and use the build-sha check below. It’s the better one anyway.

2. What would the mirror delete?

# rsync target: list the delete set
rsync -an --delete --itemize-changes build/ deploy@your-host:/srv/site/ \
  | sed -n 's/^\*deleting *//p'

# S3 target
aws s3 sync build/ s3://your-bucket/ --delete --dryrun | grep 'delete:'

Passing: empty, or every path on the list is one you’ve decided production is allowed to lose. Failing: anything else. Those files exist in production and will stop existing. To make it an exit code, keep the allowed paths in a file (an empty file is fine) and fail on anything outside it:

rsync -an --delete --itemize-changes build/ deploy@your-host:/srv/site/ \
  | sed -n 's/^\*deleting *//p' | grep -vxFf delete-allowlist.txt \
  && { echo "unexpected deletions above"; exit 1; }

This is the check for the second half of the invariant. The version and commit checks can all pass and this one still fails, because a server-only backup or upload was never in any commit to begin with.

3. Secondary hygiene: is your artifact older than your source? This compares the build to local source, not to production, so it will fail any time you’ve edited something since the last build. That’s normal during development. Just don’t mirror-deploy while it’s failing.

# GNU find
newest=$(find build -type f -printf '%T@ %p\n' | sort -n | tail -1 | cut -d' ' -f2-)
# macOS / BSD
newest=$(find build -type f -exec stat -f '%m %N' {} + | sort -n | tail -1 | cut -d' ' -f2-)

find src public -type f -newer "$newest"

Passing: no output. Failing: a list of paths, each one a change your next mirror deploy won’t include. The sort sees every file even when find splits the stat into several batches, and cut -f2- keeps filenames with spaces intact. If you control the build step, it’s simpler to touch build/.built-at as its last line and use -newer build/.built-at.

The permanent fix: record the commit and compare it

Write the commit into the artifact at build time, refuse to write it from a dirty tree, and serve it:

[ -z "$(git status --porcelain)" ] || { echo "uncommitted changes; commit first"; exit 1; }
npm run build
git rev-parse HEAD > build/build-sha.txt

The git status --porcelain guard catches untracked files too, which git diff --quiet would miss. Without it, build-sha.txt names a commit that doesn’t describe what was built.

Then have the deploy script refuse to run unless this exits 0:

git merge-base --is-ancestor \
  "$(curl -fsS https://your-site.example/build-sha.txt)" \
  "$(cat build/build-sha.txt)"

That turns “is this folder stale?” from a feeling into an exit code. It has one hole, and I’m about to recommend walking into it: a single-file push. After cp public/privacy.html build/privacy.html and an upload, production has changed and the served build-sha.txt still names the old commit. So a single-file push has to do one of two things. Either commit the change first and upload a fresh build-sha.txt alongside the file, so the next mirror deploy must descend from a commit that contains it. Or accept that production no longer matches any recorded commit, and ban full-mirror deploys until a clean rebuild resets it. The side-channel installs on 2026-07-29 did neither, and that’s how my laptop ended up three versions behind without anything saying so.

What I run now to publish one page

The fix was to stop mirroring. Two rules: deploy the narrowest thing that works, and always dry-run first. For a single static page that looks like this:

cd frontend
cp public/privacy.html build/privacy.html   # generator writes public/, nginx serves build/
rsync -avzn --no-owner --no-group --rsync-path="sudo rsync" \
  -e "ssh -i ~/.ssh/deploy-key -o IdentitiesOnly=yes" \
  build/privacy.html deploy@your-host:/srv/site/build/privacy.html
# drop -n to send, then confirm what production actually serves:
curl -s "https://your-site.example/privacy.html?cb=1" | grep -oE 'Version[^0-9]{0,40}[0-9.]+'

There’s one file on each side and no --delete, so there’s nothing to delete. The -n output has one line, and the curl afterwards checks the live version instead of trusting the upload. The whole-directory deploy script still exists, but it’s now reserved for real frontend releases that start with a fresh build from a clean tree, and it runs check #2 before it touches anything.

The rule I’d give anyone: before you mirror a directory to production, you should be able to name the clean commit it was built from, show it’s at least as new as what’s live, and read out every file the sync will delete. If you can’t do all three, push individual files, not the directory.