fix(cargo-cache): close the seed-vs-republish race the design claimed to close

The shared action's justification over zemyna's and emowheel's schemes was
that hardlink-cloning from a published snapshot closes gitdan #911 "by
construction, not by the single job slot". Review disproved that. This makes
the claim true, and corrects the README where it could only be bounded.

Finding 1 (verdict-level) — silent partial clone
------------------------------------------------
`hardlink_clone_into` ran `cp -al` with no exit-status check, and both call
sites invoked it as a condition, which suppresses `set -e` for the whole call.
A publisher's `rm -rf` of the generation it rotated away therefore unlinked
entries beneath an in-flight consumer walk, and the truncated tree was renamed
into place and reported as success.

Both layers are fixed:

* The consumer verifies its own clone. Every attempt checks `cp -al`'s status
  explicitly, the source directory's inode before and after (a wholesale
  replacement mid-walk splices two generations), and the entry count — the
  only signal for a subtree unlinked before its parent was listed, since
  `cp -al` reports no error for one it never saw. Any failure discards the
  staging tree and retries; exhausting the attempts returns a distinct status
  2 and fails the job rather than seeding a partial cache. `unshare_subtree` /
  `_unshare_files` now propagate failure too — a swallowed unshare leaves the
  clone aliasing its source, the exact corruption that step exists to prevent.
* The publisher does not unlink under a reader. A consumer publishes a
  `.reading-<snapshot>-<tag>` marker before it resolves the snapshot path; the
  publisher scans for markers after its first rename. A consumer holding the
  old generation therefore published its marker before that scan and cannot be
  missed; one arriving after the scan necessarily resolves to the new
  generation. The publisher waits for readers to drain and, on timeout,
  DEFERS reclamation rather than forcing it — the old generation is left as
  `.publish-old-<key>-<tag>` and swept by a later publish.

So correctness is closed by construction; disk reclamation is bounded, not
immediate. The residual is capped at one deferred generation per publisher
ref, and the README now says exactly that instead of the disproved claim.

Finding 2 — restore-mtimes.sh ran with no errexit
-------------------------------------------------
`set -euo pipefail` was glued to the end of a comment (`# soundness.set -euo
pipefail`), so it was entirely commented out: a partial failure of the
`git log | awk` pipeline would have produced wrong mtimes across the whole
restore instead of failing loudly. Moved to its own line. Audited every other
script for the same defect — this was the only instance. Independent
confirmation: shellcheck's two SC2164 warnings on this file's `cd "$repo_root"`
disappear now that errexit is actually in effect.

Finding 3 — lock-acquire window
-------------------------------
A just-seeded directory was unlocked until a later action step, so a
concurrent job's prune pass could evict it. `seed-target-dir.sh` now takes an
optional lock-id and writes the lock marker on every path out of the script,
including into the staging tree before its rename, so the directory carries a
lock the instant it appears under its final name. The action's acquire step
stays (it is idempotent and stamps the LRU marker).

Also hardened `prune-cache.sh` to treat a directory with live reader markers
as locked. Today no reachable configuration prunes a snapshot — only protected
refs publish them and protected refs are excluded from every pass — so this is
redundant by policy; it is here so that stops being the reason it is safe.

Verification
------------
New selftest scenario 8 races a real seed against a real publish rotation,
gating the rotation on the seed's *observed* clone progress so the window is
hit deterministically rather than on a fast machine's coin flip. Red-proven
against the unguarded scripts, three consecutive runs:

  ASSERTION FAILED: the seeded tree is truncated: 15443 entries against the
  snapshot's 493 (was 48805 before the rotation)      (15443 / 16986 / 16498)

Green after the fix, six consecutive runs, catching the clone mid-walk at
~10.5k of 48805 entries each time. Scenario 9 covers deferred reclamation and
its later sweep; scenario 10 covers an unreadable source failing loudly.

`bash scripts/selftest.sh`: 5 suites, exit 0, 75 assertions (was 63).
shellcheck over `scripts/`: no new findings, two SC2164 warnings resolved.

Docs: README's republish-safety paragraph replaced with what the code now
guarantees, including the bounded disk residual stated explicitly; new
`read-grace-seconds` / `reader-stale-seconds` inputs documented in the
`cargo-cache-publish` table; the selftest table names the new race.

Refs: daniel/gitdan#11, zemyna#911

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Sqh2vscfzisk83VuPVQX9L
This commit is contained in:
2026-08-23 16:40:21 -05:00
co-authored by Claude Opus 5
parent 248af3061e
commit f57e2a6013
9 changed files with 562 additions and 68 deletions
+20 -7
View File
@@ -49,12 +49,13 @@
# the entire benefit this scheme exists to deliver.
#
# LOCKED — a directory carrying a .ci-lock-* marker younger than
# STALE_LOCK_SECONDS is held open by a running job and is skipped by every
# pass, however dead and however tight the disk. This is what makes eviction
# safe on a runner with more than one execution slot. An older marker is
# treated as abandoned and logged as such, so an actually-still-running job
# that somehow exceeds the threshold is visible in the log rather than
# silently losing its cache mid-build.
# STALE_LOCK_SECONDS is held open by a running job, or named by a live
# .reading-<dir>-* marker (a job is hardlink-cloning it this instant), is
# skipped by every pass, however dead and however tight the disk. This is
# what makes eviction safe on a runner with more than one execution slot.
# An older marker is treated as abandoned and logged as such, so an
# actually-still-running job that somehow exceeds the threshold is visible
# in the log rather than silently losing its cache mid-build.
#
# Liveness is resolved by `git ls-remote --heads origin`, wrapped in a
# timeout. A directory name cannot be inverted back to a branch name (the
@@ -93,8 +94,20 @@ is_protected() {
}
is_locked() {
local dir="$1" now lock_file lock_age locked=1
local dir="$1" now lock_file lock_age locked=1 readers
now=$(date +%s)
# A directory being hardlink-cloned right now carries no .ci-lock-* of its
# own — a snapshot has its locks stripped by construction — so the reader
# markers are the only signal that unlinking it would truncate somebody's
# in-flight clone. Today no reachable configuration prunes a snapshot (only
# protected refs publish them, and protected refs are excluded from every
# pass), which makes this guard redundant *by policy*. It is here so that
# stops being the reason it is safe.
readers=$(live_reader_count "$ROOT" "$(basename "$dir")")
if [ "$readers" -gt 0 ]; then
echo " $(basename "$dir"): ${readers} job(s) currently cloning it — not a candidate"
locked=0
fi
for lock_file in "$dir"/.ci-lock-*; do
[ -e "$lock_file" ] || continue
lock_age=$(( now - $(stat -c '%Y' "$lock_file") ))