fix(cargo-cache): close the seed-vs-republish race the design claimed to close
The shared action's justification over zemyna's and emowheel's schemes was that hardlink-cloning from a published snapshot closes gitdan #911 "by construction, not by the single job slot". Review disproved that. This makes the claim true, and corrects the README where it could only be bounded. Finding 1 (verdict-level) — silent partial clone ------------------------------------------------ `hardlink_clone_into` ran `cp -al` with no exit-status check, and both call sites invoked it as a condition, which suppresses `set -e` for the whole call. A publisher's `rm -rf` of the generation it rotated away therefore unlinked entries beneath an in-flight consumer walk, and the truncated tree was renamed into place and reported as success. Both layers are fixed: * The consumer verifies its own clone. Every attempt checks `cp -al`'s status explicitly, the source directory's inode before and after (a wholesale replacement mid-walk splices two generations), and the entry count — the only signal for a subtree unlinked before its parent was listed, since `cp -al` reports no error for one it never saw. Any failure discards the staging tree and retries; exhausting the attempts returns a distinct status 2 and fails the job rather than seeding a partial cache. `unshare_subtree` / `_unshare_files` now propagate failure too — a swallowed unshare leaves the clone aliasing its source, the exact corruption that step exists to prevent. * The publisher does not unlink under a reader. A consumer publishes a `.reading-<snapshot>-<tag>` marker before it resolves the snapshot path; the publisher scans for markers after its first rename. A consumer holding the old generation therefore published its marker before that scan and cannot be missed; one arriving after the scan necessarily resolves to the new generation. The publisher waits for readers to drain and, on timeout, DEFERS reclamation rather than forcing it — the old generation is left as `.publish-old-<key>-<tag>` and swept by a later publish. So correctness is closed by construction; disk reclamation is bounded, not immediate. The residual is capped at one deferred generation per publisher ref, and the README now says exactly that instead of the disproved claim. Finding 2 — restore-mtimes.sh ran with no errexit ------------------------------------------------- `set -euo pipefail` was glued to the end of a comment (`# soundness.set -euo pipefail`), so it was entirely commented out: a partial failure of the `git log | awk` pipeline would have produced wrong mtimes across the whole restore instead of failing loudly. Moved to its own line. Audited every other script for the same defect — this was the only instance. Independent confirmation: shellcheck's two SC2164 warnings on this file's `cd "$repo_root"` disappear now that errexit is actually in effect. Finding 3 — lock-acquire window ------------------------------- A just-seeded directory was unlocked until a later action step, so a concurrent job's prune pass could evict it. `seed-target-dir.sh` now takes an optional lock-id and writes the lock marker on every path out of the script, including into the staging tree before its rename, so the directory carries a lock the instant it appears under its final name. The action's acquire step stays (it is idempotent and stamps the LRU marker). Also hardened `prune-cache.sh` to treat a directory with live reader markers as locked. Today no reachable configuration prunes a snapshot — only protected refs publish them and protected refs are excluded from every pass — so this is redundant by policy; it is here so that stops being the reason it is safe. Verification ------------ New selftest scenario 8 races a real seed against a real publish rotation, gating the rotation on the seed's *observed* clone progress so the window is hit deterministically rather than on a fast machine's coin flip. Red-proven against the unguarded scripts, three consecutive runs: ASSERTION FAILED: the seeded tree is truncated: 15443 entries against the snapshot's 493 (was 48805 before the rotation) (15443 / 16986 / 16498) Green after the fix, six consecutive runs, catching the clone mid-walk at ~10.5k of 48805 entries each time. Scenario 9 covers deferred reclamation and its later sweep; scenario 10 covers an unreadable source failing loudly. `bash scripts/selftest.sh`: 5 suites, exit 0, 75 assertions (was 63). shellcheck over `scripts/`: no new findings, two SC2164 warnings resolved. Docs: README's republish-safety paragraph replaced with what the code now guarantees, including the bounded disk residual stated explicitly; new `read-grace-seconds` / `reader-stale-seconds` inputs documented in the `cargo-cache-publish` table; the selftest table names the new race. Refs: daniel/gitdan#11, zemyna#911 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Sqh2vscfzisk83VuPVQX9L
This commit is contained in:
+17
-10
@@ -22,9 +22,9 @@ inputs:
|
||||
default: '10'
|
||||
restore-mtimes:
|
||||
description: >-
|
||||
Restore every tracked file''s mtime from git history. Requires a
|
||||
Restore every tracked file's mtime from git history. Requires a
|
||||
full-history checkout (fetch-depth: 0). Set to false only if the build
|
||||
does not use Cargo''s mtime-based freshness at all.
|
||||
does not use Cargo's mtime-based freshness at all.
|
||||
required: false
|
||||
default: 'true'
|
||||
prune:
|
||||
@@ -53,7 +53,7 @@ inputs:
|
||||
default: ''
|
||||
watermark-file:
|
||||
description: >-
|
||||
Name of this job''s build-watermark file inside the target dir. MUST be
|
||||
Name of this job's build-watermark file inside the target dir. MUST be
|
||||
distinct per job when two jobs share one cache key. Defaults to
|
||||
.ci-watermark-<job>-sha.
|
||||
required: false
|
||||
@@ -140,7 +140,13 @@ runs:
|
||||
# Seeds this ref's target dir from the base's published snapshot. See
|
||||
# scripts/seed-target-dir.sh — the staging-then-atomic-rename is what
|
||||
# makes concurrent jobs sharing one cache key safe by construction rather
|
||||
# than by the runner happening to have a single execution slot.
|
||||
# than by the runner happening to have a single execution slot, and the
|
||||
# clone's own consistency check plus the publish side's reader interlock
|
||||
# are what make it safe against the base republishing MID-CLONE.
|
||||
#
|
||||
# The lock id is passed here as well as acquired in the next step: the
|
||||
# seed writes it into the staging tree, so the directory carries a lock
|
||||
# the instant it appears under its final name rather than a step later.
|
||||
- id: seed
|
||||
shell: bash
|
||||
run: |
|
||||
@@ -151,13 +157,14 @@ runs:
|
||||
"${{ steps.resolve.outputs.base-key }}" \
|
||||
"${{ inputs.cache-root }}" \
|
||||
"${{ github.job }}-${{ github.run_id }}-$$" \
|
||||
"${{ inputs.seed-fallback-dir }}"
|
||||
"${{ inputs.seed-fallback-dir }}" \
|
||||
"${{ steps.resolve.outputs.lock-id }}"
|
||||
|
||||
# Marks the directory as held open, so any job's prune pass (this one
|
||||
# included) skips it, and stamps the LRU marker. The marker is touched
|
||||
# unconditionally every run: a run that hits the cache for every crate may
|
||||
# write nothing at all inside the tree, which would make a just-used
|
||||
# directory look stale to the eviction pass.
|
||||
# Re-stamps the lock the seed step already wrote (acquiring is idempotent
|
||||
# — it rewrites the timestamp) and stamps the LRU marker. The marker is
|
||||
# touched unconditionally every run: a run that hits the cache for every
|
||||
# crate may write nothing at all inside the tree, which would make a
|
||||
# just-used directory look stale to the eviction pass.
|
||||
- shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
|
||||
Reference in New Issue
Block a user