Follow-up to f57e2a6, addressing the reviewer's sharpest question: is the race
"closed by construction", or merely detected and retried? The honest answer is
"both, on different paths", and the README said only the first half.
* README now states three claims separately instead of collapsing them:
a publisher rotating a snapshot cannot tear a clone of it (by construction —
the marker ordering prevents the unlink, and the consumer's verification is
a redundant second check on that path); every OTHER way the source can
change mid-clone is detected, not prevented (the eviction pass's reader
check is check-then-delete, and a seed-fallback-dir has no interlock at all
— there, verification plus a bounded retry and a loud failure is the whole
guard); and disk reclamation is bounded rather than immediate. Overclaiming
this property once was the finding; overclaiming it twice would be worse.
* publish-snapshot-selftest.sh now covers the PUBLISHER's half of the race,
where a reader of the swap belongs. Scenario 3 only ever covered a consumer
that had already FINISHED cloning — safe for free, since its own hardlinks
keep the inodes alive. New scenario 6 covers a reader still in flight past
the grace period (the generation is left on disk, the deferral is warned
about, and an earlier consumer is still unaffected); scenario 7 covers the
sweep, so "we defer instead of forcing" cannot quietly become a disk leak.
Red-proven against 248af306's scripts:
ASSERTION FAILED: the previous generation was unlinked while a reader
still held it
The consumer's half stays in seed-target-dir-selftest.sh scenario 8, which
still red-proves at 16693 of 48805 entries against the same scripts.
* seed-target-dir-selftest.sh now asserts what happens when the retries are
EXHAUSTED, not just what hardlink_clone_into returns: an unreadable source
makes the seed script exit non-zero, name the reason, leave no target dir,
and — the one that matters — not fall through to its cold-start branch. A
corrupt-cache bug degrading into an invisible 4x-slower CI job is the
failure mode worth pinning down. Skipped when running as root, where mode
bits deny nothing.
* usage_kb: a directory we cannot read measured as the empty string, which was
then spliced into usage_gb's awk program and made it a syntax error at the
exact moment something was already going wrong. Now measures 0.
Verification: `bash scripts/selftest.sh` — 5 suites, exit 0, 82 assertions
(was 75 after f57e2a6, 63 before). shellcheck over scripts/: no new findings.
Measured the cost the reviewer asked about, on ext4, warm cache, 78,554
entries: `cp -al` 3126 ms against 44 ms for one `find | wc -l`. Two counts per
attempt is ~2.8% on top of the clone. Not measured on the CI runner's volume.
Refs: daniel/gitdan#11, zemyna#911
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Sqh2vscfzisk83VuPVQX9L