fix(prune): stop protecting publisher target dirs, so a second branch can seed #24

Closed
opened 2026-09-22 15:36:17 +00:00 by claude · 0 comments
Collaborator

Symptom

On daniel/zemyna's CI volume (193 GB root), the cache holds a snapshot plus exactly one branch's target directory. A second branch cannot seed, so its run dies at Restore the Cargo cache — and it dies by filling the disk toward 100%, which on 2026-09-21 also truncated act_runner's cached action checkouts and cost two further dead runs before anything could build.

Every PR that night needed a hand eviction between runs. That is the failure to remove.

Cause

prune-cache.sh protects both target-<ref> and snapshot-<ref> for every publisher branch:

for ref in $PROTECTED_REFS; do
  suffix=$(cache_key "$ref")
  protected_ns["target-${suffix}"]=1
  protected_ns["snapshot-${suffix}"]=1
done

So target-dev is never a pressure-pass candidate, however old, and however tight the disk. With dev's target dir permanently resident at ~66 G alongside a ~74 G snapshot, there is no room for a PR branch to seed.

The script's own header only justifies protecting the snapshot:

Evicting a publisher's snapshot doesn't free real disk (every open PR's clone keeps the data alive) but does force every subsequent PR to start cold, which is the entire benefit this scheme exists to deliver.

That argument does not reach the target directory. A publisher's target dir is a convenience cache for publisher runs; it reseeds from the snapshot, and the currently-running job's own directory is already protected separately via $OWN_DIR.

A premise in the header is false on this volume, and should be corrected with the change

The eviction-order rationale asserts snapshots share inodes with target dirs:

A snapshot is a hardlink clone of a live target dir and of every consumer cloned from it, so removing it frees almost no real bytes — its inodes stay alive through those other links … Evicting snapshots first would be nearly pure loss.

Measured on the live volume, 2026-09-21 — unique bytes meaning files with link count 1, so shared inodes are excluded from both sides:

snapshot-dev-34c6fcec    74 G apparent   66.4 G unique
target-dev-34c6fcec      66 G apparent   64.5 G unique
volume total            139 G

They share almost nothing. The cause is publish-snapshot.sh's own design: it cp -al clones and then unshares every executable, because cargo and the linker rewrite binaries in place and would otherwise corrupt the shared original. Executables are ~85% of this tree (63.09 G of a 74 G snapshot, in 191 files), so the unshare defeats the sharing for nearly all of the bytes.

Evicting a snapshot therefore frees ~66 G of real disk, not "almost none". Target-before-snapshot ordering may still be the right default — a snapshot buys every future PR a warm start, which is a real and separate argument — but the stated reason is wrong and will mislead the next reader.

Acceptance criteria

  1. target-<ref> for publisher branches is no longer protected; snapshot-<ref> still is.
  2. The currently-running job's own target dir remains protected under every pass (it already is, via $OWN_DIR — confirm, do not assume).
  3. A selftest scenario in scripts/prune-cache-selftest.sh covers it: under forced pressure, a publisher's target- dir is evicted while its snapshot- dir survives. Red-prove it — it must fail against the current script.
  4. The existing scenario asserting protected refs are never evicted is updated rather than deleted, so it still asserts the snapshot half.
  5. The header's hardlink-sharing claim is corrected or deleted, citing the measurement above rather than restating a new unsourced figure.

Notes

  • This is the lever that addresses the reported failure. A sibling attempt to bound cache growth with cargo-sweep (zemyna#1080) was closed unmerged: it keys on atime, and under relatime — which this volume uses — it would periodically wipe the warm cache entirely. See daniel/zemyna#1073 for that analysis and for the measurement showing a complete zemyna build is only 9.6 G, the 74 G snapshot having been ~87% legacy debris.
  • Consider whether the liveness pass should also reclaim a merged branch's target dir sooner, but do not widen this change to do it — file separately if it looks worthwhile.
## Symptom On `daniel/zemyna`'s CI volume (193 GB root), the cache holds a snapshot plus **exactly one** branch's target directory. A second branch cannot seed, so its run dies at `Restore the Cargo cache` — and it dies by filling the disk toward 100%, which on 2026-09-21 also truncated act_runner's cached action checkouts and cost two further dead runs before anything could build. Every PR that night needed a hand eviction between runs. That is the failure to remove. ## Cause `prune-cache.sh` protects **both** `target-<ref>` and `snapshot-<ref>` for every publisher branch: ``` for ref in $PROTECTED_REFS; do suffix=$(cache_key "$ref") protected_ns["target-${suffix}"]=1 protected_ns["snapshot-${suffix}"]=1 done ``` So `target-dev` is never a pressure-pass candidate, however old, and however tight the disk. With `dev`'s target dir permanently resident at ~66 G alongside a ~74 G snapshot, there is no room for a PR branch to seed. **The script's own header only justifies protecting the snapshot:** > Evicting a publisher's snapshot doesn't free real disk (every open PR's clone keeps the data alive) but does force every subsequent PR to start cold, which is the entire benefit this scheme exists to deliver. That argument does not reach the target directory. A publisher's target dir is a convenience cache for publisher runs; it reseeds from the snapshot, and the currently-running job's own directory is already protected separately via `$OWN_DIR`. ## A premise in the header is false on this volume, and should be corrected with the change The eviction-order rationale asserts snapshots share inodes with target dirs: > A snapshot is a hardlink clone of a live target dir and of every consumer cloned from it, so removing it frees almost no real bytes — its inodes stay alive through those other links … Evicting snapshots first would be nearly pure loss. Measured on the live volume, 2026-09-21 — unique bytes meaning files with link count 1, so shared inodes are excluded from both sides: ``` snapshot-dev-34c6fcec 74 G apparent 66.4 G unique target-dev-34c6fcec 66 G apparent 64.5 G unique volume total 139 G ``` They share almost nothing. The cause is `publish-snapshot.sh`'s own design: it `cp -al` clones and then **unshares every executable**, because cargo and the linker rewrite binaries in place and would otherwise corrupt the shared original. Executables are ~85% of this tree (63.09 G of a 74 G snapshot, in 191 files), so the unshare defeats the sharing for nearly all of the bytes. Evicting a snapshot therefore frees ~66 G of real disk, not "almost none". Target-before-snapshot ordering may still be the right default — a snapshot buys every future PR a warm start, which is a real and separate argument — but the stated reason is wrong and will mislead the next reader. ## Acceptance criteria 1. `target-<ref>` for publisher branches is **no longer** protected; `snapshot-<ref>` still is. 2. The currently-running job's own target dir remains protected under every pass (it already is, via `$OWN_DIR` — confirm, do not assume). 3. A selftest scenario in `scripts/prune-cache-selftest.sh` covers it: under forced pressure, a publisher's `target-` dir is evicted while its `snapshot-` dir survives. Red-prove it — it must fail against the current script. 4. The existing scenario asserting protected refs are never evicted is updated rather than deleted, so it still asserts the snapshot half. 5. The header's hardlink-sharing claim is corrected or deleted, citing the measurement above rather than restating a new unsourced figure. ## Notes - This is the lever that addresses the reported failure. A sibling attempt to bound cache growth with `cargo-sweep` (zemyna#1080) was closed unmerged: it keys on atime, and under `relatime` — which this volume uses — it would periodically wipe the warm cache entirely. See daniel/zemyna#1073 for that analysis and for the measurement showing a complete zemyna build is only 9.6 G, the 74 G snapshot having been ~87% legacy debris. - Consider whether the liveness pass should also reclaim a **merged** branch's target dir sooner, but do not widen this change to do it — file separately if it looks worthwhile.
claude added the bug label 2026-09-22 15:36:17 +00:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: daniel/gitdan-actions#24