31 Commits
Author SHA1 Message Date
claude 680ef8049c Merge pull request 'docs: clear PR #28's docs-drift catalog entries' (#31) from chore/docs-catalog-2026-09-23 into main
CI / shellcheck + selftests (push) Successful in 1m29s
CI / move v1 to main (push) Successful in 4s
2026-09-23 03:02:39 +00:00
claudeandClaude Opus 5.5 22dafe45e2 docs: clear PR #28's docs-drift catalog entries
CI / shellcheck + selftests (pull_request) Successful in 1m23s
CI / move v1 to main (pull_request) Skipped
Marks the token grant as verified now that #28's own merge ran
release-tag and moved v1 to that commit (confirmed against the CI
status API and the v1 tag on origin), and rewords the scenario-9
stranding framing to "a run that deferred and left no newer run
behind it" rather than a concurrency-group cancellation ci.yaml no
longer allows.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 21:53:32 -05:00
claude 0284fcd54b Merge pull request 'ci: release v1 automatically on a green merge to main' (#28) from chore/v1-release-gate into main
CI / shellcheck + selftests (push) Successful in 1m45s
CI / move v1 to main (push) Successful in 4s
2026-09-22 23:12:59 +00:00
claudeandClaude Opus 5.5 77cc5917b6 fix(release): refuse to overwrite an unrelated hand-placed v1
CI / shellcheck + selftests (pull_request) Successful in 1m47s
CI / move v1 to main (pull_request) Skipped
release-v1.sh's push_leased() only detected a lost lease after a push
was *rejected* -- but force-with-lease compares the remote ref's raw
value against the caller's expected value, not ancestry. If v1 already
sat on a hand-placed, unrelated commit when push_leased() was first
called (no race, nobody moves it mid-call), the very first push found
the ref exactly where it expected, succeeded outright, and silently
overwrote the unrelated v1 with <sha> -- skipping every ancestry check
the function has, since those only run after a rejection.

Fix: before the first push attempt, check whether the caller's
`expect` is neither an ancestor of `sha` (the ordinary stale-v1 case)
nor already covering it (nothing to do) -- and go red naming both SHAs
if so. `expect` is always a peeled commit (fetch_v1() reads
`refs/tags/v1^{commit}`), so this doesn't add a second failure mode
for an annotated v1; that tag form's existing "not a lost lease"
behavior on the first rejected push is untouched.

Surfaced by PR #28's final review. New selftest scenario 11 in
release-v1-selftest.sh, red-proven against the unfixed script (v1 was
silently moved off the stray commit); green after the fix, with the
full 7-suite gate (shellcheck + selftest.sh) passing.

Ride-alongs from the same review:
- ci.yaml:120-123 claimed the README's Versioning section documented
  what's verified about the release token's write access; it said
  nothing. Added an accurate sentence there (the grant is unobserved
  until the first merge, capped by repo/owner token-permission maxima
  and unreadable tag protections) and pointed the comment at it.
- Deleted two comment-as-decision-history paragraphs per
  comments-are-not-exposition: ci.yaml's "no job-level concurrency"
  rationale (kept one line of intent) and
  prune-cache-selftest.sh scenario 9's account of how an
  existence-only assertion used to pass with the pass-2 guard removed
  (kept a one-line statement of what it checks).
- README's release-v1-selftest.sh table row now names the new
  hand-placed-v1 scenario.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 16:14:41 -05:00
claudeandClaude Opus 5.5 ca0ee132d9 fix(ci): add a scheduled v1 sweep and lease every v1 push
CI / shellcheck + selftests (pull_request) Successful in 1m46s
CI / move v1 to main (pull_request) Skipped
The release guard in 17d87b0 was safe but not live. Gitea 1.27.2 calls
CancelPreviousJobsByJobConcurrency whenever a job's `needs` resolve
(services/actions/clear_tasks.go:91, models/actions/run_job.go:641), so
a job's place in the `release-tag-v1` group followed when its own
selftest finished, not merge order. A newer merge C2 finishing selftest
first queued behind the older C1, C1 cancelled it, saw tip = C2, and
deferred: nobody pushed, and if merges then stopped v1 stayed stale
indefinitely behind a Skipped and a Cancelled job. The "always catches
up once merges pause" claim in ci.yaml and README was false.

What now holds:

- release-sweep.yaml runs on `schedule` every 15 minutes, in its own
  workflow and concurrency group, so nothing in ci.yaml can cancel it.
  When v1 already covers main's tip it stops after a checkout and one
  merge-base. Otherwise it checks out the tip, runs the same shellcheck
  and selftest.sh as ci.yaml's selftest job, and tags the tip only if
  they pass; a failing main therefore turns the sweep red on every tick
  while v1 lags, which is #27's AC1 loud-failure half. It reads the tip
  itself because a scheduled run's github.sha is the CommitSHA recorded
  when the schedule was registered on the last push to main
  (services/actions/notifier_helper.go:569-580,
  services/actions/schedule_tasks.go:126-141), and ref is the default
  branch: schedules are registered only from it
  (notifier_helper.go:120, :531, :603-604). event_name is "schedule"
  (context.go:71 reads TriggerEvent, set at schedule_tasks.go:136).
  Cron is 5-field robfig in UTC (models/actions/schedule_spec.go:38-41).

- 15 minutes, not 10: the sweep is the fallback, not the release path,
  and every tick is a run on gitdan-ci's shared slots and a row in the
  Actions list. 96 no-op runs a day of a few seconds each is the cost;
  the lag bound it buys is one interval plus one selftest run.

- Both writers go through scripts/release-v1.sh and push with
  --force-with-lease=refs/tags/v1:<v1 as read>, so v1 cannot move
  backwards when the sweep and a merge job race. A lost lease re-reads
  v1: at or ahead of this run's gated commit is a clean skip (the other
  writer released something at least as new); still behind it is a
  retry leased on the new value, up to three attempts, since the other
  writer may have tagged an older commit and giving up there would leave
  v1 short of a commit this run did gate; anything else goes red. A
  rejection with v1 unmoved is diagnosed as a non-lease failure and goes
  red at once.

- release-tag loses its job-level concurrency group. The lease already
  gives the ordering the group was there for, and the group was what
  cancelled the one job that could have released the newest merge.
  Without it each merge's job runs, and the one whose commit is still
  the tip when it checks releases it.

The shell moves out of ci.yaml into scripts/release-v1.sh so shellcheck
and selftest.sh cover it. release-v1-selftest.sh runs it against a
scratch bare origin: sweep no-op at and ahead of the tip, tag on a
green gate, no tag and a failing sweep on a red one, the stranded trace
above followed by a catching-up sweep, and each lost-lease outcome, with
a control showing an unleased push does step v1 back. Red-proved by
seven mutations of release-v1.sh, each failing a named assertion: plain
--force, accepting any lost lease, a merge job that never defers,
ancestry reduced to equality, no non-lease diagnosis, a sweep that never
needs to run, and a retry that does not re-lease.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 15:01:05 -05:00
claudeandClaude Sonnet 5 17d87b0647 fix(ci): only ever push a commit this run actually gated; fix vacuous scenario-9 guard
## v1 could advance onto an ungated commit

`needs: selftest` gates this run's own commit, but the push targeted
origin/main's freshly-fetched tip with nothing comparing the two.
Trace: M1 merges green; M2 merges while M1's selftest is still
running; M1's release-tag job fetches tip = M2 and pushes v1 = M2,
whose own selftest may be queued, running, or red. If M2 is red, its
own job is skipped, so v1 sits on a red commit across every consuming
project until the next green merge -- with nothing red pointing at
the release itself. ci.yaml:107-109 and README.md:640-641 both
asserted this couldn't happen; ea48c03's own diff established the
precondition (tip "can be minutes stale... behind its own selftest
job") without closing it.

Fix: skip the push unless origin/main's tip IS this run's own
github.sha, checked before the existing v1-monotonicity check
(dc1e631) rather than replacing it -- the two compose (tip-mismatch
first, since it's the coarser reason to defer; ancestor-check second,
for a duplicate run whose own commit is still current). Reverted the
push target from the fetched tip back to `${{ github.sha }}`, now
that the guard makes them provably equal whenever the push fires.

## Convergence trace: does the newest commit's job still run?

Read gitea's source further at the pinned v1.27.2 tag:
PrepareToStartJobWithConcurrency (services/actions/clear_tasks.go)
calls CancelPreviousJobsByJobConcurrency on every job entering the
group, unconditionally cancelling whatever was previously
Waiting/Blocked there -- so at most one job sits queued in the group
at a time; each new arrival supersedes it. Because job-level
concurrency is only evaluated once `needs: selftest` is satisfied
(job_emitter.go re-evaluates readiness there), "arrival order" tracks
each commit's own selftest-completion time, not raw merge order -- an
older commit with a slower selftest can enter the group after a
newer one and cancel its queued slot.

That cancelled job is gone for good; it will never push. But the
commit that's genuinely current at the moment merges stop arriving is
always the one still queued when the running job finishes, because
every subsequent arrival (from every subsequent merge, not just the
"newest" one at any single instant) keeps re-superseding the queue.
So v1 always eventually catches up -- "one merge later" in the common
case, "at the next merge, whenever that happens" in the adversarial
case where a stale survivor runs, finds itself no longer current,
defers, and nothing is left queued. It cannot get stuck forever short
of the repository never receiving another merge, because every future
push re-attempts the same check against whatever's current by then.
Stated this plainly in the comment and README rather than repeating
the false "never" guarantee in softer words.

Re-derived the truth table against the new guard in a scratch
origin+clone, six cases: own commit == tip, no v1 (push); v1 already
== own commit (skip, duplicate run); tip moved past own gated commit
because a newer merge landed (skip, defers); the newer commit's own
run once nothing further has landed (push); v1 already ahead of a
now-stale gated commit (skip, tip-mismatch catches it first); tip ==
own commit but v1 independently ahead via a local-only descendant,
isolating the second (ancestor) check on its own (skip). All six
resolved as intended.

## Scenario 9's pass-3 guard was vacuous

scripts/prune-cache-selftest.sh:230 grepped 'self-clear' against
$scratch/log. summary_line() (cache-lib.sh) writes only to
$GITHUB_STEP_SUMMARY; the self-clear branch (prune-cache.sh:632)
writes 'self-clear' there and 'clearing own' to stdout (:631) --
'self-clear' never appears in $scratch/log at all, so the branch was
unreachable and the `ok` unconditional. Scenario 8 already greps the
right string against the right file (assert_log "clearing own" ...);
scenario 9 now does the same, staying on $scratch/log where it
already was -- the string was wrong, not the file.

Red-proved by capturing a real prune-cache.sh log where self-clear
genuinely fired (scenario 9's own fixture with MIN_FREE_PCT
temporarily raised to 100, in a scratch copy, reverted after) and
running both patterns against it: `grep -q 'self-clear'` -> no match
(the old check's vacuous pass, confirmed); `grep -q 'clearing own'`
-> match (the fix's correct fail). Matches the reviewer's own
measurement exactly. The real prune-cache-selftest.sh was untouched
during this experiment; only the grep string changed in the actual
commit.

bash scripts/selftest.sh: all 6 suites green (67 prune-cache
assertions, unchanged in count -- the fix corrects what scenario 9's
existing check compares, not what it asserts). shellcheck -x
--source-path=scripts scripts/*.sh: clean, as before not covering the
inline ci.yaml shell.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 14:01:28 -05:00
claudeandClaude Sonnet 5 ea48c03aeb fix(ci): target origin/main's live tip, not the run's own trigger commit
CI / shellcheck + selftests (pull_request) Successful in 1m35s
CI / move v1 to main (pull_request) Skipped
The wake-one-cancel-the-rest mechanism behind the concurrency group
(CancelPreviousJobsByJobConcurrency, models/actions/run_job.go, at
the v1.27.2 tag this instance runs) picks its survivor from
models/actions/run_job_list.go's query, which carries no `ORDER BY`
-- so under three-way contention on a shared 2-slot runner, the
*newest* commit's job can be the one cancelled while an older sibling
survives and, correctly from its own vantage, advances v1 forward
from a stale view. No push regresses v1 (the ancestor check from
dc1e631 already prevented that), but the newest merge goes silently
unreleased behind a cancelled job that reads as benign, not red --
exactly the failure #27 exists to end.

Every job now resolves `origin/main`'s tip fresh, right before the
push, instead of using `${{ github.sha }}`. Re-fetched explicitly
rather than trusted from the checkout step, which can be minutes
stale by this point behind its own selftest job. Every execution that
reaches the push step now converges on the same target regardless of
which job the concurrency group lets through, so which one wins the
wake no longer matters -- the survivor pushes where any of them
would have.

That doesn't make the push safe on its own: two jobs can still read
main at genuinely different moments if it advances between their two
fetches, so whichever read the tip earlier must not overwrite the
other's already-pushed, newer one. The ancestor check from dc1e631 is
kept for exactly this -- its target changed (origin/main's live tip,
not this job's own trigger commit) but its job didn't.

What each guard now protects against, after this change:
- concurrency group: stops two jobs from pushing at the same time --
  wasted work now that a cancelled job costs nothing, not a
  correctness backstop by itself.
- ancestor check: stops a job whose own fetch of the tip is stale
  relative to another job's already-pushed, fresher one from
  regressing v1.

Rewrote both the job-level comments and README's Versioning section,
which described "force-moves v1 to that commit" and said nothing
about the concurrency group or the skip-as-success case.

Re-derived the truth table against the new target (a live origin/main
tip, not a fixed commit) in a scratch origin+clone: no v1 yet (push),
v1 exactly at the tip (skip), the tip moved past v1 because a newer
merge landed (push, to the new tip -- not stuck on any prior commit),
v1 already ahead of the tip (skip, defensive), unrelated histories
(push, defensive). All five resolved as intended; ran the exact
condition and fetch sequence the workflow step uses, not a
simulation of it.

bash scripts/selftest.sh: all 6 suites green (67 prune-cache
assertions, unchanged). shellcheck -x --source-path=scripts
scripts/*.sh: clean -- as before, this does not cover the inline
`run:` shell in ci.yaml.

What remains unverifiable short of a real merge is unchanged from
dc1e631: the workflow step's actual execution inside a real Actions
run, and whether the built-in token has write access at all. Neither
this commit nor the one before it can exercise those.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 12:49:05 -05:00
claudeandClaude Sonnet 5 dc1e6317c6 fix(ci): make the v1 push monotonic, not just mutually exclusive
The concurrency group added in af1233f only excludes two release-tag
jobs that are simultaneously Running/Waiting/Blocked
(services/actions/clear_tasks.go:64-80,
models/actions/run_job.go:641-660 at the v1.27.2 tag this instance
runs) -- it has no notion of commit order between jobs that never
overlap. On a 2-slot runner shared across 4 repos, with a
multi-minute selftest gating each release-tag job, two merges landing
close together routinely finish in the opposite order from the pushes
that triggered them: if the newer commit's job completes and exits
first, the older commit's job later finds no live holder in the
group, is not blocked, and force-pushes v1 backward to itself. The
concurrency comment's "never regress it to an older one" and af1233f's
commit message both asserted the opposite -- true of the simultaneous
case the guard covers, false of the staggered one it doesn't, so
authored-false rather than drift.

Added a merge-base check before the push: skip if v1 already points
at this commit or a descendant of it (`--is-ancestor` treats a commit
as its own ancestor, so "at" and "ahead" are the same branch). A v1
that doesn't exist yet, or shares no history with this commit, falls
through to the push -- the guard only ever skips, never fails. Needs
`fetch-depth: 0` on the checkout: actions/checkout's own description
for that value is "all history for all branches and tags", and its
source (dist/index.js: fetchDepth <= 0 selects
getRefSpecForAllHistory, which includes the tags refspec) confirms
tags are fetched as part of that, not gated behind the separate
fetch-tags input -- so refs/tags/v1 and the history behind it are both
guaranteed present locally without a second fetch step.

Rewrote both false claims: the concurrency comment now says what the
group actually bounds (simultaneous competing pushes, not completion
order), and states plainly that the guard below is what makes the
outcome order-independent.

Verified the check's five cases (no tag yet, tag ahead of this
commit, tag behind this commit, tag equal to this commit, unrelated
history) against a real scratch git repo, extracting the exact
condition used in the workflow step -- all five resolved as intended
(skip only when v1 is already at or ahead). The workflow step itself
cannot be exercised outside a real Actions run.

bash scripts/selftest.sh: all 6 suites green (67 prune-cache
assertions, unchanged). shellcheck -x --source-path=scripts
scripts/*.sh: clean -- note this does not cover the inline `run:`
shell in ci.yaml, which shellcheck was never wired to check in this
repo (verified against the CI job itself, which shellchecks only
scripts/*.sh).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 12:24:27 -05:00
claudeandClaude Sonnet 5 af1233f14a fix(ci): serialise release-tag across concurrent merges
The workflow-level concurrency group is keyed per-commit (github.sha,
ci.yaml:19-28) so unrelated pushes never block each other -- but that
also means two merges landing close together run two concurrent
release-tag jobs, each force-pushing its own commit to v1. If the
older commit's job finishes last, v1 regresses to a stale-but-green
commit and stays there until the next merge corrects it forward.
Bounded blast radius (never a red commit, self-heals on the next
merge), but a silently wrong v1 is the exact failure #27 exists to
end.

Added a job-level `concurrency:` on release-tag with a fixed group
name and `cancel-in-progress: false`. Confirmed this is additive to
the workflow-level group, not a replacement, by reading gitea's source
at the v1.27.2 tag this instance runs (`tea api version`): run-level
and job-level concurrency are separate model fields
(ActionRunAttempt.ConcurrencyGroup vs ActionRunJob.ConcurrencyGroup),
evaluated by separate functions (EvaluateRunConcurrencyFillModel vs
EvaluateJobConcurrencyFillModel, services/actions/concurrency.go) and
enforced by separate cancellation paths (CancelPreviousJobsByRunConcurrency
in models/actions/run.go:364 vs CancelPreviousJobsByJobConcurrency in
models/actions/run_job.go:641) -- job_emitter.go's checkRunConcurrency
checks both groups independently (services/actions/job_emitter.go:210-236).
A fixed, non-sha group name is what serialises release-tag across
commits without touching the per-sha grouping every other job still
relies on. `cancel-in-progress: false` matters because Gitea's default
queue behaviour (documented same as GitHub's: a new queued job
supersedes an older *queued* one in the same group, but not a
*running* one) already gives the newest commit's push priority once
the running job clears; cancelling the running job too would abandon
whichever merge is currently pushing mid-flight, the same defect
wearing different clothes.

Re-verified: YAML parses, shellcheck clean, `bash scripts/selftest.sh`
all 6 suites green (67 prune-cache assertions, unchanged by this
commit).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 12:07:01 -05:00
claudeandClaude Sonnet 5 21b444121d test(prune): red-prove $OWN_DIR protection under genuine disk pressure
Scenario 9 asserted only that target-$OWN existed, checked right after
a no-pressure run (3b) where nothing was ever a candidate for
eviction, and prune-cache.sh's self-clear step unconditionally
recreates an empty $OWN_DIR whenever the run ends under the percentage
floor regardless of what pass 2 did to it. Either way the existence
check passed whether or not the pass-2 guard (protected_reason,
prune-cache.sh:221) actually protected the directory. Moving that
guard's $OWN_DIR check into the pass-1-only predicate — the exact
mutation gitdan-actions#26 describes — left all 64 assertions green,
confirmed here before the fix.

Rewritten to run under a real, shrinking `df` (the scenario-17
pattern: a fake df that re-measures the fixture with `du` on every
call, so eviction genuinely lowers the reported pressure), with
MIN_FREE_PCT=0 and a clone-headroom floor sized so self-clear's
percentage check can never fire — only pass 2's guard decides the
outcome. The fixture gives $OWN_DIR real content and an older
timestamp than a sibling target dir, sized so the requirement is met
by evicting exactly one of them. With the guard removed, $OWN_DIR is
the one evicted (LRU-oldest, and self-clear is structurally disabled
by MIN_FREE_PCT=0 so there is nothing left to recreate it) — the
scenario now fails loudly on the same mutation that left it green
before. Restored and reverified green (67 assertions, up from 64) with
the guard intact.

bash scripts/selftest.sh: all 6 suites green. shellcheck -x
--source-path=scripts scripts/*.sh: clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 11:52:01 -05:00
claudeandClaude Sonnet 5 7f18cb2436 feat(ci): advance v1 automatically once the gate is green on main
Nothing moved v1 when main advanced, so a merged change was inert
until someone remembered to retag by hand — it happened on 2026-09-22
(PR #25 merged, v1 stayed on the previous release) and was only caught
because a person asked whether the tag had moved.

Option 1 from gitdan-actions#27 (automate it) over option 2 (fail loud
while it lags): a `release-tag` job, gated with `needs: selftest` so a
broken build never reaches it, force-moves v1 to the pushed commit
using the run's built-in GITHUB_TOKEN. If that token turns out not to
have write access, the push fails and the job goes red in the Actions
UI — a loud failure either way, not the silent one this replaces.
Whether the token actually has write access here is unverified short
of a real merge; that merge is the next step for this branch.

README's Versioning section documented the old manual step as a
deliberate decision "never something a merge does by itself" — that
claim is now false, so it's rewritten to describe the automated job
and keeps the manual command as the recovery path for when the job
can't push.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 11:51:51 -05:00
claude 21dffdb725 Merge pull request 'fix(prune): stop protecting publisher target dirs under disk pressure' (#25) from fix/prune-unprotect-target into main
CI / shellcheck + selftests (push) Successful in 1m49s
2026-09-22 16:07:52 +00:00
claudeandClaude Sonnet 5 dc473f0d0c fix(prune): stop protecting publisher target dirs under disk pressure
CI / shellcheck + selftests (pull_request) Successful in 1m48s
A publisher branch's target-<ref> was protected identically to its
snapshot-<ref>, so it was never a pressure-pass candidate however old
and however tight the disk. With one branch's target dir permanently
resident alongside its snapshot, a second branch had no room to seed,
and every PR that night needed a hand eviction between runs
(daniel/gitdan-actions#24).

A publisher's target dir is a convenience cache its own next run
reseeds from the snapshot, so losing it under pressure is cheap;
nothing downstream depends on it surviving. Only the snapshot stays
protected in the pressure and self-clear passes.

Unprotecting the target dir outright surfaced a second bug the fix
would otherwise have shipped: a branch's tip is trivially an ancestor
of itself, so once a protected ref's target dir was no longer skipped
before reaching the merged-branch check, pass 1 read it as "merged
into itself" and deleted it unconditionally on every run, independent
of disk pressure. is_protected_from_liveness keeps a protected ref's
target dir out of pass 1 alone, so it stays an ordinary pressure-pass
candidate without ever reaching that check. Scenario 3b in the
selftest red-proves this against the unprotect-only version of the
fix.

Also corrects the header's inode-sharing claim, measured false on the
live volume by daniel/zemyna#1073: publish-snapshot.sh unshares every
executable after its cp -al, and executables are most of the tree by
bytes, so a snapshot eviction is a real, large disk cost rather than
the near-free one the old text described — the target-before-snapshot
ordering still holds, now for the warm-start reason alone plus that
cost.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 10:42:08 -05:00
claude 0184df25a2 Merge pull request 'fix(prune): reclaim merged branches and size the volume for the clone' (#21) from fix/prune-merged-live-and-seed-headroom into main
CI / shellcheck + selftests (push) Successful in 1m24s
2026-09-07 05:41:12 +00:00
claudeandClaude Fable 5.1 38a6387936 docs(cache): recommend delete-on-merge, and say what ancestry still covers
CI / shellcheck + selftests (pull_request) Successful in 1m27s
`daniel/zemyna` enabled `default_delete_branch_after_merge` after this
branch was written, so the deleted-branch signal will fire there on
future merges. That makes the setting worth recommending — it is the
cheapest case for this scheme, decidable from `ls-remote` with no
checkout, no objects and no walk — and it does not make the ancestry
signal redundant.

Three things it leaves behind, now enumerated in the README's eviction
section rather than implied: every branch merged before the setting was
turned on, of which zemyna carried 48 and which nothing retroactively
deletes; every merge whose deletion the forge declines or is never asked
to make, since it is best-effort and silent and an API merge without the
flag never asks; and every repo that has not enabled it, which is the
default.

Two present-tense claims about one repo's configuration are reworded
into the conditions they were standing in for, in prune-cache.sh's
header and beside is_merged_dead, plus the two in the selftest that
asserted the forge keeps branches rather than describing the fixture.
No behaviour change; the suite is green and unchanged at 61 assertions.

Refs #20.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXMQCJ5Eg5f9G9cfYzyh4Z
2026-09-07 00:27:11 -05:00
claudeandClaude Fable 5.1 24f87a6b98 docs(cache): describe both liveness signals and the derived requirement
CI / shellcheck + selftests (pull_request) Skipped
The eviction section described one way for a branch to be dead and a
percentage threshold that no longer exists as a gate. Rewritten around
what the pass actually does: two dead-branch signals with the limits of
the ancestry one stated, a requirement measured off the clone's mutable
set, and a failure that names its shortfall.

The 10% figure is deleted rather than corrected — it was the default of
a gate, and min-free-percent is now an additional floor defaulting to 0,
so there is no percentage left to state. The inputs table, the prune and
liveness rows, the selftest coverage row and the fetch-depth comment in
the usage example all say what they now mean; fetch-depth: 0 has a
second reason to be required, since a shallow checkout cannot answer
ancestry.

Refs #20.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXMQCJ5Eg5f9G9cfYzyh4Z
2026-09-07 00:21:12 -05:00
claudeandClaude Fable 5.1 07ba53ca79 fix(prune): reclaim merged branches, and size the volume for the clone
Two assumptions in the eviction pass did not hold on this forge, and
between them a volume filled up three times in three days with nothing
reclaimed automatically. Both are replaced here; the pass also moves
ahead of the seed, which is the only order in which its work can help
the run performing it.

LIVENESS. Pass 1 evicted a cache only when its branch was gone from
origin. Gitea keeps a PR's branch after the merge unless the repo opts
into delete-on-merge, and zemyna does not — so ls-remote reports fifty
merged branches and the signal fires for none of them. A second signal
is added beside it: a branch still on origin whose tip is an ancestor of
a protected branch's tip holds no commit that branch does not, so its
cache will never be read again and goes in the same unconditional pass.

Ancestry is answered from the commits in the job's own checkout, so the
answer "cannot tell" exists and stays distinct from "not merged" at both
granularities. A shallow checkout withholds the signal entirely, since a
missing object is its normal case rather than evidence. A single branch
whose tip is not in the checkout is kept, with a warning naming it. A
squash or rebase merge leaves no ancestry and reads as live until the
branch is deleted. All three are missed reclamations, which cost disk;
the other direction costs a branch its cache mid-build.

HEADROOM. Passes 2 and 3 gated on a percentage of the volume, which
cannot express the failure they have to prevent: a clone runs out of
disk while unsharing its mutable paths, and how much that needs is a
property of the snapshot rather than of the disk. Staging failed at 34 G
free and passed at 74 G, so a 10% floor — 19 G here — never fired first.
The requirement is now measured per run off the very source the seed
will read, the pass evicts oldest-first until it is met and stops there,
and falling short of it fails with the shortfall and every directory it
kept, rather than letting the seed fail seconds later against a staging
path that names none of that. min-free-percent survives as an additional
floor, defaulting to 0, and falling short of that one is still a warning
and a self-clear.

ORDERING. The prune step ran after the seed, so each run freed space for
the next one. It now runs between resolve and seed. Two things that
makes newly reachable are closed: the source about to be cloned is
excluded from every pass by name, and a concurrent job's target dir
already carries its lock from the instant it appears under its final
name, so nothing is seen unlocked that is in use.

Red-proven: sixteen assertions across the four new scenarios fail
against the pre-fix scripts, including the zemyna layout evicting
nothing where it should evict exactly one directory, and the headroom
scenario exiting 0 where it should exit 1.

Refs #20.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXMQCJ5Eg5f9G9cfYzyh4Z
2026-09-07 00:21:12 -05:00
claudeandClaude Fable 5.1 a960c8f91b refactor(cache): name the mutable set once, and measure it
The set of paths a hardlink clone has to real-copy — dep-info,
build-script metadata, linked outputs, the pruned directories that hold
them — was spelled out inline in unshare_mutable_paths, in four find
invocations. Nothing else needed it, so one spelling was enough.

Something else needs it now: the prune has to know what a clone will
cost before it happens, and a sizer with its own copy of the predicates
would drift from the copier silently and in the dangerous direction — an
under-measured clone is one that starts and runs out of disk halfway
through unsharing. So the directory names become one array and the file
rules one dispatcher, applied by a callback per side, with each rule's
rationale moved to the rule rather than left at the old call site.

mutable_set_kb measures that set off a snapshot, skipping the subtrees
already measured whole so nothing is counted twice; clone_headroom_kb
scales it by a hand-written margin and floor for what the measurement
cannot see (cp -al materialising every directory for real, and the
unshare holding one subtree twice at its peak). Both residuals are named
where the function is, in both directions.

seed_source_candidates moves the seed's source-preference list into
cache-lib for the same reason: the prune ahead of it has to resolve the
same source the seed will clone, and two agreeing derivations are one
edit away from disagreeing.

No behaviour change — the copier applies the same rules to the same
tree, verified by the hardlink-clone suite's inode partition in both
directions.

Refs #20.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXMQCJ5Eg5f9G9cfYzyh4Z
2026-09-07 00:20:45 -05:00
claude 0b099ecf62 Merge pull request 'fix(ci): trigger on edited so un-drafting actually lifts the draft skip' (#18) from fix/ci-edited-trigger into main
CI / shellcheck + selftests (push) Successful in 1m24s
2026-09-02 18:54:43 +00:00
claudeandClaude Opus 5 5abd0a9968 docs(ci): name the timing fields this forge actually returns
CI / shellcheck + selftests (pull_request) Successful in 1m30s
The accepted-cost derivation cited `run_started_at` and `updated_at`, which are
GitHub's field names. This Gitea's runs payload has neither -- it returns
`started_at` and `completed_at`, and omits `run_started_at`, `updated_at` and
`created_at` entirely (confirmed by dumping the keys of a run object). The
figures are unaffected: the script that produced them fell through to the real
fields, so 83 s median over 77-106 s stands.

The derivation was named so a reader could re-take the measurement, and as
written it returned nothing when followed literally, which defeats the only
reason it was there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LjbhSqQf3pwnPA6MVaWcWL
2026-09-02 13:03:45 -05:00
claudeandClaude Opus 5 8217f53d4a docs(ci): measure the accepted cost instead of comparing it unmeasured
CI / shellcheck + selftests (pull_request) Successful in 1m21s
The accepted-cost paragraph claimed this repo's job made the `edited` trade
worse than in the sibling repos that shipped it first, reasoning from the
toolchain install, the Cargo-driving suites and `timeout-minutes: 20`. The
measurement inverts it: last twelve non-skipped runs here are 83 s median
(77-106), against ~118 s for daniel/gitdan and ~330 s for daniel/emowheel --
this is the cheapest of the three, and 20 minutes is a hang ceiling, not a
duration. This PR's own runs measured 87 s and 90 s.

The comparison is dropped rather than re-pointed; the absolute figure replaces
it, with the derivation named (`run_started_at` to `updated_at` off the Actions
API) so a reader can re-take it. The neighbouring concurrency claim was read
off this repo's own file and is unchanged -- it was the checked half of a
paragraph whose other half was not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LjbhSqQf3pwnPA6MVaWcWL
2026-09-02 12:55:24 -05:00
claudeandClaude Opus 5 5b6acd7b3b docs(ci): date the cancellation observation to 1.26.0, not the current version
CI / shellcheck + selftests (pull_request) Successful in 1m30s
The concurrency note said "this Gitea (1.26.0)" while `GET /version` now
returns 1.27.2 — the instance was upgraded on 2026-08-25/26 (daniel/gitdan's
runbook, gitdan#46). Swapping the number would have asserted the cancellation
was observed on 1.27.2, which nobody has checked: the runs it cites were seen
before the upgrade. The comment now dates the observation to 1.26.0, names the
current version, and says persistence is unverified — which is why every commit
gets its own group rather than trusting `cancel-in-progress`. That reasoning is
unchanged; only the implied currency was wrong.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LjbhSqQf3pwnPA6MVaWcWL
2026-09-02 12:49:56 -05:00
claudeandClaude Opus 5 b791896e03 fix(ci): trigger on edited so un-drafting actually lifts the draft skip
CI / shellcheck + selftests (pull_request) Successful in 1m27s
`ready_for_review` does not exist as a pull_request action on this Gitea, so
it never fired and the draft skip never lifted: a PR opened as `WIP:` carried
its skip decision to merge unless a later push happened to create a run. Draft
here is not a persisted column — it is the `WIP:` title prefix, derived by
`issue.IsWorkInProgress` — so un-drafting is a title edit, which fires a plain
`edited`.

Ported from daniel/emowheel commit 08da820 (see daniel/emowheel#71) and
daniel/gitdan (see daniel/gitdan#92), where the identical change landed first.

Accepted cost: Gitea populates no `changes` field for a title-or-body edit, so
the workflow cannot tell an un-drafting edit from an ordinary body PATCH, and
every body edit on a non-draft PR now starts a real run. That is a heavier
trade here than in the sibling repos -- this job installs two Rust toolchains
and drives a real Cargo, at `timeout-minutes: 20` -- but the `concurrency:`
block groups `pull_request` runs on `github.ref` with `cancel-in-progress:
true`, so a burst of edits collapses to one run. Still-draft PRs are unchanged.

Docs: the workflow comment is rewritten and shortened, and README's Development
section retires the empty-commit workaround it prescribed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LjbhSqQf3pwnPA6MVaWcWL
2026-09-02 12:34:24 -05:00
claude 14bea98626 Merge pull request 'fix(hardlink): name the mutable set directly, under both build-dir layouts' (#16) from fix/layout-v2-selection into main
CI / shellcheck + selftests (push) Successful in 1m46s
2026-08-27 20:03:40 +00:00
claude 4bb880b7a7 chore(ci): trigger CI after un-WIP
CI / shellcheck + selftests (pull_request) Successful in 1m21s
2026-08-27 14:52:11 -05:00
claude fb3c72aa28 docs(hardlink): withdraw the nlink claim, which asserted more than was measured
CI / shellcheck + selftests (pull_request) Skipped
The comments said hardlink count at link time was ruled out as the
discriminator between a rewritten executable and an intact one, citing a
lib+bin crate whose two test binaries were both `nlink == 1` and appeared to
behave differently. Re-checked on review: the intact one had not been rebuilt
at all — same content, same inode — so it demonstrated nothing, and forcing
both to rebuild rewrote both.

What was actually observed is narrower and now says so: every executable
measured intact had an uplift hardlink twin Cargo must re-create anyway, every
one measured rewritten had none, and whether the twin is the mechanism or a
correlate was not determined. The rule does not rest on the answer — exempting
twinned executables would recover none of the bytes this change newly copies.

This PR exists because a claim outlived its evidence; it should not ship one.
2026-08-27 14:50:53 -05:00
claude 0553a6b956 fix(hardlink): resolve an ambiguous out/ toward unsharing, and pin the linked-output scenario on a shape that exhibits it
CI / shellcheck + selftests (pull_request) Skipped
Two review findings on #16.

The `out/` discriminator keyed on Cargo's record of a build-script execution,
which Cargo writes only AFTER the script exits successfully. A build script
that populates OUT_DIR and then fails leaves a unit with no record at all, so
its OUT_DIR read as a compile unit's artifact directory and stayed shared —
a regression against the old `-name build` selection, which real-copied that
state by construction. Reproduced on cargo 1.93.1 stable.

An `out` directory now stays shared only when two independent signals agree:
it holds an `.rlib`/`.rmeta` of its own, and its unit carries no execution
record. Either one missing real-copies it. The cost is unchanged to the byte —
the newly-unshared directories hold only executables and `*.d`, both already
privately owned by the file rules.

The live linked-test-binary scenario was built on the lib+bin probe crate,
whose test binaries relink to a fresh inode — a shape gitdan-actions#17 records
as measured safe. Both halves passed green against the unfixed selection on
dep-info mutations the previous scenario already covers. It now builds a
bin-only crate with a unit test, reads only executables, and skips loudly with
a warning rather than passing quietly if the toolchain does not exhibit the
rewrite at all.

Also: drop a clause asserting the linker writes in place "whenever the path has
no other hard link", which this change's own evidence denies; move the
load-bearing comment block back above `unshare_mutable_paths`; correct a
superseded 99.998% figure; and record both cost rows in the README rather than
only the flattering whole-tree one.
2026-08-27 14:39:29 -05:00
claude e5b26a9368 fix(hardlink): name the mutable set directly, under both build-dir layouts
CI / shellcheck + selftests (pull_request) Skipped
`unshare_mutable_paths` selected `.fingerprint` and `build` directories. Under
Cargo's build-dir layout v2 the first clause matches nothing and the second
matches the whole tree, because v2 regroups artifacts under `build/` alongside
the metadata. Measured on one scratch crate: 39.3% of the tree real-copied
under v1, 99.996% under v2.

The selection now names the mutable set rather than the container it used to
live in: fingerprint directories under either spelling, layout v2's `run/`
directories, layout v1's loose build-script run metadata, and the `out`
directories that are a build script's OUT_DIR rather than a compile unit's
artifact directory. The two are told apart structurally, by Cargo's record of
the build-script execution sitting beside the OUT_DIR and nowhere else.

Verifying that turned up a second, layout-independent hazard: a linked
executable is written through whatever inode is already at its path, so a
`cargo test --no-run` inside a `cp -al` clone rewrites the source's own test
binary. Reproduced on cargo 1.93.1 stable, 1.96.0-nightly, 1.98.0-nightly and
1.100.0-nightly, under both layouts. Every executable is now real-copied;
`.rlib`, `.rmeta` and `incremental/` are what stay shared.

hardlink-clone-selftest.sh gains two file-only layout fixtures that pin the
partition in both directions without a compiler, and a live scenario that
relinks a test binary.
2026-08-27 14:12:01 -05:00
claude 4114996954 Merge pull request 'fix(hardlink): content freshness moved switches, it was not withdrawn' (#15) from chore/checksum-freshness into main
CI / shellcheck + selftests (push) Successful in 1m19s
2026-08-27 03:20:03 +00:00
claude fa3cef53e0 docs(ci): the nightly does enable content freshness, as of 2026-08-26
CI / shellcheck + selftests (pull_request) Successful in 1m36s
The README's CI section still described the nightly toolchain step as
buying nothing: "which no nightly currently enables, so it is skipped and
the step is kept only against the day upstream restores it". That
sentence predates this branch and states as fact the exact reading the
rest of the PR retracts.

Three things in this PR falsify it. The corrected `env:` block at
README.md:110-117 records that since cargo PR #17382 (2026-08-22) the
`-Z` gate only unlocks the feature and `build.fingerprint` selects it, so
setting both turns it on. The corrected workflow comment in
.gitea/workflows/ci.yaml says "as of 2026-08-26 it does". And this
branch's own CI run printed `=== checksum-freshness mode: on ===` and
`hardlink-clone-selftest: 4 assertions passed` on 1.100.0-nightly
(787af2b8c 2026-08-25) — the scenario is not skipped, it runs.

The paragraph also carried no date, which is the failure mode every other
block this PR touched was rewritten to prevent. The replacement is dated
and names the toolchain, matching the corrected blocks elsewhere.

It deliberately stops short of "the scenario always runs": the suite
still settles the question by experiment on every run and still skips
loudly when it cannot measure, so the step is not unconditionally
exercised. Saying otherwise would trade one overclaim for its mirror.

Docs-only; no behaviour change.
2026-08-26 20:02:02 -05:00
claude 554310186f fix(hardlink): content freshness moved switches, it was not withdrawn
CI / shellcheck + selftests (pull_request) Successful in 1m17s
`unshare_mutable_paths`' comment recorded that upstream had stopped
rewriting `dep-<target>` in place, on a measurement taken against
1.100.0-nightly. It had not. Two unrelated cargo changes landed within
four days of each other and between them moved the switch that turns the
behaviour on and the path it writes to:

  - cargo PR #17382 (2026-08-22) demoted `-Z checksum-freshness` to a
    gate and gave `build.fingerprint` the choice, defaulting to `mtime`.
    Setting only the gate is accepted and does nothing, which is exactly
    the result that was read as a withdrawal.
  - build-dir layout v2 (cargo PR #17354, stable 1.100.0 on 2026-11-12,
    nightly default since 1.99) moved the file from
    `.fingerprint/<unit>/dep-*` to `build/<pkg>/<hash>/fingerprint/dep-*`.

Measured 2026-08-26 on 1.100.0-nightly (e8cb624d5): same toolchain, same
clone procedure, one env var apart — with the gate alone a `cp -al` clone
mutates only the build/ and *.d families; add
`CARGO_BUILD_FINGERPRINT=content` and the source's dep-info file is
mutated through the shared inode again. The hazard is intact.

So the suite now exports both switches, and its strongest scenario runs
again on current nightlies — verified passing against both layouts. Its
control note learned the v2 path too: it looked for the v1 path only, and
so printed "does NOT rewrite ... in place" three lines beneath a listing
that showed the rewrite.

Every claim these comments make is now dated and cited, because the
defect being fixed is a comment that cited one measurement and silently
stopped reproducing.

Layout v2 also drags the artifacts under `build/`, which collapses this
function's real-copy set from 21.3% of a target dir to 99.998%. That is a
live cost, not a correctness problem, and it is filed as gitdan-actions#14
rather than fixed here.

Part of daniel/gitdan#62.
2026-08-26 19:48:07 -05:00
14 changed files with 2039 additions and 225 deletions
+58 -20
View File
@@ -5,27 +5,27 @@ on:
branches: [main]
pull_request:
branches: [main]
# Spelled out only to keep `ready_for_review` in the list — naming any type
# replaces the whole default set, so the other three have to be restated.
# It is inert on this instance (draft state here is the `WIP:` title
# prefix, so un-drafting is a title edit and raises no
# `ready_for_review` action) and costs nothing.
# Spelled out only to keep `edited` in the list — naming any type replaces
# the whole default set, so the other three have to be restated.
#
# The consequence, which is the part that bites: the `if:` guard below is
# evaluated when a run is CREATED, and un-drafting creates no run. A PR
# opened as a draft keeps its skip decision until something else produces
# one. Push an empty commit after un-WIP'ing.
types: [opened, synchronize, reopened, ready_for_review]
# `edited`, not `ready_for_review`: draft here is the `WIP:` title prefix,
# so un-drafting is a title edit and no ready-for-review action is ever
# raised. That edit is what creates the run which lifts the `if:` skip
# below (decided once, when a run is CREATED). Do not swap it back and do
# not drop this to the bare default — either restores the bug. The accepted
# cost and the evidence are in README's Development section.
types: [opened, synchronize, reopened, edited]
# gitdan-ci runs four repos' CI on two capacity slots, and the compiler-backed
# suites below are multi-minute. A superseded run costs a slot in front of
# somebody's build, so drop it.
#
# `push` groups on `github.sha` rather than `github.ref`: a constant per-branch
# group is what let this Gitea (1.26.0) cancel two of daniel/gitdan's merge
# runs outright while `cancel-in-progress` was gated away from `push` entirely
# — see the long note in that repo's ci.yaml for the evidence. Giving every
# commit its own group leaves that behaviour nothing to act on.
# group is what let this Gitea cancel two of daniel/gitdan's merge runs
# outright while `cancel-in-progress` was gated away from `push` entirely —
# see the long note in that repo's ci.yaml. Observed on 1.26.0; the instance
# is 1.27.2 now and whether it persists is unverified, which is why every
# commit gets its own group instead of trusting the flag.
concurrency:
group: ${{ github.workflow }}-${{ github.event_name == 'pull_request' && github.ref || github.sha }}
cancel-in-progress: true
@@ -73,12 +73,17 @@ jobs:
# resolves freshness by CONTENT — the mode where the dep-info file
# carries per-source checksums, which is the mutation that turns a
# hardlink clone into silent stale-artifact reuse rather than a slow
# build. As of 1.100.0-nightly (2026-08-25) no nightly provides it:
# `-Z checksum-freshness` is still accepted and freshness is still
# resolved by mtime, so the suite's probe reports that by name and skips
# the scenario. This step therefore buys nothing today and is kept
# anyway — it costs about twenty seconds, and the day upstream restores
# the behaviour the coverage comes back with no edit here. See daniel/gitdan#62.
# build. This step is what supplies it, and as of 2026-08-26 it does:
# the scenario ran and passed against 1.100.0-nightly.
#
# It briefly did not. Cargo PR #17382 (2026-08-22) demoted
# `-Z checksum-freshness` to a gate and gave `build.fingerprint` the
# choice, defaulting to `mtime`, so the suite — which set only the gate —
# measured a genuine INACTIVE and skipped its strongest scenario. That
# read as "upstream withdrew content freshness" and was written up here
# as this step buying nothing. It was a moved switch, not a withdrawal;
# the suite now exports both and the coverage is back. See
# daniel/gitdan#62 for the investigation.
- name: Install Rust nightly
uses: dtolnay/rust-toolchain@nightly
- name: Install Rust toolchain
@@ -96,3 +101,36 @@ jobs:
# file.
- name: Selftests
run: bash scripts/selftest.sh
release-tag:
name: move v1 to main
# `needs: selftest` is what makes this "after the gate is green": a failed
# selftest skips this job, so v1 never advances onto a broken build. The
# `if:` restricts it to an actual push to main.
#
# No job-level `concurrency:` -- the lease in release-v1.sh already keeps
# v1 from moving backwards, and release-sweep.yaml picks up anything this
# job defers or misses.
needs: selftest
if: ${{ github.event_name == 'push' && github.ref == 'refs/heads/main' }}
runs-on: ubuntu-latest
timeout-minutes: 2
# Requests write access from the run's built-in token -- whether that
# grant actually lets it push here is unobserved until the first merge
# (see README's Versioning section). Without this the checkout below
# still succeeds -- it's the push that would be rejected, which is a red
# job, not a silent no-op.
permissions:
contents: write
steps:
# Full history, so the ancestry checks in release-v1.sh can see how
# this commit relates to v1.
- uses: actions/checkout@v4
with:
token: ${{ secrets.GITHUB_TOKEN }}
fetch-depth: 0
# Releases this run's own commit only while it is still main's tip, and
# only forward -- see release-v1.sh.
- name: Move v1 to this commit if it is still main's tip
run: bash scripts/release-v1.sh merge "${{ github.sha }}"
+69
View File
@@ -0,0 +1,69 @@
name: Release sweep
# Keeps v1 from lagging main when ci.yaml's release-tag job defers or never
# runs. Each tick either finds v1 already covering main's tip and exits, or
# gates the tip exactly as ci.yaml's selftest job does and moves v1 to it. A
# red run here means v1 is behind a main that fails its gate.
#
# Gitea registers schedules from the default branch only, so this fires once
# it is on main.
on:
schedule:
- cron: '*/15 * * * *'
# A tick that arrives while another is still gating waits behind it rather
# than gating the same tip twice.
concurrency:
group: release-sweep
cancel-in-progress: false
jobs:
sweep:
name: move v1 to main if it lags
runs-on: ubuntu-latest
timeout-minutes: 25
permissions:
contents: write
steps:
- uses: actions/checkout@v4
with:
token: ${{ secrets.GITHUB_TOKEN }}
fetch-depth: 0
# A scheduled run's github.sha is main as of the last push, not
# necessarily its tip, so the tip is read here instead.
- name: Check whether v1 lags main
id: check
run: bash scripts/release-v1.sh sweep-check
# Everything below runs only when v1 lags, and gates the tip itself,
# not the commit this run was created from.
- name: Check out main's tip
if: steps.check.outputs.needed == 'true'
run: git checkout -q --detach "${{ steps.check.outputs.tip }}"
- name: Install shellcheck
if: steps.check.outputs.needed == 'true'
uses: taiki-e/install-action@v2
with:
tool: shellcheck
# Same toolchains, same order, as ci.yaml's selftest job.
- name: Install Rust nightly
if: steps.check.outputs.needed == 'true'
uses: dtolnay/rust-toolchain@nightly
- name: Install Rust toolchain
if: steps.check.outputs.needed == 'true'
uses: dtolnay/rust-toolchain@stable
- name: shellcheck
if: steps.check.outputs.needed == 'true'
run: shellcheck -x --source-path=scripts scripts/*.sh
- name: Selftests
if: steps.check.outputs.needed == 'true'
run: bash scripts/selftest.sh
- name: Move v1 to the gated tip
if: steps.check.outputs.needed == 'true'
run: bash scripts/release-v1.sh push "${{ steps.check.outputs.tip }}" "${{ steps.check.outputs.v1 }}"
+236 -51
View File
@@ -29,16 +29,63 @@ on the runner having a single execution slot.
One thing neither project had, and the reason the clone is not a plain
`cp -al`: **a build inside a hardlink clone does mutate the directory it was
cloned from.** Cargo writes real artifacts by replacing them, but writes its
cloned from.** rustc renames its own outputs into place, but Cargo writes its
metadata — and build scripts write their `OUT_DIR` — with a plain truncating
write, straight through the shared inode. Under
`CARGO_UNSTABLE_CHECKSUM_FRESHNESS` the file that gets corrupted is
`.fingerprint/<unit>/dep-<target>`, which holds the per-source checksums that
decide freshness, and the failure is silent stale-artifact reuse rather than a
slow build. `scripts/hardlink-clone-selftest.sh` reproduces it as an explicit
control and asserts the fix. The fix is to hardlink the artifacts (the GB) and
real-copy the metadata (the MB) — about 3.7% of a Bevy-sized target directory,
against 100% for a full copy.
write, straight through the shared inode. When Cargo resolves freshness by
content the file that gets corrupted is its dep-info fingerprint —
`.fingerprint/<unit>/dep-<target>`, or
`build/<pkg>/<hash>/fingerprint/dep-<target>` under Cargo's build-dir layout v2
— which holds the per-source checksums that decide freshness, and the failure
is silent stale-artifact reuse rather than a slow build.
One more family joins them, and it is not metadata: **anything the linker
writes**. rustc writes an `.rlib` or `.rmeta` to a temporary and renames it
into place, but a linked executable is written through whatever inode is
already at its path — so a `cargo test --no-run` in a raw `cp -al` clone
rewrites the source's own test binary. Measured 2026-08-27 on cargo 1.93.1
stable, 1.96.0-nightly, 1.98.0-nightly and 1.100.0-nightly, under both
build-dir layouts.
`scripts/hardlink-clone-selftest.sh` reproduces both as explicit controls and
asserts the fix. The fix is to hardlink what rustc renames into place — the
`.rlib`, `.rmeta` and `incremental/` bulk — and real-copy the metadata and the
linker outputs.
**The selection names that set directly, under either build-dir layout.**
Layout v2 regroups everything per build unit under
`build/<pkg>/<hash>/{fingerprint,out,run}/`, artifacts included, so there is no
`.fingerprint` and no `deps` to key off and `build/` is no longer a proxy for
"metadata" — it is the whole tree. The one place the two layouts genuinely
differ is that under v2 a build script's `OUT_DIR` and a compile unit's rlib
are both a directory called `out`.
**Ambiguity there resolves toward unsharing**, because over-unsharing costs
bytes and under-unsharing costs corruption. An `out` directory stays shared
only when two independent signals agree it is a compile unit's: it holds an
`.rlib`/`.rmeta` of its own, *and* its unit carries no record of a build-script
execution beside it (`run/` under v2, a loose `root-output` under v1). The
execution record alone is not enough — Cargo writes it only after the script
succeeds, so a build script that populates `OUT_DIR` and then fails leaves a
unit that reads as a compile unit. v2 is the nightly default and stabilises in
cargo 1.100.0 on 2026-11-12.
### What it costs
Real-copied share of a 5.5 GB Bevy target directory, before and after the
linker-output rule landed:
| tree | before | after |
|---|---|---|
| **excluding `incremental/`** — the figure to plan against, since the quick-start below sets `CARGO_INCREMENTAL: 0` | **36.4%** | **57.0%** |
| whole tree, `incremental/` included (a local dev checkout, not CI) | 9.0% | 14.1% |
The first row is the one a CI consumer gets. The increase is the linker-output
rule, not the layout work: on a scratch crate the layout fix alone takes v2
from 99.996% to 0.2%.
The copy is paid per clone and does not amortise — a fresh `cp -al` leaves
every file with `nlink >= 2`, so the `-links +1` filter cannot skip anything —
and a clone happens twice per job, once seeding and once publishing.
---
@@ -64,10 +111,14 @@ jobs:
steps:
- uses: actions/checkout@v4
with:
# REQUIRED. The mtime restore walks every commit that ever touched a
# tracked file; a depth-1 checkout makes every file resolve to the
# tip commit and the cache stops working. The action fails loudly
# rather than silently degrading if this is missing.
# REQUIRED, for two things. The mtime restore walks every commit
# that ever touched a tracked file; a depth-1 checkout makes every
# file resolve to the tip commit and the cache stops working, and
# the action fails loudly rather than silently degrading. The prune
# also decides whether a branch has been merged by asking this
# checkout for ancestry, which a shallow one cannot answer — there
# it withholds that half of the pass and says so, so merged-but-
# undeleted branches keep their caches.
fetch-depth: 0
- uses: https://gitdan.com/daniel/gitdan-actions/cargo-cache@v1
@@ -98,10 +149,14 @@ env:
CARGO_INCREMENTAL: 0 # per-run bloat on a persistent volume
CARGO_PROFILE_DEV_DEBUG: line-tables-only
CARGO_PROFILE_TEST_DEBUG: line-tables-only
# Nightly only. Content-addressed freshness instead of mtime-based — a
# strictly stronger guarantee, complementary to the mtime restore (which
# still covers directory-form `rerun-if-changed` build-script watches).
# Nightly only, and BOTH are needed. Content-addressed freshness instead of
# mtime-based — a strictly stronger guarantee, complementary to the mtime
# restore (which still covers directory-form `rerun-if-changed` build-script
# watches). Since cargo PR #17382 (2026-08-22) the `-Z` gate below only
# unlocks the feature; `build.fingerprint` selects it and defaults to
# `mtime`, so the gate on its own is accepted and does nothing.
CARGO_UNSTABLE_CHECKSUM_FRESHNESS: "true"
CARGO_BUILD_FINGERPRINT: "content"
```
---
@@ -192,8 +247,8 @@ not collapsed into one:
pass still deciding about one is never mistaken for a pass that died holding
it. That settle window is a bound rather than a construction, and it is the
only part of this that is. Until the rest of it was structural it was merely
policy — snapshots belong to protected refs, protected refs are never
eviction candidates — a property held by vigilance rather than by
policy — snapshots belong to protected refs, and a protected ref's snapshot
is never an eviction candidate — a property held by vigilance rather than by
construction.
- **Every other way the source can change mid-clone is detected, not
prevented.** A `seed-fallback-dir` pointing at a directory something else
@@ -212,14 +267,82 @@ not collapsed into one:
disk, normally as the cache restored under its own name and otherwise as one
set aside for a later pass to reclaim.
**Eviction** runs three passes: caches for branches that no longer exist on
origin are removed unconditionally; then, only if free space is under the
threshold, live caches are evicted oldest-first; then, as a last resort, this
run's own cache. Protected refs and any cache held open by a running job are
never candidates. Within the pressure pass, `target-*` directories are evicted
before `snapshot-*` ones — the reverse of the obvious order, because a
snapshot is hardlinked to everything cloned from it, so removing one frees
almost no real bytes while costing every future PR its warm start.
**Eviction** runs three passes, ahead of the seed so that what it frees is
available to the clone that follows. Caches for branches that are DEAD are
removed unconditionally; then, only if free space is under the requirement,
live caches are evicted oldest-first; then, as a last resort, this run's own
cache. A protected ref's *snapshot*, the source this run is about to clone,
and any cache held open by a running job are never candidates in the pressure
or self-clear passes. A protected ref's own *target* dir is an ordinary
pressure-pass candidate, since it is a convenience cache the publisher's next
run reseeds from the snapshot — but it is excluded from the liveness pass
alone, because a branch's tip is trivially an ancestor of itself, and without
that exclusion the merged-branch signal would read a publisher's own target
dir as merged into itself and delete it every run, unconditionally. Within the
pressure pass, `target-*` directories are evicted before `snapshot-*` ones:
losing a target dir is cheap for exactly that reseeding reason, while
evicting a snapshot forces every subsequent PR to start cold and, measured on
the live volume (daniel/zemyna#1073), frees real disk rather than the
near-nothing a shared-inode hardlink clone would suggest — `publish-snapshot.sh`
unshares every executable after its `cp -al`, and executables are most of the
tree by bytes.
**A branch is dead in two ways, and neither signal makes the other
redundant.** The first is that the branch is gone from origin. The second is
ancestry: a branch still on origin whose tip is an ancestor of a protected
branch's tip holds no commit that branch does not, so its cache will never be
read again and goes in the same pass.
**Turn delete-on-merge on** (`default_delete_branch_after_merge`, per repo) —
it is the setting this scheme is cheapest under, because a deleted branch is
decidable from `ls-remote` alone, with no checkout, no objects and no walk.
The ancestry signal is what covers the rest, and the rest is not a corner:
- **Every branch merged before the setting was turned on.** They stay on
origin forever; nothing retroactively deletes them. zemyna carried 48 of
them at the time the setting was enabled, and ancestry is the only thing
that reclaims a cache dir belonging to any of them.
- **Every merge the deletion declines or fails.** Gitea's delete is
best-effort and silent: it declines for a protected branch and for one
another open PR still uses, and an API merge that omits the flag — which
`tea pulls merge` does — simply never asks.
- **Repos that have not enabled it**, which is the default.
That is the difference between reclaiming nothing and reclaiming a 40 GB
directory per merged PR on a full volume (issue 20).
Ancestry is answered from the commits in the job's own checkout, so it is only
answered where they are there to answer it — and "cannot tell" is never folded
into "dead", at either granularity. A shallow checkout withholds the signal
entirely, because a missing object is its normal case rather than evidence; a
single branch whose tip is not in the checkout is kept, with a warning naming
it. A squash or rebase merge leaves no ancestry at all, so its branch reads as
live until it is deleted. All three are missed reclamations, which cost disk;
the alternative direction costs a branch its cache while it is still being
built on.
**How much free space is enough is measured, not chosen.** What the seed is
about to do is hardlink-clone a snapshot and then real-copy that clone's
*mutable set* — the dep-info, build-script metadata and linked outputs that a
build would otherwise write through a shared inode. The rest stays hardlinked
and costs nothing. So the requirement is derived per run, from that set,
measured off the very snapshot the seed will read using the same enumeration
`unshare_mutable_paths` copies from; the pass evicts oldest-first until it is
met and then stops. A percentage of the volume cannot express this: how much a
clone needs is a property of the snapshot, and a threshold sized for a
different failure is one that never fires before the seed refuses. `cp -al`
still materialises every directory for real, and the unshare stages each
subtree through a sibling copy, so the measurement carries a margin —
`CACHE_CLONE_HEADROOM_PERCENT` and `CACHE_CLONE_HEADROOM_FLOOR_KB`, both
hand-written defaults, both erring toward asking for more.
A run that cannot reach the derived requirement after evicting everything
eligible **fails, naming the shortfall and every directory it kept instead**.
The seed would otherwise fail seconds later, reporting a staging path and
nothing about which cache was holding the space — which is the failure this
pass now pre-empts. `min-free-percent` is an additional floor on top and
nothing more: it defaults to `0`, it only ever raises the requirement, and
falling short of it is still a warning and a self-clear rather than a failure.
**File mtimes.** `actions/checkout` stamps every file with "now", which makes
every crate look changed to Cargo's mtime-based freshness check — a persistent
@@ -360,11 +483,11 @@ directory; the line just doesn't say which.
|---|---|---|
| `cache-root` | `/cache` | mount point of the persistent volume inside the job container |
| `cache-lineage` | *(empty)* | one directory level under `cache-root`, for a second job building the same ref for a different target or profile — see [Multiple jobs in one workflow](#multiple-jobs-in-one-workflow) |
| `protected-branches` | `dev main` | refs that publish snapshots and are never evicted |
| `min-free-percent` | `10` | prune when free space drops below this |
| `protected-branches` | `dev main` | refs that publish snapshots, whose snapshots are never evicted (their target dirs are ordinary pressure-pass candidates) |
| `min-free-percent` | `0` | an ADDITIONAL free-space floor, as a percentage of the volume. The gate is derived per run from what the seed is about to clone; this only ever raises it |
| `restore-mtimes` | `true` | restore tracked-file mtimes from git history |
| `prune` | `true` | run the eviction pass |
| `liveness-prune` | `true` | within eviction, remove caches for branches gone from origin |
| `prune` | `true` | run the eviction pass — before the seed, so what it frees is available to the clone |
| `liveness-prune` | `true` | within eviction, remove caches for branches that are dead: gone from origin, or merged into a protected branch |
| `own-ref` | *(auto)* | override; defaults to `github.head_ref`, else `github.ref_name` |
| `base-ref` | *(auto)* | override; defaults to `github.base_ref` (empty on push) |
| `seed-fallback-dir` | *(empty)* | absolute path to seed from when no snapshot exists — for migrating off an existing flat cache |
@@ -511,13 +634,39 @@ entry, another permission — is a `v2`, not a `v1` move. Everything else moves
`v1`: correctness fixes, new optional inputs, and anything internal to
`scripts/`.
**Moving the tag is a release step, and it is the operator's.** Merging to
`main` ships nothing to anybody. `v1` is a lightweight tag and does not follow
a branch, so until it is re-pointed every consumer keeps fetching the commit it
already named, whatever `main` now says. The gap is deliberate: re-pointing
`v1` changes what another repository's CI executes on its next run, so it is a
decision taken once, knowingly, after the merge — never something a merge does
by itself.
**Moving the tag is automatic, gated on the same build that gates a PR.** Two
jobs move it, both through `scripts/release-v1.sh`, both with the run's
built-in `GITHUB_TOKEN`:
- **`release-tag`** in `.gitea/workflows/ci.yaml` runs on every push to
`main`, `needs: selftest`, and moves `v1` to that run's own commit — but only
while that commit is still `main`'s tip. A run whose merge has already been
overtaken defers, as a successful no-op, rather than release a commit it
never gated.
- **`release-sweep.yaml`** runs every 15 minutes. When `v1` already points at
`main`'s tip or a descendant of it, it exits after a checkout and one
comparison. Otherwise it runs the same shellcheck and selftests against the
tip and moves `v1` there only if they pass.
Both jobs request `contents: write` on the run's built-in token, and that
grant is now **verified**: PR #28's own merge ran `release-tag` successfully
and moved `v1` to that merge's commit. The grant is still capped by the
repo's and owner's maximum token permissions, and branch/tag protections on
`v1` can't be read without admin access. A rejected push is a red job, not a
silent no-op.
So `v1` trails a green `main` by at most about one sweep interval plus one
selftest run, and **a `main` that fails its gate shows up as a red sweep on every
tick until it is fixed** — as does a push the token is not allowed to
make. Both jobs push with `--force-with-lease` on the `v1` they read, so
neither can move `v1` backwards over the other; a job that loses the lease to
a newer `v1` finishes green.
This used to be a manual step, treated as a deliberate release decision taken
once, knowingly, after the merge — in practice it was still forgotten
(gitdan-actions#27): PR #25 merged to `main` and `v1` stayed on the previous
release until someone asked whether it had moved. The manual form below is
still the recovery path, for when neither job can push:
```bash
git fetch origin
@@ -528,9 +677,8 @@ git ls-remote --tags origin v1 # must equal git rev-parse origin/main
**Downstream** are emowheel, which pins `cargo-cache@v1` and
`cargo-cache-publish@v1` across its CI workflow, and zemyna, migrating to the
same pin. Both pick a move up on their next run with no change on their side,
which is the whole point of the moving pointer and also the reason the move is
not automatic.
same pin. Both pick a move up automatically on their next run with no change
on their side, which is the whole point of the moving pointer.
---
@@ -545,16 +693,52 @@ bash scripts/selftest.sh --fast # fixture-only suites, no compiler
Both run in CI — `.gitea/workflows/ci.yaml`, one job, on pushes to `main` and
on PRs that were non-draft when the run was created. It installs shellcheck
and both a stable and a nightly Rust toolchain (nightly so
`hardlink-clone-selftest.sh` can run its content-freshness scenario — which no
nightly currently enables, so it is skipped and the step is kept only against
the day upstream restores it) and references no credentials; the scratch
workspaces the compiler-backed suites build use path dependencies only, so
nothing reaches crates.io. It runs the full suite rather than `--fast`,
`hardlink-clone-selftest.sh` can run its content-freshness scenario, which as
of 2026-08-26 a nightly does enable — 1.100.0-nightly (787af2b8c 2026-08-25)
resolves freshness by content given both `CARGO_UNSTABLE_CHECKSUM_FRESHNESS`
and `CARGO_BUILD_FINGERPRINT: content`, per cargo PR #17382; the suite still
settles that by experiment on every run and skips the scenario loudly when it
cannot measure) and references no credentials; the scratch workspaces the
compiler-backed suites build use path dependencies only, so nothing reaches
crates.io. It runs the full suite rather than `--fast`,
because the two compiler-backed suites are the ones that check this scheme
against real Cargo instead of against a fixture. Draft (`WIP:`-titled) PRs
skip it, and un-drafting does **not** un-skip them — the guard is evaluated
when a run is created and un-drafting creates none, so push an empty commit
after un-WIP'ing.
against real Cargo instead of against a fixture.
**Draft (`WIP:`-titled) PRs skip it; un-drafting un-skips them, through
`edited`** — no empty commit needed. The skip is decided when a run is
*created*, so lifting it needs an event that creates one, and un-drafting on
this Gitea is a title edit: `edited` is in the workflow's `pull_request` types
for exactly that reason. `ready_for_review` held that slot first and never
fired — this Gitea has no draft column and no ready-for-review event at all,
draft being computed from the title prefix — so a PR opened as `WIP:` carried
its skip decision all the way to merge unless some later push happened to
create a run. Do not swap the type back and do not drop the `types:` list to
its bare default; either restores the bug.
**Accepted cost: a body edit on an already-non-draft PR now triggers a real
run.** Gitea populates no `changes` field for a title-or-body edit, unlike
GitHub, so the workflow cannot tell the edit that un-drafts a PR from an
ordinary body PATCH — a closing-reference fix-up, say. The price is around
**90 seconds** of a runner shared across four repos on two capacity slots: the
last twelve non-skipped runs of this job, `started_at` to `completed_at` off
the Actions API, are 83 s median over 77–106 s. `timeout-minutes: 20` is a
ceiling for a hung suite, not a duration. It is also bounded: the workflow's
`concurrency:` block groups `pull_request` runs on `github.ref` with
`cancel-in-progress: true`, so a burst of edits collapses to one run rather
than N. A still-draft PR pays nothing extra — the `if:` guard skips those
exactly as before.
The evidence for `edited` comes from `daniel/emowheel`, which hit the identical
bug, shipped the identical wrong fix, and corrected it in commit `08da820` (see
`daniel/emowheel#71`): Gitea 1.27.2's `HookIssueAction` enum has no
`ready_for_review` entry and no notifier emits one, while
`issue.IsWorkInProgress` derives draft from the title. Live since: emowheel PR
#141 was un-drafted at 19:01:13 on 2026-08-31 and run 2801 was created two
seconds later on the **same** head SHA as the two runs skipped before it — a
run created by the title edit alone, with no push, that executed and passed.
`daniel/gitdan` ported the same one-word change. That it behaves the same way
*here* has not been demonstrated in this repo; the next `WIP:` PR opened
against `main` is the test.
This repo is consumed by three other repos' CI at `@v1`, a moving tag, so a
change here reaches all of them at once. That is what the gate is for.
@@ -562,10 +746,11 @@ change here reaches all of them at once. That is what the gate is for.
| suite | covers |
|---|---|
| `cache-root-selftest.sh` | that a lineage nests one level and nothing else moves: no lineage resolves byte-for-byte to the cache root, two lineages on one cache key get disjoint target dirs, seed/publish/prune all stay inside their own lineage, a PR layers over its own lineage's base snapshot — **and one rejection per lineage name a reader elsewhere would stop seeing**, plus the publish-side mismatch guard |
| `hardlink-clone-selftest.sh` | that a build in a clone cannot mutate its source — with a control proving a raw `cp -al` does. Needs a real compiler, **and a nightly that actually resolves freshness by content for its last scenario**: the source's-next-build check reasons about content rather than mtime, so under mtime freshness it would assert a bug. Whether the toolchain does is settled by experiment on a throwaway crate, not by asking it — 1.100.0-nightly accepts `-Z checksum-freshness` and rebuilds on mtime anyway. The experiment reports **three** outcomes, not two: active, measured-inactive, and *not measured*. Its answer codes are `0` and `3`, deliberately clear of every status bash generates for its own errors — so nothing that goes wrong inside the probe, including an expansion failure no guard can catch, can be read as an answer. The scenario is skipped for the last two alike, but a failure to measure is never reported as a measurement. The control also reports which mutation families the running Cargo exhibits — a note, not an assertion, since that set moves upstream. |
| `hardlink-clone-selftest.sh` | that a build in a clone cannot mutate its source — with a control proving a raw `cp -al` does. Needs a real compiler, **and a nightly that actually resolves freshness by content for its last scenario**: the source's-next-build check reasons about content rather than mtime, so under mtime freshness it would assert a bug. Whether the toolchain does is settled by experiment on a throwaway crate, not by asking it — accepting `-Z checksum-freshness` stopped implying it on 2026-08-22, when cargo PR #17382 demoted the flag to a gate and gave `build.fingerprint` (default `mtime`) the choice; the suite now exports both and 1.100.0-nightly measures ACTIVE again. The experiment reports **three** outcomes, not two: active, measured-inactive, and *not measured*. Its answer codes are `0` and `3`, deliberately clear of every status bash generates for its own errors — so nothing that goes wrong inside the probe, including an expansion failure no guard can catch, can be read as an answer. The scenario is skipped for the last two alike, but a failure to measure is never reported as a measurement. The control also reports which mutation families the running Cargo exhibits — a note, not an assertion, since that set moves upstream. It also pins the SELECTION itself against both of Cargo's build-dir layouts, from two file-only fixtures that need no compiler — so the layout the installed Cargo does not happen to write is still covered — and asserts the partition in both directions: every file Cargo rewrites in place is privately owned, and every `.rlib`/`.rmeta` still shares its inode. The second half is what the old suite never checked beyond `shared > 0`, and it is what a layout change silently inverts. |
| `seed-target-dir-selftest.sh` | seed-source preference, lock-file stripping, two jobs racing on one cache key, **and one scenario per check a hardlink clone is validated against**: a source rotated wholesale, a subtree silently lost from the walk, a copy that reports failure over a tree both other checks read as whole, and a source identity that resolved at neither end — plus a staging tree that could not be privately owned being discarded rather than published, and the publisher's log showing it waited on the consumer's own reader-lock marker before reclaiming a rotated snapshot |
| `publish-snapshot-selftest.sh` | the atomic swap, that a live consumer survives a republish, and the publisher's side of the rotation race: deferred reclamation under a live reader, and its sweep once the reader is gone |
| `prune-cache-selftest.sh` | liveness, protection, locking, eviction order, self-clear, **and that a cache a job claims *inside* the check-to-unlink window survives it** — against a real scratch `origin` |
| `prune-cache-selftest.sh` | liveness in both its forms — a branch deleted from origin, and one still on it whose tip is already merged — plus protection, locking, eviction order, self-clear, **that a cache a job claims *inside* the check-to-unlink window survives it**, and that a requirement derived from the clone's mutable set evicts exactly enough and then fails rather than under-delivering. Against a real scratch `origin`, including a genuinely shallow clone of it and a `df` that answers from the fixture's own size, since a fixed one cannot show a pass stopping |
| `release-v1-selftest.sh` | that `v1` reaches `main`'s tip only through a gate and never moves backwards: the sweep's no-op, tag and red-gate cases, the stranded-defer trace the sweep exists to recover, each lost-lease outcome — a newer `v1` skipped cleanly (with a control showing an unleased push steps it back), an older one retried, an unrelated one and a server rejection red — and a `v1` hand-placed on an unrelated commit before any push is ever attempted, also red. Against a real scratch `origin`; the other writer is sequenced between check and push, not raced |
| `restore-mtimes-selftest.sh` | the merge hazard and the watermark that closes it, including the two-jobs-one-namespace case. Needs a real compiler. |
Every suite runs the actual script, not a reimplementation of its logic, and
+57 -20
View File
@@ -24,14 +24,25 @@ inputs:
default: ''
protected-branches:
description: >-
Space-separated refs that publish snapshots and are never evicted.
Space-separated refs that publish snapshots. A protected ref's
snapshot is never evicted; its own target dir is an ordinary
pressure-pass candidate, reseeded from the snapshot on its next run.
These are the branches PR caches layer over.
required: false
default: 'dev main'
min-free-percent:
description: 'Prune when free space on the cache volume drops below this percentage.'
description: >-
An ADDITIONAL free-space floor, as a percentage of the cache volume.
The prune's own requirement is derived per run from what the seed is
about to clone — the mutable set it has to real-copy out of the source
snapshot, which is a property of that snapshot and not of the volume —
and this floor only ever raises it. 0, the default, leaves the derived
requirement as the only gate. Set it to keep headroom for something
other than the clone (the build's own output, another job on the same
volume); it will not make the clone fit, because it does not know how
big the clone is.
required: false
default: '10'
default: '0'
restore-mtimes:
description: >-
Restore every tracked file's mtime from git history. Requires a
@@ -40,13 +51,20 @@ inputs:
required: false
default: 'true'
prune:
description: 'Run the eviction pass (dead-branch liveness + disk pressure).'
description: >-
Run the eviction pass (dead-branch liveness + disk pressure). It runs
BEFORE the seed step, so what it frees is available to the clone that
step makes.
required: false
default: 'true'
liveness-prune:
description: >-
Within the prune pass, remove caches for branches that no longer exist
on origin. Set to false on a runner that cannot reach origin.
Within the prune pass, remove caches for branches that are dead: gone
from origin, or still on origin with a tip already merged into a
protected branch. The second signal is what reclaims anything at all on
a forge that keeps branches after merge, and it needs the protected
branches' commits in the checkout — a shallow one withholds it and says
so. Set to false on a runner that cannot reach origin.
required: false
default: 'true'
own-ref:
@@ -169,6 +187,39 @@ runs:
echo "CI_WATERMARK_FILE=${WATERMARK}"
} >> "$GITHUB_ENV"
# Runs BEFORE the seed, which is the only order in which its work can
# help: the eviction it performs is what makes room for the clone the seed
# step is about to make, and the requirement it evicts against is measured
# off the snapshot that clone will read. Running afterwards — where this
# step used to be — meant every run freed space for the NEXT one and the
# seed met whatever the last run happened to leave.
#
# Two things this ordering has to be safe against, and is:
#
# The source it is about to read is excluded from every pass by name
# (see protected_reason in prune-cache.sh), so pass 1 cannot take the
# snapshot out from under the seed that follows it.
# A concurrent job's target dir carries its lock from the instant it
# appears under its final name — the seed writes it into the staging
# tree before the rename — so there is no window in which running this
# earlier sees an unlocked directory somebody is using.
- if: ${{ inputs.prune == 'true' }}
shell: bash
env:
STALE_LOCK_SECONDS: ${{ inputs.stale-lock-seconds }}
CACHE_LIVENESS: ${{ inputs.liveness-prune }}
run: |
set -euo pipefail
SCRIPTS=$(cd "${{ github.action_path }}/.." && pwd)/scripts
bash "${SCRIPTS}/prune-cache.sh" \
"${{ steps.resolve.outputs.cache-root }}" \
"${{ steps.resolve.outputs.target-dir }}" \
"${{ inputs.protected-branches }}" \
"${{ inputs.min-free-percent }}" \
"${{ steps.resolve.outputs.cache-key }}" \
"${{ steps.resolve.outputs.base-key }}" \
"${{ inputs.seed-fallback-dir }}"
# Seeds this ref's target dir from the base's published snapshot. See
# scripts/seed-target-dir.sh — the staging-then-atomic-rename is what
# makes concurrent jobs sharing one cache key safe by construction rather
@@ -228,17 +279,3 @@ runs:
CARGO_TARGET_DIR: ${{ steps.resolve.outputs.target-dir }}
CI_WATERMARK_FILE: ${{ steps.resolve.outputs.watermark-file }}
run: bash "$(cd "${{ github.action_path }}/.." && pwd)/scripts/restore-mtimes.sh"
- if: ${{ inputs.prune == 'true' }}
shell: bash
env:
STALE_LOCK_SECONDS: ${{ inputs.stale-lock-seconds }}
CACHE_LIVENESS: ${{ inputs.liveness-prune }}
run: |
set -euo pipefail
SCRIPTS=$(cd "${{ github.action_path }}/.." && pwd)/scripts
bash "${SCRIPTS}/prune-cache.sh" \
"${{ steps.resolve.outputs.cache-root }}" \
"${{ steps.resolve.outputs.target-dir }}" \
"${{ inputs.protected-branches }}" \
"${{ inputs.min-free-percent }}"
+334 -36
View File
@@ -317,39 +317,186 @@ _unshare_files() {
xargs -0 -r -n 64 bash -c 'rc=0; for f; do cp -p -- "$f" "$f.unshare.$$" && mv -f -- "$f.unshare.$$" "$f" || rc=1; done; exit $rc' _
}
# True when a directory holds a compiled library artifact of its own.
#
# The glob is left unquoted and unmatched-glob-safe on purpose: with nullglob
# off an unmatched pattern stays literal and the `-e` test fails, which is the
# answer wanted.
_holds_compiled_artifact() {
local f
for f in "$1"/*.rlib "$1"/*.rmeta; do
[ -e "$f" ] && return 0
done
return 1
}
# The directory names that select a mutable subtree, as one find predicate.
#
# Named once because THREE readers have to agree on it: the selection in
# _mutable_dirs, the prune that skips those subtrees when it sizes the file
# rules, and anything later that measures what a clone will cost. Two
# spellings of this list would size a different tree than the one copied, and
# the direction that fails is silent — an under-measured clone runs out of
# disk mid-unshare, which is gitdan-actions#20.
_MUTABLE_DIR_NAMES=( -name .fingerprint -o -name fingerprint -o -name run -o -name out )
# _mutable_file_rules <fn>
#
# The FILE half of the mutable set, applied one rule at a time:
#
# <fn> <label> <maxdepth|-> <find-predicate...>
#
# Same reason as the array above — `unshare_mutable_paths` copies these and
# `mutable_set_kb` measures them, and a rule that exists in only one of the
# two is exactly the under-estimate the headroom gate cannot survive. The
# maxdepth is a separate field because GNU find wants it ahead of every other
# predicate, so it cannot live inside the predicate vector.
#
# `-type f` is each caller's to add: the sizer needs it inside the `-o`
# alternation it builds, the copier ahead of it.
_mutable_file_rules() {
local fn="$1"
"$fn" 'dep-info files' - -name '*.d' || return 1
# Layout v1's build-script run metadata, which v2 groups under `run/` and v1
# leaves loose in the run unit's directory. `invoked.timestamp` is empty and
# carries its meaning in its mtime, which a shared inode carries too.
"$fn" 'build-script run metadata' - \
\( -name output -o -name root-output -o -name stderr -o -name invoked.timestamp \) || return 1
# Linked outputs. Unlike an rlib or an rmeta — which rustc writes to a
# temporary and renames into place — an executable or shared object is
# written by the LINKER, and the linker writes THROUGH an existing inode.
# Measured 2026-08-27 on cargo 1.93.1 stable, 1.96.0-nightly, 1.98.0-nightly
# (layout v1) and 1.100.0-nightly (e8cb624d5, layout v2), mold and the
# default linker alike: a `cargo test --no-run` binary in a `cp -al` clone
# rewrote the SOURCE's copy of itself in place, under both layouts.
#
# What separates that from the executables measured INTACT is not
# established. Every intact case observed was one Cargo has to re-create
# anyway to maintain an uplift hardlink — a bin target's
# `deps/<bin>-<hash>`, twinned at `<profile>/<bin>`. Whether the twin is the
# mechanism or a correlate of it was not determined, and the rule below does
# not depend on the answer: exempting twinned executables would recover no
# bytes this function newly copies. See gitdan-actions#17.
#
# The executable bit is the discriminator because it is the linker's own
# output that is at risk, not the directory it happens to land in — `.rlib`,
# `.rmeta` and `incremental/` stay shared and they are the bytes that matter.
"$fn" 'linked outputs' - -perm -u+x || return 1
"$fn" '.rustc_info.json' 3 -name '.rustc_info.json' || return 1
return 0
}
# The directories `unshare_mutable_paths` replaces, under either layout.
#
# All four names are pruned, so nothing selected here can contain anything else
# selected here and the caller never unshares a subtree twice.
#
# `out` is the one that needs deciding rather than naming, and it is the whole
# difficulty of layout v2: a compile unit's rlib and a build script's OUT_DIR
# are both a directory called `out`, one directory apart, and they need
# opposite treatment.
#
# AMBIGUITY RESOLVES TOWARD UNSHARING, and that direction is the rule rather
# than a default: over-unsharing costs bytes, under-unsharing costs corruption.
# So an `out` directory is left shared only when TWO independent signals agree
# it is a compile unit's artifact directory, and either one missing is enough
# to real-copy it:
#
# 1. it holds an `.rlib`/`.rmeta` of its own — the artifact whose sharing is
# the entire point of the clone; and
# 2. its unit directory has no record of a build-script execution beside it
# (`run/` under layout v2, a loose `root-output` under v1).
#
# Signal 2 alone was the first cut of this and it is NOT sufficient, because
# Cargo writes `root-output` only AFTER the script exits successfully. A build
# script that populates `OUT_DIR` and then FAILS leaves a unit with no record
# at all, which reads as "compile unit" — and the old `-name build` selection
# covered that state by real-copying `build/` wholesale, so trusting signal 2
# alone was a regression against it. Reproduced on cargo 1.93.1 stable: the
# clone's build wrote through the shared inode into the source's OUT_DIR.
# Signal 1 closes it, because a failed build script's OUT_DIR holds no rlib.
#
# The residual is a build script that writes a file NAMED `*.rlib`/`*.rmeta`
# into `OUT_DIR` and has never once succeeded. Nothing bounds that away; it is
# simply far narrower than what it replaces.
#
# Requiring signal 1 also means a bin, test or build-script COMPILE unit's
# `out` is real-copied rather than shared — at no cost in bytes, since
# everything in one is an executable or a `*.d`, and both are privately owned
# by the file rules below either way.
_mutable_dirs() {
local root="$1" d unit
while IFS= read -r d; do
if [ "${d##*/}" = out ]; then
unit="${d%/out}"
if ! [ -d "$unit/run" ] && ! [ -e "$unit/root-output" ] \
&& _holds_compiled_artifact "$d"; then
continue
fi
fi
printf '%s\n' "$d"
done < <(find "$root" -type d \
\( "${_MUTABLE_DIR_NAMES[@]}" \) \
-prune -print 2>/dev/null)
}
# _mutable_file_rules' callback for the copying side. The root travels in a
# variable rather than an argument because the callback's own signature is the
# rule's, and every rule has to reach the same tree.
_unshare_one_rule() {
local label="$1" maxdepth="$2"; shift 2
local -a depth=()
[ "$maxdepth" = - ] || depth=(-maxdepth "$maxdepth")
_unshare_files "$_MUTABLE_ROOT" ${depth[@]+"${depth[@]}"} -type f "$@" || {
echo "::error::unshare_mutable_paths: failed to unshare ${label} under ${_MUTABLE_ROOT}" >&2
return 1
}
return 0
}
# THE load-bearing function of this whole design.
#
# A hardlink clone is only safe if every write the clone's build performs
# lands on a NEW inode, leaving the source's data untouched. That is true for
# compilation artifacts — rustc and the linker replace `deps/*.rlib`,
# `*.rmeta`, and binaries rather than truncating them in place — and it is
# NOT true for the metadata Cargo and build scripts write with a plain
# truncating write. Measured directly (Linux, ext4, cargo 1.9x nightly:
# `cp -al` a warm target dir, change a source file, build in the clone, diff
# the source) the following files in the SOURCE were mutated through the
# shared inode:
# lands on a NEW inode, leaving the source's data untouched. That is true of
# rustc's own outputs — it writes an `.rlib` or `.rmeta` to a temporary and
# renames it into place — and it is NOT true of the metadata Cargo and build
# scripts write with a plain truncating write, nor of anything the LINKER
# produces. Measured directly (Linux, ext4: `cp -al` a warm target dir, change
# a source file, build in the clone, diff the source) the following files in
# the SOURCE were mutated through the shared inode:
#
# <profile>/.fingerprint/<unit>/dep-<target> (only under
# CARGO_UNSTABLE_CHECKSUM_FRESHNESS,
# where this file carries the
# per-source blake3 checksums.
# NOT reproduced on
# 1.100.0-nightly (2026-08-25),
# measured by this repo's own CI
# — upstream appears to have
# stopped writing it in place.
# Kept in the unshared set
# anyway: it costs 22 MB of a
# 6.9 GB tree, and the failure
# it guards is a wrong answer,
# not a slow one.)
# <profile>/build/<pkg>/output, root-output (Cargo build-script metadata)
# <profile>/.fingerprint/<unit>/dep-<target> (build-dir layout v1) — or,
# <profile>/build/<pkg>/<hash>/fingerprint/dep-<target>
# (build-dir layout v2; see the
# dated note below for which
# Cargo writes which). Only
# when Cargo resolves freshness
# by CONTENT, where this file
# carries the per-source blake3
# checksums.
# <profile>/build/<pkg>/output, root-output (Cargo build-script metadata;
# `<pkg>/<hash>/run/root-output`
# under layout v2)
# <profile>/build/<pkg>/out/** (whatever the build script
# writes into OUT_DIR — build
# scripts overwhelmingly use a
# plain fs::write)
# <profile>/deps/*.d, <profile>/*.d (Cargo's post-processed
# dep-info)
# <profile>/deps/<test>-<hash> (a linked TEST binary; under
# layout v2,
# `build/<pkg>/<hash>/out/`)
#
# THE LINKED-OUTPUT CASE IS NOT LAYOUT-SPECIFIC AND WAS NOT PART OF THIS
# FUNCTION UNTIL 2026-08-27 (gitdan-actions#14). A `cargo test --no-run` inside
# a raw `cp -al` clone rewrote the source's own test binary in place on cargo
# 1.93.1 stable, 1.96.0-nightly, 1.98.0-nightly (layout v1) and 1.100.0-nightly
# (layout v2) alike. Cargo re-creates the path first when it also has to uplift
# the result — a bin target's `deps/<bin>-<hash>` has a hardlink twin at
# `<profile>/<bin>` — and other crate shapes relinked to a fresh inode for
# reasons this measurement did not pin down. Since the safe cases could not be
# enumerated, every executable is treated as mutable; `.rlib`, `.rmeta` and
# `incremental/` are what stay shared, and they are the bytes worth sharing.
#
# The checksum-freshness case is not a cosmetic one. Reproduced end to end:
# branch B clones base's cache, builds its own content, and thereby rewrites
@@ -359,10 +506,54 @@ _unshare_files() {
# sources, reports `Fresh`, and reuses a binary built from the PRE-merge code.
# That is silent stale-artifact reuse — a wrong answer, not a slow one.
#
# So: hardlink the artifacts (the GB), real-copy the metadata (the MB).
# Measured on a 6.9 GB Bevy workspace target dir, the unshared set is
# .fingerprint 22 MB + build/ 237 MB + a handful of dep-info files — about
# 3.7% of the tree, against 100% for a plain `cp -a`.
# WHAT UPSTREAM CHANGED, AND WHAT IT DID NOT (measured 2026-08-26, daniel/gitdan#62).
#
# An earlier revision of this comment recorded that the dep-info write was
# "NOT reproduced on 1.100.0-nightly (2026-08-25) — upstream appears to have
# stopped writing it in place". That reading was wrong, and the way it was
# wrong is the reason this paragraph is dated. Two unrelated upstream changes
# landed within days of each other, and between them they moved both the
# switch that turns the behaviour on and the path it writes to:
#
# 1. The ON-SWITCH MOVED. cargo PR #17382 `feat(config): Add build.fingerprint`
# (merged 2026-08-22) demoted `-Z checksum-freshness` to a gate: it now
# only UNLOCKS the feature, and `build.fingerprint` SELECTS it, defaulting
# to `"mtime"`. So `CARGO_UNSTABLE_CHECKSUM_FRESHNESS=true` on its own is
# accepted and does nothing, which is exactly the "flag accepted, mtime
# anyway" result that was mistaken for a withdrawal. Content freshness
# needs BOTH, and with both it is entirely intact:
#
# CARGO_UNSTABLE_CHECKSUM_FRESHNESS=true CARGO_BUILD_FINGERPRINT=content
#
# Measured on cargo 1.100.0-nightly (e8cb624d5 2026-08-22): with the gate
# alone a `cp -al` clone mutates only the build/ and *.d families; add
# `CARGO_BUILD_FINGERPRINT=content` and the source's dep-info file is
# mutated through the shared inode again. Same toolchain, same clone, one
# env var apart. The hazard was never removed — it was switched off.
#
# 2. THE PATH MOVED. Build-dir layout v2 (`-Z build-dir-new-layout`, cargo
# 1.91) became the nightly default in cargo 1.99 (PR #17258) and was
# stabilised by PR #17354, merged 2026-08-18, shipping in cargo 1.100.0
# stable on 2026-11-12. Under v2 there is no `<profile>/.fingerprint` and
# no `<profile>/deps` at all: everything is regrouped per build unit under
# `<profile>/build/<pkg>/<hash>/{fingerprint,out,run}/`, artifacts
# included. Bracketed locally: cargo 1.97.1 and 1.98.0-nightly write v1,
# 1.100.0-nightly writes v2.
#
# Until 2026-08-27 the selection was `-name .fingerprint -o -name build`,
# which under a v2 Cargo matched nothing on its first clause and the entire
# tree on its second, because the artifacts moved under `build/` too. The guard
# held by accident and the saving did not: on one scratch crate (serde +
# serde_json + regex plus a build script), same sources both ways —
#
# cargo 1.98.0-nightly (layout v1) 64.9 MB unshared of 165.1 MB — 39.3%
# 1.100.0-nightly (layout v2) 110.4 MB unshared of 110.4 MB — 99.996%
#
# The selection now names the mutable set directly rather than by the container
# it used to live in, so it holds under both layouts; the same crate measures
# 45.3% (v1) and 38.4% (v2), both dominated by the linked-output rule above
# rather than by the layout. On a real 5.5 GB Bevy target dir the whole change
# moves the real-copied share from 9.0% to 14.1%.
#
# `incremental/` is deliberately left shared: rustc writes each incremental
# session to a fresh `s-*-working` directory and finalises it with a rename,
@@ -372,12 +563,12 @@ _unshare_files() {
unshare_mutable_paths() {
local root="$1" d
[ -d "$root" ] || return 0
_MUTABLE_ROOT="$root"
# The list is materialised in full before anything is replaced: each
# replacement deletes and recreates a directory, and a live `find` walk over
# a tree being mutated underneath it is a needless hazard. `-prune` keeps a
# match's own contents out of the list.
# a tree being mutated underneath it is a needless hazard.
local -a dirs=()
mapfile -t dirs < <(find "$root" -type d \( -name .fingerprint -o -name build \) -prune -print 2>/dev/null)
mapfile -t dirs < <(_mutable_dirs "$root")
for d in "${dirs[@]}"; do
[ -n "$d" ] || continue
unshare_subtree "$d" || {
@@ -385,14 +576,121 @@ unshare_mutable_paths() {
return 1
}
done
_unshare_files "$root" -type f -name '*.d' || {
echo "::error::unshare_mutable_paths: failed to unshare dep-info files under ${root}" >&2
return 1
_mutable_file_rules _unshare_one_rule || return 1
return 0
}
_unshare_files "$root" -maxdepth 3 -type f -name '.rustc_info.json' || {
echo "::error::unshare_mutable_paths: failed to unshare .rustc_info.json under ${root}" >&2
return 1
# ---------------------------------------------------------------------------
# What the seed will clone, and what that clone costs in disk
# ---------------------------------------------------------------------------
# The sources seed-target-dir.sh considers, most specific first, as
# `<dir>:<label>` lines.
#
# Read by the seed, which clones the first one that exists, and by the prune
# that has to size the volume for that clone BEFORE it happens. One derivation
# rather than two agreeing ones: a prune that sizes a different tree than the
# seed clones is measuring nothing, and nothing downstream would say so.
seed_source_candidates() {
local root="$1" own_key="$2" base_key="$3" fallback="${4:-}"
[ -n "$base_key" ] && printf '%s:base snapshot\n' "$(snapshot_dir_for "$root" "$base_key")"
printf '%s:own snapshot\n' "$(snapshot_dir_for "$root" "$own_key")"
[ -n "$fallback" ] && printf '%s:fallback dir\n' "$fallback"
return 0
}
# The directory the seed will actually hardlink-clone on this run, or nothing
# at all when it will not clone: its own target dir already exists (the seed
# reuses it and returns before the candidate list is consulted), or no
# candidate exists (it starts cold).
seed_clone_source() {
local root="$1" own_key="$2" base_key="$3" fallback="${4:-}" entry src
[ -d "$(target_dir_for "$root" "$own_key")" ] && return 0
while IFS= read -r entry; do
src="${entry%%:*}"
if [ -d "$src" ]; then printf '%s' "$src"; return 0; fi
done < <(seed_source_candidates "$root" "$own_key" "$base_key" "$fallback")
return 0
}
# _mutable_file_rules' callback for the measuring side.
#
# One find per rule, skipping the subtrees `_mutable_dirs` already selects
# whole — those are measured by the `du` in mutable_set_kb, and counting a
# file twice would inflate the requirement into evicting caches nothing
# needed. `%k` is allocated 1K blocks, the same unit `du -sk` reports, so the
# two halves add.
_size_one_rule() {
local maxdepth="$2"; shift 2
local -a depth=()
[ "$maxdepth" = - ] || depth=(-maxdepth "$maxdepth")
find "$_MUTABLE_ROOT" ${depth[@]+"${depth[@]}"} \
\( "${_MUTABLE_DIR_NAMES[@]}" \) -prune -o \
-type f \( "$@" \) -printf '%k\n' 2>/dev/null
return 0
}
# The kilobytes `unshare_mutable_paths` will really-copy out of <dir> — the
# part of a hardlink clone that costs new disk, as opposed to the `.rlib`,
# `.rmeta` and `incremental/` bytes that stay shared with the source.
#
# Measured off the SAME enumeration the copier uses (`_mutable_dirs` and
# `_mutable_file_rules`), which is the only thing that makes this a
# measurement rather than an estimate.
#
# Two residuals, both named because neither is bounded away:
#
# OVER by any file matching two rules at once — an executable named
# `output`, say. Rare, and small.
# UNDER by the `*.d` files inside an `out` directory that _mutable_dirs
# leaves SHARED (the compile-unit case: it holds an .rlib and has no
# build-script record beside it). Such a directory holds a library artifact
# by definition, so what is missed is dep-info, not executables. Also under
# by a `.rustc_info.json` deeper than the copier's own maxdepth, which is
# kilobytes.
#
# The margin in clone_headroom_kb is what covers the under-count; it is not
# there to make the measurement optional.
mutable_set_kb() {
local root="$1"
[ -d "$root" ] || { printf '0'; return 0; }
_MUTABLE_ROOT="$root"
{
_mutable_dirs "$root" | tr '\n' '\0' | xargs -0 -r du -sk 2>/dev/null | awk '{print $1}'
_mutable_file_rules _size_one_rule
} | awk '{s += $1} END { printf "%d", s + 0 }'
return 0
}
# How much free space the seed needs on the volume before it clones <dir>.
#
# CACHE_CLONE_HEADROOM_PERCENT scales the measured mutable set;
# CACHE_CLONE_HEADROOM_FLOOR_KB is added on top. BOTH DEFAULTS ARE
# HAND-WRITTEN — nothing measures them, and they are separate because they
# cover different things:
#
# The percentage covers what scales with the tree: `unshare_subtree` stages
# each mutable directory through a sibling copy before dropping the shared
# original, so at its peak one subtree is held twice, and the under-count
# named on mutable_set_kb scales with the tree too.
# The floor covers what does not: `cp -al` materialises every DIRECTORY for
# real (only files are linked), and a Bevy-sized target dir has hundreds of
# thousands of them.
#
# Both are overridable, and the direction of error is deliberate. Over-asking
# evicts a cache that would have fitted, costing one branch a cold start;
# under-asking lets the clone start and run out of disk halfway through the
# unshare, which fails the job with an error naming a staging path — the
# failure gitdan-actions#20 is filed about.
CACHE_CLONE_HEADROOM_PERCENT="${CACHE_CLONE_HEADROOM_PERCENT:-150}"
CACHE_CLONE_HEADROOM_FLOOR_KB="${CACHE_CLONE_HEADROOM_FLOOR_KB:-2097152}"
clone_headroom_kb() {
local src="${1:-}" kb
if [ -z "$src" ] || [ ! -d "$src" ]; then printf '0'; return 0; fi
kb=$(mutable_set_kb "$src")
awk -v k="$kb" -v pct="$CACHE_CLONE_HEADROOM_PERCENT" -v floor="$CACHE_CLONE_HEADROOM_FLOOR_KB" \
'BEGIN { printf "%d", (k * pct / 100) + floor }'
return 0
}
+384 -33
View File
@@ -5,8 +5,10 @@
#
# That assumption is FALSE for a plain `cp -al`. Measured, and asserted below
# as an explicit control: build in a raw `cp -al` clone and the source's
# `.fingerprint/<unit>/dep-*` (under CARGO_UNSTABLE_CHECKSUM_FRESHNESS),
# `build/<pkg>/output`, `build/<pkg>/out/**` and `deps/*.d` all change,
# dep-info file (`.fingerprint/<unit>/dep-*` under Cargo's build-dir layout
# v1, `build/<pkg>/<hash>/fingerprint/dep-*` under v2 — and under content
# freshness only, see the probe below), `build/<pkg>/output`,
# `build/<pkg>/out/**` and `deps/*.d` all change,
# because Cargo and build scripts write those with a plain truncating write
# rather than the write-then-rename Cargo uses for real artifacts.
#
@@ -26,8 +28,6 @@ set -euo pipefail
script_dir=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)
. "$script_dir/cache-lib.sh"
command -v cargo >/dev/null || { echo "SKIP: no cargo on PATH"; exit 0; }
scratch=$(mktemp -d)
trap 'rm -rf "$scratch"' EXIT
pass_count=0
@@ -57,6 +57,14 @@ version = "0.1.0"
edition = "2021"
[workspace]
TOML
# A binary as well as a library, because the two are written differently and
# only one of them is safe to share: rustc writes an rlib to a temporary and
# renames it into place, while the LINKER writes an executable through the
# existing inode. Without a bin target this suite never relinks anything and
# cannot see that difference.
cat > "$dir/src/main.rs" <<'RS'
fn main() { println!("{}", probe::f()); }
RS
cat > "$dir/build.rs" <<'RS'
use std::{env, fs, path::PathBuf};
fn main() {
@@ -68,6 +76,219 @@ fn main() {
RS
}
# ---------------------------------------------------------------------------
# Both build-dir layouts, without a compiler
# ---------------------------------------------------------------------------
#
# Every other scenario in this file runs whichever layout the installed Cargo
# happens to write, so on any one machine it exercises exactly ONE of the two —
# and gitdan-ci's is v1. The two fixtures below reproduce both directory shapes
# from files alone, clone them through the real `hardlink_clone_into()`, and
# assert file by file which side of the partition each one lands on.
#
# Shapes taken from a scratch crate (serde + serde_json + regex, plus a build
# script) built on 2026-08-27: cargo 1.98.0-nightly (a335d47ff 2026-06-26)
# writes v1, cargo 1.100.0-nightly (e8cb624d5 2026-08-22) writes v2.
#
# v2 is where the partition is easy to get wrong, and the fixtures are built to
# say so: a build script's OUT_DIR and a compile unit's rlib are BOTH a
# directory called `out`, one directory apart, and they need opposite
# treatment.
mkfile() { mkdir -p "$(dirname "$1")"; printf '%s' "$2" > "$1"; }
# Big enough that the byte-fraction assertion below measures something.
artifact_bytes=$(head -c 4096 /dev/zero | tr '\0' 'A')
# Spec lines are `<shared|private>|<path relative to the tree root>`.
assert_partition() {
local label="$1" src="$2" spec="$3"
local clone="${src}-clone" want rel si ci total=0 copied=0 sz
hardlink_clone_into "$src" "$clone" "selftest-$label" \
|| fail "$label: hardlink_clone_into refused the destination"
while IFS='|' read -r want rel; do
[ -n "${rel:-}" ] || continue
[ -e "$clone/$rel" ] || fail "$label: $rel is missing from the clone"
si=$(stat -c '%i' "$src/$rel"); ci=$(stat -c '%i' "$clone/$rel")
case "$want" in
shared)
[ "$si" = "$ci" ] \
|| fail "$label: $rel was real-copied, but it is an artifact and must stay shared" ;;
private)
[ "$si" != "$ci" ] \
|| fail "$label: $rel still shares an inode with the source, so a build in the clone can rewrite it" ;;
*) fail "$label: unknown spec verb '$want'" ;;
esac
done <<< "$spec"
ok "$label: every file landed on the right side of the partition"
# The cost model, asserted rather than assumed. Selecting too much is not a
# correctness bug, which is exactly why nothing caught layout v2 taking the
# selection from 39.3% of one scratch crate's tree to 99.996% of it
# (gitdan-actions#14):
# a hardlink clone that real-copies everything is a `cp -a` with extra steps.
# The bound is loose on purpose. It is not a budget — the honest figure moves
# with how much of a tree is linker output, and these fixtures are mostly
# that by construction — it is a floor under "still a hardlink clone at all".
while IFS= read -r rel; do
# `.cargo-*lock*` is stripped from every clone by design, so it has no
# counterpart to compare against.
[ -e "$clone/$rel" ] || continue
sz=$(stat -c '%s' "$src/$rel")
total=$((total + sz))
[ "$(stat -c '%i' "$src/$rel")" = "$(stat -c '%i' "$clone/$rel")" ] || copied=$((copied + sz))
done < <(cd "$src" && find . -type f -printf '%P\n')
[ "$total" -gt 0 ] || fail "$label: fixture has no bytes to measure"
[ $((copied * 100 / total)) -lt 90 ] \
|| fail "$label: the clone real-copied $((copied * 100 / total))% of its source's bytes — the hardlink saving is gone"
ok "$label: clone real-copies $((copied * 100 / total))% of ${total} B (${copied} B), the rest is shared"
}
# Layout v2: no `.fingerprint`, no `deps`. Everything regroups per build unit
# under `build/<pkg>/<hash>/{fingerprint,out,run}`, artifacts included — which
# is what took `-name build` from "the metadata" to "the whole tree".
#
# THE THREE UNIT KINDS ARE THE POINT. `1bf...` is a compile unit: its `out`
# holds the rlib. `08c...` is the build script's own compile unit: its `out`
# holds the build-script binary. `da9...` is the build-script RUN unit: its
# `out` IS the OUT_DIR, and it is the only one of the three whose `out` is
# mutable. The `run/` directory beside it is the structural difference.
v2="$scratch/layout-v2"
mkfile "$v2/CACHEDIR.TAG" 'Signature: 8a477f597d28d172'
mkfile "$v2/.rustc_info.json" '{"rustc_fingerprint":1}'
mkfile "$v2/debug/.cargo-lock" ''
mkfile "$v2/debug/libprobe.rlib" "$artifact_bytes"
mkfile "$v2/debug/libprobe.d" '/probe/src/lib.rs:'
for f in dep-lib-probe lib-probe lib-probe.json invoked.timestamp; do
mkfile "$v2/debug/build/probe/1bf5493368dce3cd/fingerprint/$f" "$f"
done
mkfile "$v2/debug/build/probe/1bf5493368dce3cd/out/libprobe-1bf5493368dce3cd.rlib" "$artifact_bytes"
mkfile "$v2/debug/build/probe/1bf5493368dce3cd/out/libprobe-1bf5493368dce3cd.rmeta" "$artifact_bytes"
mkfile "$v2/debug/build/probe/1bf5493368dce3cd/out/probe-1bf5493368dce3cd.d" '/probe/src/lib.rs:'
for f in build-script-build-script-build build-script-build-script-build.json \
dep-build-script-build-script-build invoked.timestamp; do
mkfile "$v2/debug/build/probe/08c7dda6eacd6dca/fingerprint/$f" "$f"
done
mkfile "$v2/debug/build/probe/08c7dda6eacd6dca/out/build_script_build" "$artifact_bytes"
chmod +x "$v2/debug/build/probe/08c7dda6eacd6dca/out/build_script_build"
mkfile "$v2/debug/build/probe/08c7dda6eacd6dca/out/build_script_build.d" '/probe/build.rs:'
# A test binary: the same `out` directory as the rlib above, and the largest
# thing in a real tree that a linker writes.
for f in dep-test-lib-probe test-lib-probe test-lib-probe.json invoked.timestamp; do
mkfile "$v2/debug/build/probe/6a091d813b2be60d/fingerprint/$f" "$f"
done
mkfile "$v2/debug/build/probe/6a091d813b2be60d/out/probe-6a091d813b2be60d" "$artifact_bytes"
chmod +x "$v2/debug/build/probe/6a091d813b2be60d/out/probe-6a091d813b2be60d"
mkfile "$v2/debug/build/probe/6a091d813b2be60d/out/probe-6a091d813b2be60d.d" '/probe/src/lib.rs:'
for f in run-build-script-build-script-build run-build-script-build-script-build.json; do
mkfile "$v2/debug/build/probe/da96cf45111f80dd/fingerprint/$f" "$f"
done
mkfile "$v2/debug/build/probe/da96cf45111f80dd/out/gen.txt" 'generated from 24 bytes'
for f in invoked.timestamp root-output stdout stderr; do
mkfile "$v2/debug/build/probe/da96cf45111f80dd/run/$f" "$f"
done
# A build script that wrote into OUT_DIR and then FAILED: Cargo records the
# run only on success, so this unit has `out/` populated and no `run/` at all.
# Reading "no execution record" as "compile unit" left this shared, which is
# the one state the old `-name build` selection covered and the first cut of
# this one did not.
mkfile "$v2/debug/build/probe/f00ded00f00ded00/out/gen.txt" 'half-written'
mkfile "$v2/debug/incremental/probe-abc/s-xyz/dep-graph.bin" "$artifact_bytes"
assert_partition "layout v2" "$v2" "$(cat <<'SPEC'
private|.rustc_info.json
private|debug/libprobe.d
private|debug/build/probe/1bf5493368dce3cd/fingerprint/dep-lib-probe
private|debug/build/probe/1bf5493368dce3cd/fingerprint/lib-probe
private|debug/build/probe/1bf5493368dce3cd/fingerprint/lib-probe.json
private|debug/build/probe/1bf5493368dce3cd/fingerprint/invoked.timestamp
private|debug/build/probe/1bf5493368dce3cd/out/probe-1bf5493368dce3cd.d
private|debug/build/probe/08c7dda6eacd6dca/fingerprint/dep-build-script-build-script-build
private|debug/build/probe/08c7dda6eacd6dca/fingerprint/invoked.timestamp
private|debug/build/probe/08c7dda6eacd6dca/out/build_script_build.d
private|debug/build/probe/da96cf45111f80dd/fingerprint/run-build-script-build-script-build
private|debug/build/probe/da96cf45111f80dd/out/gen.txt
private|debug/build/probe/da96cf45111f80dd/run/root-output
private|debug/build/probe/da96cf45111f80dd/run/stdout
private|debug/build/probe/da96cf45111f80dd/run/stderr
private|debug/build/probe/da96cf45111f80dd/run/invoked.timestamp
private|debug/build/probe/f00ded00f00ded00/out/gen.txt
shared|debug/libprobe.rlib
shared|debug/build/probe/1bf5493368dce3cd/out/libprobe-1bf5493368dce3cd.rlib
shared|debug/build/probe/1bf5493368dce3cd/out/libprobe-1bf5493368dce3cd.rmeta
private|debug/build/probe/08c7dda6eacd6dca/out/build_script_build
private|debug/build/probe/6a091d813b2be60d/out/probe-6a091d813b2be60d
private|debug/build/probe/6a091d813b2be60d/out/probe-6a091d813b2be60d.d
private|debug/build/probe/6a091d813b2be60d/fingerprint/dep-test-lib-probe
shared|debug/incremental/probe-abc/s-xyz/dep-graph.bin
SPEC
)"
# Layout v1: one `.fingerprint` and one `deps` per profile; `build/<pkg>-<hash>`
# holds the build script's compiled binary in one unit directory and its run
# metadata plus OUT_DIR in another.
v1="$scratch/layout-v1"
mkfile "$v1/CACHEDIR.TAG" 'Signature: 8a477f597d28d172'
mkfile "$v1/.rustc_info.json" '{"rustc_fingerprint":1}'
mkfile "$v1/debug/.cargo-lock" ''
mkfile "$v1/debug/libprobe.rlib" "$artifact_bytes"
mkfile "$v1/debug/libprobe.d" '/probe/src/lib.rs:'
mkfile "$v1/debug/deps/libprobe-1bf5493368dce3cd.rlib" "$artifact_bytes"
mkfile "$v1/debug/deps/libprobe-1bf5493368dce3cd.rmeta" "$artifact_bytes"
mkfile "$v1/debug/deps/probe-1bf5493368dce3cd.d" '/probe/src/lib.rs:'
# A test binary, which layout v1 leaves in `deps/` beside the rlibs.
mkfile "$v1/debug/deps/probe-6a091d813b2be60d" "$artifact_bytes"
chmod +x "$v1/debug/deps/probe-6a091d813b2be60d"
mkfile "$v1/debug/deps/probe-6a091d813b2be60d.d" '/probe/src/lib.rs:'
for f in dep-lib-probe lib-probe lib-probe.json invoked.timestamp; do
mkfile "$v1/debug/.fingerprint/probe-1bf5493368dce3cd/$f" "$f"
done
mkfile "$v1/debug/build/probe-08c7dda6eacd6dca/build-script-build" "$artifact_bytes"
mkfile "$v1/debug/build/probe-08c7dda6eacd6dca/build_script_build-08c7dda6eacd6dca" "$artifact_bytes"
chmod +x "$v1/debug/build/probe-08c7dda6eacd6dca/build-script-build" \
"$v1/debug/build/probe-08c7dda6eacd6dca/build_script_build-08c7dda6eacd6dca"
mkfile "$v1/debug/build/probe-08c7dda6eacd6dca/build_script_build-08c7dda6eacd6dca.d" '/probe/build.rs:'
for f in invoked.timestamp output root-output stderr; do
mkfile "$v1/debug/build/probe-da96cf45111f80dd/$f" "$f"
done
mkfile "$v1/debug/build/probe-da96cf45111f80dd/out/gen.txt" 'generated from 24 bytes'
# The same never-succeeded build script under layout v1.
mkfile "$v1/debug/build/probe-f00ded00f00ded00/out/gen.txt" 'half-written'
mkfile "$v1/debug/incremental/probe-abc/s-xyz/dep-graph.bin" "$artifact_bytes"
assert_partition "layout v1" "$v1" "$(cat <<'SPEC'
private|.rustc_info.json
private|debug/libprobe.d
private|debug/deps/probe-1bf5493368dce3cd.d
private|debug/.fingerprint/probe-1bf5493368dce3cd/dep-lib-probe
private|debug/.fingerprint/probe-1bf5493368dce3cd/lib-probe
private|debug/.fingerprint/probe-1bf5493368dce3cd/lib-probe.json
private|debug/.fingerprint/probe-1bf5493368dce3cd/invoked.timestamp
private|debug/build/probe-08c7dda6eacd6dca/build_script_build-08c7dda6eacd6dca.d
private|debug/deps/probe-6a091d813b2be60d
private|debug/deps/probe-6a091d813b2be60d.d
private|debug/build/probe-da96cf45111f80dd/invoked.timestamp
private|debug/build/probe-da96cf45111f80dd/output
private|debug/build/probe-da96cf45111f80dd/root-output
private|debug/build/probe-da96cf45111f80dd/stderr
private|debug/build/probe-da96cf45111f80dd/out/gen.txt
private|debug/build/probe-f00ded00f00ded00/out/gen.txt
shared|debug/libprobe.rlib
shared|debug/deps/libprobe-1bf5493368dce3cd.rlib
shared|debug/deps/libprobe-1bf5493368dce3cd.rmeta
private|debug/build/probe-08c7dda6eacd6dca/build-script-build
private|debug/build/probe-08c7dda6eacd6dca/build_script_build-08c7dda6eacd6dca
shared|debug/incremental/probe-abc/s-xyz/dep-graph.bin
SPEC
)"
echo
command -v cargo >/dev/null || {
echo "SKIP: no cargo on PATH — the fixture scenarios above ran, the live-Cargo ones cannot"
echo "hardlink-clone-selftest: ${pass_count} assertions passed"
exit 0
}
crate_dir="$scratch/probe"
mkcrate "$crate_dir"
cd "$crate_dir"
@@ -97,10 +318,19 @@ CONTENT_B='pub fn f() -> u32 { 22222 } pub fn g() -> u32 { 7 }'
# short-circuits before `-Z` is
# even parsed.
# `cargo +nightly -Z checksum-freshness — answers "is this flag still
# locate-project` accepted". 1.100.0-nightly
# (2026-08-25) accepts it and does
# not resolve freshness by content
# anyway.
# locate-project` accepted", which since cargo PR
# #17382 (2026-08-22) is a
# different question from "is
# content freshness on". That PR
# demoted the flag to a gate and
# gave `build.fingerprint` the
# choice, defaulting to `mtime` —
# so 1.100.0-nightly accepts the
# flag and resolves freshness by
# mtime unless
# CARGO_BUILD_FINGERPRINT=content
# is set too. Measured 2026-08-26;
# see daniel/gitdan#62.
#
# The scenario at the end of this file depends on one thing and it is neither
# of those: that changed content with an OLDER mtime rebuilds. Under mtime
@@ -165,7 +395,9 @@ CONTENT_B='pub fn f() -> u32 { 22222 } pub fn g() -> u32 { 7 }'
# broken" across the three repos consuming this action — the same category
# error the three-state split exists to prevent, one level up. What would
# change the answer is not-measured becoming the everyday CI outcome; it is not
# (gitdan-ci reports a measured INACTIVE, by measurement).
# — gitdan-ci's outcome is a measurement either way. It reported a measured
# INACTIVE until 2026-08-26, for the reason recorded at the `export` below, and
# an ACTIVE once both switches were set.
CHECKSUM_MODE="off"
CHECKSUM_REASON="no nightly on PATH accepting -Z checksum-freshness"
CARGO_BIN=(cargo)
@@ -189,6 +421,14 @@ checksum_freshness_probe() {
}
if cargo +nightly -Z checksum-freshness locate-project > /dev/null 2>&1; then
export CARGO_UNSTABLE_CHECKSUM_FRESHNESS=true
# BOTH, since cargo PR #17382 (2026-08-22): the -Z flag only unlocks the
# feature and `build.fingerprint` selects it, defaulting to `mtime`. Setting
# the gate alone is what made this suite report a measured INACTIVE on
# 1.100.0-nightly and skip its strongest scenario (daniel/gitdan#62). Safe to
# export unconditionally — a Cargo that does not know the key ignores it
# silently, verified 2026-08-26 on 1.93.1 stable and 1.96.0-nightly, both of
# which still measure ACTIVE from the gate alone.
export CARGO_BUILD_FINGERPRINT=content
# Errexit off across the call, so the subshell can arm its own — see the
# header. `probe_rc` is read before it is restored.
probe_rc=0
@@ -203,11 +443,11 @@ if cargo +nightly -Z checksum-freshness locate-project > /dev/null 2>&1; then
CHECKSUM_REASON=""
;;
3)
unset CARGO_UNSTABLE_CHECKSUM_FRESHNESS
CHECKSUM_REASON="this nightly accepts -Z checksum-freshness but resolves freshness by mtime"
unset CARGO_UNSTABLE_CHECKSUM_FRESHNESS CARGO_BUILD_FINGERPRINT
CHECKSUM_REASON="this nightly accepts -Z checksum-freshness and build.fingerprint=content but still resolves freshness by mtime"
;;
*)
unset CARGO_UNSTABLE_CHECKSUM_FRESHNESS
unset CARGO_UNSTABLE_CHECKSUM_FRESHNESS CARGO_BUILD_FINGERPRINT
CHECKSUM_MODE="unmeasured"
CHECKSUM_REASON="the probe exited ${probe_rc}, which is not one of its answer codes, so this was NOT MEASURED — this toolchain may or may not resolve freshness by content"
# Loud, because the cost is silently lost coverage on a machine that
@@ -248,22 +488,31 @@ printf '%s\n' "$ctl_mutated" | sed 's/^/ /'
# is to prove the hazard exists at all, which the non-empty set above already
# does; this line records WHICH families a given Cargo exhibits.
#
# `.fingerprint/*/dep-*` is the worst of them — it carries the per-source
# The dep-info file is the worst of them — it carries the per-source
# checksums, so mutating it through a shared inode turns a hardlink clone into
# silent stale-artifact reuse rather than a slow build. It was measured on
# cargo 1.9x nightly (see unshare_mutable_paths in cache-lib.sh) and is NOT
# reproduced on 1.100.0-nightly (2026-08-25), where the control mutates only
# the build/ and *.d families. Failing on its absence would mean this suite
# goes red whenever upstream stops doing something we never wanted it to do —
# and it would go red in the CONTROL, where a failure reads as "the hazard is
# gone" rather than "upstream changed". Nothing is lost by reporting it: the
# fix scenario below asserts the source is byte-identical after a full rebuild
# in the clone, which covers every family this Cargo has, named or not.
# silent stale-artifact reuse rather than a slow build. Failing on its absence
# would mean this suite goes red whenever upstream stops doing something we
# never wanted it to do — and it would go red in the CONTROL, where a failure
# reads as "the hazard is gone" rather than "upstream changed". Nothing is lost
# by reporting it: the fix scenario below asserts the source is byte-identical
# after a full rebuild in the clone, which covers every family this Cargo has,
# named or not.
#
# THE PATTERN MUST MATCH BOTH LAYOUTS, and that is not a detail. Cargo's
# build-dir layout v2 moved the file from `<profile>/.fingerprint/<unit>/dep-*`
# to `<profile>/build/<pkg>/<hash>/fingerprint/dep-*` (stabilised by cargo PR
# #17354, cargo 1.100.0, stable 2026-11-12; nightly default since 1.99). An
# earlier cut of this line looked for the v1 path only, so on 2026-08-26,
# against 1.100.0-nightly with content freshness genuinely on, it printed
# "does NOT rewrite ... in place" directly beneath a control listing that
# showed the rewrite. A reporting line that can contradict the data three
# lines above it is worse than no line at all. `fingerprint/.*dep-` matches
# either layout and neither `.d` family.
if [ "$CHECKSUM_MODE" = "on" ]; then
if printf '%s' "$ctl_mutated" | grep -q '\.fingerprint/.*/dep-'; then
echo " note: this cargo DOES rewrite .fingerprint/*/dep-* in place under checksum freshness"
if printf '%s' "$ctl_mutated" | grep -q 'fingerprint/.*dep-'; then
echo " note: this cargo DOES rewrite its dep-info fingerprint file in place under content freshness"
else
echo " note: this cargo does NOT rewrite .fingerprint/*/dep-* in place; only the build/ and *.d families appear above"
echo " note: this cargo does NOT rewrite its dep-info fingerprint file in place; only the build/ and *.d families appear above"
fi
fi
@@ -274,27 +523,44 @@ build_base "$base_fix"
before=$(snapshot_tree "$base_fix")
hardlink_clone_into "$base_fix" "$clone_fix" "selftest" || fail "hardlink_clone_into reported the destination already existed"
# The clone's contract, asserted before anything builds in it: artifacts
# share inodes (that is what makes the clone near-free), and every file Cargo
# rewrites in place does not (that is what makes it sound). Checking after a
# rebuild would prove nothing — the rebuild replaces those files anyway.
shared=0; unshared=0
# The clone's contract, asserted before anything builds in it and asserted in
# BOTH directions: every file Cargo rewrites in place is privately owned (that
# is what makes the clone sound), and every artifact still shares its inode
# (that is what makes it near-free). Checking after a rebuild would prove
# nothing — the rebuild replaces those files anyway.
#
# The mutable-family patterns cover both layouts: `.fingerprint/` is v1's,
# `fingerprint/` and `run/` are v2's, and `gen.txt` is this crate's build
# script's OUT_DIR product, which under v2 sits in a directory called `out`
# beside sibling units whose `out` holds artifacts.
shared=0; unshared=0; shared_bytes=0; copied_bytes=0
while IFS= read -r f; do
rel="${f#"$base_fix"/}"
[ -e "$clone_fix/$rel" ] || continue
sz=$(stat -c '%s' "$f")
if [ "$(stat -c '%i' "$f")" = "$(stat -c '%i' "$clone_fix/$rel")" ]; then
case "$rel" in
*/.fingerprint/*|*/build/*|*.d|.rustc_info.json)
*/.fingerprint/*|*/fingerprint/*|*/run/*|*/out/gen.txt|*/output|*/root-output|*/stderr|*/invoked.timestamp|*.d|.rustc_info.json)
fail "mutable path still shares an inode with the source: $rel" ;;
esac
shared=$((shared + 1))
shared=$((shared + 1)); shared_bytes=$((shared_bytes + sz))
else
unshared=$((unshared + 1))
case "$rel" in
*.rlib|*.rmeta)
fail "artifact was real-copied rather than shared: $rel" ;;
esac
unshared=$((unshared + 1)); copied_bytes=$((copied_bytes + sz))
fi
done < <(find "$base_fix" -type f)
[ "$shared" -gt 0 ] || fail "nothing is shared — the clone degenerated into a full copy"
[ "$unshared" -gt 0 ] || fail "nothing was unshared — unshare_mutable_paths did not run"
total_bytes=$((shared_bytes + copied_bytes))
ok "fresh clone shares ${shared} artifact files and privately owns ${unshared} mutable ones"
# Reported, not asserted. This crate has no dependencies, so almost all of its
# bytes are the two executables — a ratio that says nothing about a real tree.
# The fixtures above are where the cost model is gated, because there the
# composition is fixed.
echo " note: this clone real-copies $((copied_bytes * 100 / total_bytes))% of ${total_bytes} B"
printf '%s\n' "$CONTENT_B" > src/lib.rs
CARGO_TARGET_DIR="$clone_fix" "${CARGO_BIN[@]}" build -q
@@ -307,6 +573,91 @@ if [ -n "$fix_mutated" ]; then
fi
ok "no file in the source changed after a full rebuild in the clone"
echo
echo "=== a linked TEST binary, which nothing uplifts and nothing replaces ==="
# The one artifact family that is NOT safe to share, and the reason
# `unshare_mutable_paths` privately owns every executable. rustc writes an
# rlib to a temporary and renames it in; the LINKER writes an executable
# through whatever inode is already at the path.
#
# THE CRATE SHAPE IS LOAD-BEARING AND WAS WRONG ONCE. An earlier cut of this
# scenario reused the lib+bin probe crate above, whose test binaries relink to
# a FRESH inode — a shape gitdan-actions#17 records as measured safe. Both
# halves then passed green against the unfixed selection, on the strength of
# dep-info mutations the previous scenario already covers, and the scenario
# pinned nothing. A bin-only crate with a unit test does exhibit the rewrite,
# on cargo 1.93.1 stable and on 1.98.0-nightly and 1.100.0-nightly, so that is
# what this builds. The shape is chosen by measurement rather than derived:
# what separates a rewritten executable from an intact one is not established,
# so the only crate shape this scenario may rest on is one observed to exhibit
# the rewrite.
mkbincrate() {
local dir="$1" marker="$2"
mkdir -p "$dir/src"
cat > "$dir/Cargo.toml" <<'TOML'
[package]
name = "binprobe"
version = "0.1.0"
edition = "2021"
[workspace]
TOML
cat > "$dir/src/main.rs" <<RS
fn main() { println!("${marker}"); }
#[cfg(test)]
mod t { #[test] fn a() { assert_eq!("${marker}".len() > 0, true); } }
RS
}
bin_dir="$scratch/binprobe"
mkbincrate "$bin_dir" MARKER_AAAA
cd "$bin_dir"
base_exe="$scratch/base-exe"; clone_exe_ctl="$scratch/clone-exe-ctl"; clone_exe="$scratch/clone-exe"
# Only the executables are read here. The families the other scenarios cover
# would satisfy a "something changed" assertion on their own, which is exactly
# how the earlier cut of this passed while pinning nothing.
source_exe_digest() {
(cd "$1" && find . -type f -executable -print0 | sort -z | xargs -0 -r sha1sum) 2>/dev/null
}
CARGO_TARGET_DIR="$base_exe" "${CARGO_BIN[@]}" test --no-run -q > /dev/null 2>&1 \
|| fail "the bin-only probe crate failed to build"
before=$(source_exe_digest "$base_exe")
cp -al "$base_exe" "$clone_exe_ctl"
strip_cargo_locks "$clone_exe_ctl"
mkbincrate "$bin_dir" MARKER_BBBB
CARGO_TARGET_DIR="$clone_exe_ctl" "${CARGO_BIN[@]}" test --no-run -q > /dev/null 2>&1
exe_ctl_mutated=$(mutated_paths "$before" "$(source_exe_digest "$base_exe")")
# THREE OUTCOMES, as the freshness probe above has, and for the same reason: a
# scenario that cannot tell "the fix works" from "the hazard never fired" is
# not a gate. If this Cargo does not rewrite the source's test binary, the
# assertion below would pass for a toolchain reason rather than a code one, so
# it is skipped LOUDLY instead of passing quietly.
if [ -z "$exe_ctl_mutated" ]; then
echo "::warning::hardlink-clone-selftest: this toolchain did not rewrite the source's test binary through a raw cp -al clone, so the linked-output scenario proves nothing here and was SKIPPED. That is a statement about this Cargo, not about unshare_mutable_paths."
else
ok "control: a raw cp -al clone rewrites the source's own linked test binary"
printf '%s\n' "$exe_ctl_mutated" | sed 's/^/ /'
# Rebuild the base from the original marker so it is warm and consistent
# again, then do the same thing through the real clone.
mkbincrate "$bin_dir" MARKER_AAAA
CARGO_TARGET_DIR="$base_exe" "${CARGO_BIN[@]}" test --no-run -q > /dev/null 2>&1
before=$(snapshot_tree "$base_exe")
hardlink_clone_into "$base_exe" "$clone_exe" "selftest-exe" \
|| fail "hardlink_clone_into refused the destination"
mkbincrate "$bin_dir" MARKER_BBBB
CARGO_TARGET_DIR="$clone_exe" "${CARGO_BIN[@]}" test --no-run -q > /dev/null 2>&1
exe_mutated=$(mutated_paths "$before" "$(snapshot_tree "$base_exe")")
if [ -n "$exe_mutated" ]; then
printf '%s\n' "$exe_mutated" | sed 's/^/ /' >&2
fail "a test build in the clone mutated the source through a shared inode"
fi
ok "no file in the source changed after a full test build in the clone"
fi
cd "$crate_dir"
echo
if [ "$CHECKSUM_MODE" = "on" ]; then
echo "=== the whole point: the source's next build is still correct ==="
+277 -9
View File
@@ -11,8 +11,13 @@
# Waiting for pressure to notice means paying for dead caches until then.
# 2. LIVE BRANCH SURVIVES despite being OLDER than the dead one — liveness,
# not age, is what decides pass 1.
# 3. PROTECTED REFS NEVER EVICTED under forced disk pressure, even when
# their caches are the oldest on disk and would rank first for LRU.
# 3. A PROTECTED REF'S SNAPSHOT NEVER EVICTED under forced disk pressure,
# even when it is the oldest on disk and would rank first for LRU — its
# own TARGET dir is an ordinary candidate and goes (gitdan-actions#24).
# 3b. AND A PROTECTED REF'S TARGET DIR SURVIVES PASS 1 ANYWAY: its tip is
# trivially an ancestor of itself, so the merged-branch signal must never
# be allowed to evaluate it, or pass 1 — unconditional, not gated on
# pressure — would delete it every run.
# 4. LOCKED CACHE PROTECTED even when dead, old, and under pressure — and
# the pass's closing summary agrees with the decline it just logged,
# rather than reporting that it found nothing.
@@ -28,7 +33,10 @@
# so evicting it frees almost nothing while costing every future PR its
# warm start.
# 8. SELF-CLEAR REPORTS LOUDLY to the job summary, not just a log warning.
# 9. OWN CACHE NEVER EVICTED by a sibling pass.
# 9. OWN CACHE NEVER EVICTED by a sibling pass, genuinely under pressure —
# against a real, shrinking `df` (gitdan-actions#26). Checks CONTENTS,
# not just existence, so pass 3's self-clear can't mask a missed
# pass-2 guard.
# 10. SCOPED TO THE CACHE ROOT — a decoy outside it (standing in for another
# project's volume) is never touched.
# 11. A LIVE READER MARKER PROTECTS A CACHE the same way a lock file does — a
@@ -46,10 +54,36 @@
# empty a tree its owner may still restore under a live cache name. Its
# fixture is an OLD directory renamed a moment ago — production's shape,
# and what lets it tell the two timestamps apart.
# 15. A MERGED-BUT-UNDELETED BRANCH IS DEAD TOO. A branch the forge did not
# delete at merge stays on `ls-remote` forever, so scenario 1's signal
# never fires for it — which is how three 40 GB caches sat on a full
# volume until somebody removed them by hand (gitdan-actions#20). A
# branch whose tip is an ancestor of a protected branch's tip is pruned
# like a deleted one; an unmerged branch beside it is not.
# 16. AND "CANNOT TELL" IS STILL NOT DEATH, at both granularities: a branch
# whose tip is not in this checkout is kept with a warning naming it,
# and a shallow checkout — where a missing object is the normal case —
# withholds the whole signal rather than reading it as "nothing merged".
# The deleted-branch signal keeps working in both.
# 17. THE FREE-SPACE REQUIREMENT IS MEASURED OFF THE SOURCE, not taken as a
# percentage of the volume: the pass evicts until the clone the seed is
# about to make fits, and stops there rather than draining the volume.
# When it cannot get there it FAILS, naming the shortfall and every
# directory it kept instead — because the seed would otherwise fail
# seconds later against a staging path that names nothing.
# 18. AND ON THE LAYOUT THAT PRODUCED THE BUG: three equal-sized caches, one
# of them a merged-but-undeleted branch's, with disk to spare. Exactly
# that one goes. Equal sizes and no pressure are the point — nothing but
# the merge state can be what decides.
set -euo pipefail
script_dir=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)
. "$script_dir/cache-lib.sh"
prune="$script_dir/prune-cache.sh"
# Two scenarios below put a stub of a real tool on PATH for one command.
# Captured once, here, rather than read back at each of those sites: a `$PATH`
# read after the first of them is indistinguishable, to a static check, from
# reading the modification the subshell lost.
outer_path="$PATH"
scratch=$(mktemp -d)
trap 'rm -rf "$scratch"' EXIT
@@ -113,17 +147,84 @@ assert_kept "$root/target-$LIVE" "live branch survives despite an older marker t
assert_log "no matching branch on origin" "eviction reason reported"
echo
echo "=== 3: protected refs never evicted under forced pressure ==="
echo "=== 3: protected refs' SNAPSHOTS never evicted under forced pressure ==="
reset_cache
run_prune "1000000 1000" # 0.1% free
assert_kept "$root/target-$DEV" "dev's target dir survives disk pressure"
assert_kept "$root/snapshot-$DEV" "dev's snapshot survives disk pressure"
assert_kept "$root/target-$MAIN" "main's target dir survives disk pressure"
assert_kept "$root/snapshot-$MAIN" "main's snapshot survives disk pressure"
# Their TARGET dirs are ordinary candidates and go under the same pressure —
# gitdan-actions#24: protecting them starved every second branch of room to
# seed. Asserted here (gone, not kept) so this scenario still red-proves the
# snapshot half if a future change reintroduces target protection.
assert_gone "$root/target-$DEV" "dev's target dir is an ordinary pressure-pass candidate"
assert_gone "$root/target-$MAIN" "main's target dir is an ordinary pressure-pass candidate"
echo
echo "=== 9: own cache never evicted by a sibling pass ==="
assert_kept "$root/target-$OWN" "this run's own cache survives"
echo "=== 3b: a protected ref's target dir is NOT pruned by pass 1's liveness ==="
# The regression this fix could introduce and the selftest above cannot see:
# a protected branch's tip is trivially an ancestor of itself, so once its
# target dir stopped being excluded from pass 1 altogether, is_merged_dead
# read it as "merged into itself" and pass 1 — unconditional, not gated on
# pressure — deleted it on every single run. Plenty of free space, so only
# pass 1 can be responsible for anything gone here.
reset_cache
run_prune "1000000 900000" # 90% free: no pressure at all
assert_kept "$root/target-$DEV" "dev's target dir survives pass 1 despite being its own ancestor"
assert_kept "$root/target-$MAIN" "and so does main's"
if grep -q "merged into" "$scratch/log"; then
fail "a protected ref's own target dir was evaluated by the merged-branch signal at all"
fi
ok "no protected ref's own target dir reaches the merged-branch check"
echo
echo "=== 9: own cache survives a sibling pass genuinely under pressure ==="
# A real, shrinking `df` (the scenario-17 pattern), not CACHE_DF_OVERRIDE:
# eviction has to actually free space for "pressure eases once enough is
# freed" to mean anything. MIN_FREE_PCT=0 and a clone-headroom floor (not
# the percentage floor) drive the requirement, so the requirement is an
# exact, chosen KB rather than a percentage of a volume size this fixture
# would otherwise have to reverse-engineer.
rm -rf "$root"; mkdir -p "$root"
blob_kb=4096
mkdir -p "$root/target-$OWN"
head -c $((blob_kb * 1024)) /dev/zero > "$root/target-$OWN/blob"
touch -d '2020-01-01' "$root/target-$OWN/.cache-last-used"
mkdir -p "$root/target-$LIVE"
head -c $((blob_kb * 1024)) /dev/zero > "$root/target-$LIVE/blob"
touch -d '2021-01-01' "$root/target-$LIVE/.cache-last-used"
# The clone-headroom lookup's base-snapshot candidate — never read for its
# content (CACHE_CLONE_HEADROOM_PERCENT=0 below), only for existing so the
# floor alone becomes the requirement.
mkdir -p "$root/snapshot-$DEV"
real_du=$(command -v du)
used9=$($real_du -sk "$root" | awk '{print $1}')
cap9=$(( used9 + 2048 )) # 2 MB to spare: under the requirement, over nothing else
mkdir -p "$scratch/bin9"
cat > "$scratch/bin9/df" <<DFEOF
#!/usr/bin/env bash
used=\$($real_du -sk "$root" | awk '{print \$1}')
echo "Filesystem 1024-blocks Used Available Capacity Mounted-on"
echo "fake $cap9 \$used \$(( $cap9 - used )) 50% $root"
DFEOF
chmod +x "$scratch/bin9/df"
# Floor sits strictly between "0 evicted" (2048 KB free) and "1 evicted"
# (2048 + blob_kb free) — satisfiable by evicting exactly one candidate.
PATH="$scratch/bin9:$outer_path" \
CACHE_CLONE_HEADROOM_PERCENT=0 CACHE_CLONE_HEADROOM_FLOOR_KB=$(( 2048 + blob_kb / 2 )) \
GITHUB_STEP_SUMMARY="$scratch/summary" \
bash "$prune" "$root" "$root/target-$OWN" "dev main" 0 \
"$(cache_key unused-clone-probe)" "$DEV" "" \
> "$scratch/log" 2>&1 \
|| { cat "$scratch/log"; fail "prune-cache.sh exited non-zero"; }
assert_kept "$root/target-$OWN" "this run's own cache directory survives a genuinely pressured sibling pass"
assert_kept "$root/target-$OWN/blob" "and its contents survive — not a recreated empty directory"
assert_gone "$root/target-$LIVE" "the sibling is evicted instead, to make the same room"
if grep -q 'clearing own' "$scratch/log"; then
fail "own cache was cleared by pass 3, not genuinely spared by pass 2 — this scenario proves nothing"
fi
ok "the requirement was met by pass 2 alone; pass 3 never ran"
echo
echo "=== 7: target dirs evicted before snapshots ==="
@@ -235,7 +336,7 @@ done
exec "$real_du" "\$@"
EOF
chmod +x "$scratch/bin/du"
( PATH="$scratch/bin:$PATH"; run_prune "1000000 900000" )
( PATH="$scratch/bin:$outer_path"; run_prune "1000000 900000" )
[ -e "$root/.reading-target-$DEAD-racer" ] || fail "the racing marker was never published — scenario 12 proves nothing"
assert_kept "$root/target-$DEAD" "a cache claimed inside the eviction window is not unlinked"
assert_kept "$root/target-$DEAD/blob" "the reprieved cache still has its contents"
@@ -286,5 +387,172 @@ assert_kept "$aside" "an aside younger than the settle window is not reclaimed"
assert_kept "$aside/blob" "and is left intact, not part-way emptied"
assert_log "may still be evicting it" "the deferral gives its actual reason"
echo
echo "=== 15: a merged-but-undeleted branch is dead too ==="
# A branch the forge did not delete at merge, which `ls-remote` then reports
# forever. Built the way that happens: a branch merged into dev with a merge
# commit, still pushed, beside one branched at the same point and NOT merged.
git="git -c user.email=t@t -c user.name=t -c commit.gpgsign=false"
(
cd "$work"
git checkout -q dev
git checkout -q -b feat/merged
$git commit -q --allow-empty -m merged
git checkout -q dev
$git merge -q --no-ff feat/merged -m "merge feat/merged"
git checkout -q -b feat/unmerged
$git commit -q --allow-empty -m unmerged
git checkout -q dev
git push -q origin dev feat/merged feat/unmerged
)
MERGED=$(cache_key feat/merged); UNMERGED=$(cache_key feat/unmerged)
reset_cache
mk "target-$MERGED" '2030-01-01'
mk "snapshot-$MERGED" '2030-01-01'
mk "target-$UNMERGED" '2020-01-01' # older, deliberately: merge state decides, not age
run_prune "1000000 900000" # 90% free: no pressure at all
assert_log "merged-branch detection anchored on" "the pass says what it anchored ancestry on"
assert_gone "$root/target-$MERGED" "a merged branch's cache is pruned though its branch is still on origin"
assert_gone "$root/snapshot-$MERGED" "and so is its snapshot"
assert_kept "$root/target-$UNMERGED" "an unmerged branch's cache survives, though it is the older of the two"
assert_log "merged into dev" "the eviction names the branch it was merged into"
echo
echo "=== 16: 'cannot tell' is not death, per branch and per checkout ==="
# A branch whose tip this checkout has never seen. Pushed from a second clone,
# so `ls-remote` reports a SHA that `$work` holds no object for — which is
# what "cannot determine" actually looks like, rather than a stubbed failure.
other="$scratch/other"; git clone -q "$origin" "$other"
(
cd "$other"
git checkout -q -b feat/elsewhere origin/dev
git -c user.email=t@t -c user.name=t -c commit.gpgsign=false commit -q --allow-empty -m elsewhere
git push -q origin feat/elsewhere
)
ELSEWHERE=$(cache_key feat/elsewhere)
reset_cache
mk "target-$ELSEWHERE" '2030-01-01'
mk "target-$MERGED" '2030-01-01'
run_prune "1000000 900000"
assert_kept "$root/target-$ELSEWHERE" "a branch whose tip is not in this checkout is kept, not classified dead"
assert_log "cannot tell merged from live" "and the undecidable branch is named, not silently skipped"
assert_gone "$root/target-$MERGED" "while a branch it CAN decide is still pruned in the same pass"
# A shallow checkout, where a missing object is the ordinary case rather than
# a signal — so the whole merged half is withheld. The deleted-branch half is
# unaffected, which is what keeps this a narrowing rather than an outage.
shallow="$scratch/shallow"; git clone -q --depth 1 -b dev "file://$origin" "$shallow"
[ "$(git -C "$shallow" rev-parse --is-shallow-repository)" = true ] \
|| fail "the fixture clone is not shallow — scenario 16's second half proves nothing"
reset_cache
mk "target-$MERGED" '2030-01-01'
(
cd "$shallow"
CACHE_DF_OVERRIDE="1000000 900000" GITHUB_STEP_SUMMARY="$scratch/summary" \
bash "$prune" "$root" "$root/target-$OWN" "dev main" 10
) > "$scratch/log" 2>&1 || { cat "$scratch/log"; fail "prune-cache.sh exited non-zero in a shallow checkout"; }
assert_log "checkout is shallow" "a shallow checkout withholds the merged signal and says why"
assert_kept "$root/target-$MERGED" "and keeps a merged branch's cache rather than guessing"
assert_gone "$root/target-$DEAD" "while the deleted-branch signal still fires"
echo
echo "=== 17: the free-space requirement is measured off the clone's source ==="
# A `df` that answers from the cache root's actual size, because the property
# under test is that the pass STOPS once the requirement is met — which a
# fixed CACHE_DF_OVERRIDE cannot express, since evicting never changes it.
mkdir -p "$scratch/bin17"
real_du=$(command -v du)
build_seed_fixture() {
rm -rf "$root"; mkdir -p "$root"
# The source the seed is about to clone. 8 MB of dep-info, which
# unshare_mutable_paths has to real-copy, beside 16 MB of .rlib that it
# leaves hardlinked — so a requirement derived from the SIZE of the source
# would be three times the one derived from its mutable set.
mkdir -p "$root/snapshot-$DEV/debug/.fingerprint/unit" "$root/snapshot-$DEV/debug/deps"
head -c $((8 * 1024 * 1024)) /dev/zero > "$root/snapshot-$DEV/debug/.fingerprint/unit/dep-lib"
head -c $((16 * 1024 * 1024)) /dev/zero > "$root/snapshot-$DEV/debug/deps/libx.rlib"
touch -d '2020-01-01' "$root/snapshot-$DEV/.cache-last-used"
# Three live, unmerged branches' caches of 4 MB each, oldest first.
local i=0
for b in a b c; do
i=$((i + 1))
mkdir -p "$root/target-$(cache_key "feat/$b")"
head -c $((4 * 1024 * 1024)) /dev/zero > "$root/target-$(cache_key "feat/$b")/blob"
touch -d "202${i}-01-01" "$root/target-$(cache_key "feat/$b")/.cache-last-used"
done
# A volume with 2 MB to spare: under the requirement, over nothing else.
cap=$(( $($real_du -sk "$root" | awk '{print $1}') + 2048 ))
cat > "$scratch/bin17/df" <<DFEOF
#!/usr/bin/env bash
used=\$($real_du -sk "$root" | awk '{print \$1}')
echo "Filesystem 1024-blocks Used Available Capacity Mounted-on"
echo "fake $cap \$used \$(( $cap - used )) 50% $root"
DFEOF
chmod +x "$scratch/bin17/df"
}
(
cd "$work"
for b in a b c; do
git checkout -q dev
git checkout -q -b "feat/$b"
git -c user.email=t@t -c user.name=t -c commit.gpgsign=false commit -q --allow-empty -m "$b"
done
git checkout -q dev
git push -q origin feat/a feat/b feat/c
)
# run_seeded_prune <own-ref> <base-ref> — the form cargo-cache/action.yml uses:
# the same pass, told what the seed step it now runs ahead of will clone. The
# headroom knobs are pinned so the arithmetic is the fixture's, not the
# defaults' (whose 2 GiB floor would dwarf any fixture on a test host).
seeded_rc=0
run_seeded_prune() {
seeded_rc=0
PATH="$scratch/bin17:$outer_path" \
CACHE_CLONE_HEADROOM_PERCENT=100 CACHE_CLONE_HEADROOM_FLOOR_KB=1024 \
GITHUB_STEP_SUMMARY="$scratch/summary" \
bash "$prune" "$root" "$root/target-$(cache_key "$1")" "dev main" 0 \
"$(cache_key "$1")" "$(cache_key "$2")" "" \
> "$scratch/log" 2>&1 || seeded_rc=$?
}
build_seed_fixture
run_seeded_prune feat/own dev
[ "$seeded_rc" = 0 ] || { cat "$scratch/log"; fail "prune-cache.sh exited ${seeded_rc} with the requirement satisfiable"; }
assert_log "measured from its mutable set" "the requirement says where it came from"
assert_gone "$root/target-$(cache_key feat/a)" "the oldest cache is evicted to make room for the clone"
assert_gone "$root/target-$(cache_key feat/b)" "and the next oldest, because one was not enough"
assert_kept "$root/target-$(cache_key feat/c)" "and the pass STOPS there rather than draining the volume"
assert_kept "$root/snapshot-$DEV" "the source the seed is about to clone is never a candidate"
# Nothing eligible: every sibling is held open by a running job. The pass
# cannot reach the requirement, and the seed that follows would fail against a
# staging path naming none of this.
build_seed_fixture
for b in a b c; do date +%s > "$root/target-$(cache_key "feat/$b")/.ci-lock-ci-1"; done
run_seeded_prune feat/own dev
[ "$seeded_rc" = 1 ] || { cat "$scratch/log"; fail "expected exit 1 when the clone cannot fit, got ${seeded_rc}" ; }
ok "a clone that cannot be made to fit fails the pass rather than the seed"
assert_log "short by" "the failure names the shortfall"
assert_log "held open by a running job" "and what was kept instead of it, with the reason"
assert_kept "$root/target-$(cache_key feat/a)" "a locked cache is still not evicted, however tight the disk"
echo
echo "=== 18: the layout that produced the bug ==="
# zemyna's volume on 2026-09-07: the base branch's snapshot and target dir,
# plus one target dir for a branch merged the day before and never deleted.
# Equal sizes and 90% free, so neither age nor pressure nor size can be what
# decides — only the merge state.
rm -rf "$root"; mkdir -p "$root"
mk "snapshot-$DEV" '2026-09-01'
mk "target-$DEV" '2026-09-01'
mk "target-$MERGED" '2026-09-06'
run_prune "1000000 900000"
assert_kept "$root/snapshot-$DEV" "the base snapshot stays"
assert_kept "$root/target-$DEV" "and the base target dir stays"
assert_gone "$root/target-$MERGED" "and the merged-but-undeleted branch's cache is the one reclaimed"
[ "$(grep -c 'pruned dead-branch cache' "$scratch/log")" = 1 ] \
|| { cat "$scratch/log"; fail "expected exactly one eviction on the zemyna layout"; }
ok "exactly one directory is evicted, and it is that one"
echo
echo "prune-cache-selftest: ${pass_count} assertions passed"
+321 -46
View File
@@ -1,8 +1,16 @@
#!/usr/bin/env bash
# Eviction for the per-ref cache directories on the persistent volume.
#
# Usage: prune-cache.sh <cache-root> <own-target-dir> <protected-branches> <min-free-percent>
# Usage: prune-cache.sh <cache-root> <own-target-dir> <protected-branches> \
# <min-free-percent> [own-key] [base-key] [fallback-dir]
# protected-branches space-separated raw refs (e.g. "dev main")
# min-free-percent a FLOOR, not the gate — 0 to rely on the derived
# requirement alone (the default)
# own-key, base-key, fallback-dir
# the same three the seed step resolves its source from.
# Given them, this pass sizes the volume for the clone
# that step is about to make; without them it has only
# the percentage floor, and says so.
#
# Optional environment:
# STALE_LOCK_SECONDS age past which a .ci-lock-* marker is treated as
@@ -16,19 +24,42 @@
#
# Three passes, in order:
#
# 1. LIVENESS — every target-*/snapshot-* directory whose branch no longer
# exists on origin is removed UNCONDITIONALLY, not gated on free space.
# A directory for a branch deleted days ago is pure loss: nothing will
# ever read it again, since a merged PR's branch cannot be reopened.
# Waiting for disk pressure to notice means paying for it until then.
# Skipped entirely, loudly, if the liveness signal itself is
# unavailable — "couldn't determine" is never folded into "dead".
# 2. PRESSURE — if free space is still under the threshold, evict remaining
# (now necessarily live) directories oldest-first until it clears.
# 3. SELF-CLEAR — if pass 2 still isn't enough, wipe this run's own target
# dir and pay a cold rebuild, reported to the job summary as well as the
# log, because a warning on a green run is what lets a silently 4x-slower
# job go unnoticed.
# 1. LIVENESS — every target-*/snapshot-* directory whose branch is DEAD is
# removed UNCONDITIONALLY, not gated on free space. A directory for a
# branch nothing will build again is pure loss; waiting for disk pressure
# to notice means paying for it until then. Two signals make a branch
# dead, and the second exists because the first alone is inert wherever
# a merged branch stays on origin — the default, and still the outcome
# whenever delete-on-merge declines or is not asked (gitdan-actions#20):
#
# DELETED — the branch is no longer on origin at all.
# MERGED — the branch is still on origin, but its tip is an ancestor
# of a protected branch's tip, so every commit it holds is
# already on the branch its cache would be re-cloned from.
#
# Skipped entirely, loudly, if the signal itself is unavailable —
# "couldn't determine" is never folded into "dead", for either signal
# and at either granularity: a checkout that cannot answer ancestry at
# all skips the merged half, and a single branch whose tip is not in the
# checkout is kept with a warning naming it.
# 2. PRESSURE — if free space is under the requirement, evict remaining
# (now necessarily live) directories oldest-first until it clears. The
# requirement is what the seed step is about to need to clone its source,
# MEASURED off that source (see clone_headroom_kb in cache-lib.sh), and a
# percentage floor only if one is configured. A percentage cannot express
# this: the failure it has to prevent is a clone running out of disk
# part-way through unsharing its mutable paths, and how much that needs
# is a property of the snapshot, not of the volume.
# 3. SELF-CLEAR — if pass 2 still isn't enough for the percentage floor,
# wipe this run's own target dir and pay a cold rebuild, reported to the
# job summary as well as the log, because a warning on a green run is
# what lets a silently 4x-slower job go unnoticed. It is not a way out of
# the derived requirement: a run that has an own target dir to wipe is a
# run whose seed reuses it and clones nothing, so that requirement is
# zero. Falling short of a NON-ZERO derived requirement fails the job
# here, naming the shortfall and what was kept instead of it — the seed
# would otherwise fail seconds later against a half-unshared staging
# tree, which is the failure this pass exists to pre-empt.
#
# Reactive-only, with no hard cap on cache size: a workspace's natural working
# set is what it is, and bounding the footprint preemptively means wiping
@@ -36,21 +67,32 @@
# physics, not an arbitrary GB number.
#
# EVICTION ORDER, and why it is the reverse of the obvious one: within the
# pressure pass, `target-*` directories are evicted BEFORE `snapshot-*` ones.
# A snapshot is a hardlink clone of a live target dir and of every consumer
# cloned from it, so removing it frees almost no real bytes — its inodes stay
# alive through those other links — while costing every future PR its warm
# start. Evicting snapshots first would be nearly pure loss. Target
# directories are where a branch's own divergent artifacts actually live, so
# they are what freeing space means.
# pressure pass, `target-*` directories are evicted BEFORE `snapshot-*` ones
# (a publisher's own target dir included — see PROTECTED below). A publisher
# branch's target dir is a convenience cache that its own next run reseeds
# from the snapshot, so losing it is cheap; evicting the snapshot instead
# forces every subsequent PR to start cold. Measured on the live volume
# (daniel/zemyna#1073), that cold start is not the near-free move it looks
# like either: a snapshot shares almost nothing with the target dirs cloned
# from it, because publish-snapshot.sh unshares every executable after its
# `cp -al` (cargo and the linker rewrite binaries in place, so a shared
# original would corrupt under them), and executables are the great majority
# of the tree by bytes. So a snapshot eviction is a real, large disk cost as
# well as a cold-start one — reserved for last because both costs are larger
# than a target dir's.
#
# Two exclusions every pass respects:
#
# PROTECTED — the publisher branches' target and snapshot directories, and
# this run's own target dir, are never candidates in any pass. Evicting a
# publisher's snapshot doesn't free real disk (every open PR's clone keeps
# the data alive) but does force every subsequent PR to start cold, which is
# the entire benefit this scheme exists to deliver.
# PROTECTED — a publisher branch's SNAPSHOT directory, and this run's own
# target dir, are never candidates in the pressure or self-clear passes. A
# publisher branch's own TARGET dir is not protected there: it is an
# ordinary pressure-pass candidate, evicted oldest-first like any other,
# because nothing downstream depends on it surviving — the publisher's own
# next run reseeds it from the snapshot. It IS excluded from pass 1 alone
# (is_protected_from_liveness): a branch's tip is trivially an ancestor of
# itself, so without this exclusion the merged-branch signal would read a
# publisher's own target dir as "merged into itself" and pass 1 —
# unconditional, not gated on pressure — would delete it every run.
#
# LOCKED — a directory carrying a .ci-lock-* marker younger than
# STALE_LOCK_SECONDS is held open by a running job, or named by a live
@@ -75,16 +117,34 @@
# than reimplementing a lookalike is what makes the classification sound; any
# drift between two spellings would silently misclassify every directory.
#
# The merged half reads the TIP SHA out of that same `ls-remote` output and
# asks `git merge-base --is-ancestor` against each protected branch's tip,
# using the objects in this job's own checkout. Ancestry is only decidable
# where the objects are there to decide it, so the answer "I cannot tell"
# exists and is distinct from "not merged" everywhere it can arise:
#
# - a shallow checkout makes a MISSING object prove nothing, so the whole
# signal is withheld rather than read as "no branch is merged";
# - a protected tip that is not in the checkout is not used as an anchor;
# - a branch tip that is not in the checkout is kept, loudly.
#
# A squash or rebase merge leaves no ancestor relationship at all, so its
# branch reads as live here. That is a missed reclamation, not a wrong one,
# and the deleted-branch signal still covers it once the branch is removed.
#
# Only ever globs inside <cache-root>. Another project's volume is a different
# Docker named volume and is not mounted in this container at all, so "stays
# scoped to this repo's cache" holds structurally, not by convention.
set -euo pipefail
. "$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/cache-lib.sh"
ROOT="${1:?usage: prune-cache.sh <cache-root> <own-target-dir> <protected-branches> <min-free-percent>}"
ROOT="${1:?usage: prune-cache.sh <cache-root> <own-target-dir> <protected-branches> <min-free-percent> [own-key] [base-key] [fallback-dir]}"
OWN_DIR="${2:?}"
PROTECTED_REFS="${3:-}"
MIN_FREE_PCT="${4:-10}"
MIN_FREE_PCT="${4:-0}"
OWN_KEY="${5:-}"
BASE_KEY="${6:-}"
FALLBACK="${7:-}"
# Mirrored by daniel/gitdan's ci-cache-reclaim.sh, whose copy must be >= this
# one — raising this without raising theirs first lets that script treat a lock
# this side still honours as abandoned. Same direction and same reasoning as
@@ -102,21 +162,81 @@ STALE_LOCK_SECONDS="${STALE_LOCK_SECONDS:-7200}"
# the volume is under pressure.
EVICTION_ASIDE_SETTLE_SECONDS="${EVICTION_ASIDE_SETTLE_SECONDS:-60}"
# What the seed step is about to do, resolved through the same function that
# step resolves it with (cache-lib.sh's seed_source_candidates). Empty when it
# will clone nothing at all: its own target dir already exists and it reuses
# it, or no source exists and it starts cold. Either way the derived
# requirement is zero, because nothing is about to be copied.
#
# gb() is for reporting only. Every comparison below is in KB, because
# `read_df` reports KB and rounding a threshold to a tenth of a GB either
# passes a run that cannot fit or evicts a cache the run did not need.
gb() { awk -v k="${1:-0}" 'BEGIN { printf "%.1f", k / 1048576 }'; }
SEED_SRC=""
CLONE_KB=0
if [ -n "$OWN_KEY" ]; then
SEED_SRC=$(seed_clone_source "$ROOT" "$OWN_KEY" "$BASE_KEY" "$FALLBACK")
if [ -n "$SEED_SRC" ]; then
# A full walk of the source, and the reason this pass moved ahead of the
# seed rather than staying where it was: the number is only useful before
# the clone it describes.
CLONE_KB=$(clone_headroom_kb "$SEED_SRC")
echo "clone requirement: seeding from $(basename "$SEED_SRC") needs $(gb "$CLONE_KB") GB free (measured from its mutable set)"
else
echo "clone requirement: none — this run reuses its own cache or starts cold, so nothing will be copied"
fi
else
echo "clone requirement: not derivable — no cache key was passed to this pass; the ${MIN_FREE_PCT}% floor is the only gate"
fi
declare -A protected_ns=()
# A protected ref's own target dir is excluded from pass 1 ONLY (see
# is_protected_from_liveness below), never from the pressure pass. It must
# stay out of pass 1 for a reason that has nothing to do with disk: a
# protected ref's tip is trivially an ancestor of itself, so without this
# is_merged_dead would read dev's own target dir as "merged into dev" and
# pass 1 — unconditional, not gated on pressure — would delete it on every
# single run.
declare -A protected_target_ns=()
for ref in $PROTECTED_REFS; do
suffix=$(cache_key "$ref")
protected_ns["target-${suffix}"]=1
protected_ns["snapshot-${suffix}"]=1
protected_target_ns["target-${suffix}"]=1
done
is_protected() {
# Prints why <dir> is off limits to the pressure/self-clear passes, or
# nothing when it is a candidate there. The reason is not decoration: it is
# what the failure report at the bottom lists against each directory it kept
# while running out of space.
#
# SEED_SRC is the third exclusion and the one this script did not used to
# need. The prune ran after the seed, so the source had already been cloned
# and the reader marker over it was gone; running BEFORE the seed puts the
# directory this run is about to read squarely in the candidate set, and
# pass 1 would take it the moment its branch merged.
protected_reason() {
local dir="$1" name
name=$(basename "$dir")
[ "$dir" = "$OWN_DIR" ] && return 0
[ -n "${protected_ns[$name]:-}" ] && return 0
[ "$dir" = "$OWN_DIR" ] && { printf 'this run own cache'; return 0; }
[ -n "$SEED_SRC" ] && [ "$dir" = "$SEED_SRC" ] && { printf 'the source this run is about to clone'; return 0; }
[ -n "${protected_ns[$name]:-}" ] && { printf 'a protected branch snapshot'; return 0; }
return 1
}
is_protected() { protected_reason "$1" >/dev/null; }
# Pass 1 (liveness) only: also excludes a protected ref's own target dir, for
# the self-ancestor reason above protected_target_ns documents. The pressure
# pass does not call this — is_protected is what it uses, via protected_reason
# directly.
is_protected_from_liveness() {
local dir="$1" name
is_protected "$dir" && return 0
name=$(basename "$dir")
[ -n "${protected_target_ns[$name]:-}" ]
}
# is_locked <dir> [name]
#
# `name` is the directory's own name for reporting and for the reader-marker
@@ -277,6 +397,11 @@ done
echo "=== pass 1: liveness (unconditional, not gated on free space) ==="
declare -A live_ns=()
# The tip SHA origin reports for the branch each directory name belongs to.
# Same output, same loop, one field over — reading it from a second `git` call
# would be reading a different instant.
declare -A live_tip=()
declare -A remote_tip_of=()
LIVENESS_AVAILABLE=0
LIVENESS_REASON=""
if [ "${CACHE_LIVENESS:-true}" = "0" ] || [ "${CACHE_LIVENESS:-true}" = "false" ]; then
@@ -289,9 +414,13 @@ elif remote_heads=$(timeout 20 git ls-remote --heads origin 2>&1); then
case "$line" in *refs/heads/*) ;; *) continue ;; esac
branch="${line#*refs/heads/}"
[ -n "$branch" ] || continue
sha="${line%%[[:space:]]*}"
suffix=$(cache_key "$branch")
live_ns["target-${suffix}"]=1
live_ns["snapshot-${suffix}"]=1
live_tip["target-${suffix}"]="$sha"
live_tip["snapshot-${suffix}"]="$sha"
remote_tip_of["$branch"]="$sha"
branch_count=$((branch_count + 1))
done <<< "$remote_heads"
echo "liveness: ${branch_count} live branches on origin"
@@ -299,6 +428,80 @@ else
LIVENESS_REASON="git ls-remote --heads origin failed or timed out"
fi
# The merged half of pass 1, and whether this checkout can answer it at all.
# Every branch of this decision that ends in "no" ends in the signal being
# WITHHELD, never in a directory being classified dead by default.
MERGED_AVAILABLE=0
MERGED_REASON=""
PROTECTED_TIPS=()
declare -A merged_verdict=()
declare -A merged_into=()
MERGED_INTO=""
if [ "$LIVENESS_AVAILABLE" = "1" ]; then
if ! git rev-parse --git-dir >/dev/null 2>&1; then
MERGED_REASON="not inside a git checkout"
elif [ "$(git rev-parse --is-shallow-repository 2>/dev/null || echo unknown)" != "false" ]; then
# In a shallow clone an absent commit is the normal case, so `--is-ancestor`
# answers about the graph that was fetched rather than the one that exists.
MERGED_REASON="the checkout is shallow, so a commit missing from it says nothing about ancestry"
else
for ref in $PROTECTED_REFS; do
tip="${remote_tip_of[$ref]:-}"
[ -n "$tip" ] || continue
git cat-file -e "${tip}^{commit}" 2>/dev/null || continue
PROTECTED_TIPS+=("${ref}:${tip}")
done
if [ "${#PROTECTED_TIPS[@]}" -gt 0 ]; then
MERGED_AVAILABLE=1
echo "liveness: merged-branch detection anchored on ${PROTECTED_TIPS[*]%%:*}"
else
MERGED_REASON="none of the protected branch tips (${PROTECTED_REFS:-none configured}) is present in this checkout"
fi
fi
[ "$MERGED_AVAILABLE" = "1" ] || \
echo "::warning::liveness: ${MERGED_REASON} — merged-but-undeleted branches keep their caches this run"
fi
# is_merged_dead <dir-name>
#
# True when the branch this directory belongs to is still on origin but every
# commit it holds is already on a protected branch — a merged PR whose branch
# the forge did not delete, which the deleted-branch signal above can never
# see. Enabling delete-on-merge narrows this to the branches merged before it
# was enabled, the ones its deletion declines (protected, or used by another
# open PR), and the merges that never ask (an API merge without the flag);
# see README's eviction section.
#
# Memoised per tip because target-<key> and snapshot-<key> share one branch,
# and because the "cannot tell" warning belongs to the branch rather than to
# each of its directories.
is_merged_dead() {
local name="$1" tip="${live_tip[$1]:-}" entry ref psha
MERGED_INTO=""
[ "$MERGED_AVAILABLE" = "1" ] || return 1
[ -n "$tip" ] || return 1
case "${merged_verdict[$tip]:-}" in
dead) MERGED_INTO="${merged_into[$tip]}"; return 0 ;;
live|unknown) return 1 ;;
esac
if ! git cat-file -e "${tip}^{commit}" 2>/dev/null; then
merged_verdict["$tip"]=unknown
echo "::warning::prune: ${name}: its branch tip ${tip} is not in this checkout — cannot tell merged from live, keeping it"
return 1
fi
for entry in "${PROTECTED_TIPS[@]}"; do
ref="${entry%%:*}"; psha="${entry#*:}"
if git merge-base --is-ancestor "$tip" "$psha" 2>/dev/null; then
merged_verdict["$tip"]=dead
merged_into["$tip"]="$ref"
MERGED_INTO="$ref"
return 0
fi
done
merged_verdict["$tip"]=live
return 1
}
if [ "$LIVENESS_AVAILABLE" = "1" ]; then
pruned_any=0
# Tracked separately so the line below cannot contradict the decline lines
@@ -308,13 +511,21 @@ if [ "$LIVENESS_AVAILABLE" = "1" ]; then
for dir in "$ROOT"/target-* "$ROOT"/snapshot-*; do
[ -d "$dir" ] || continue
name=$(basename "$dir")
is_protected "$dir" && continue
[ -n "${live_ns[$name]:-}" ] && continue
is_protected_from_liveness "$dir" && continue
if [ -z "${live_ns[$name]:-}" ]; then
why="no matching branch on origin"
why_summary="branch no longer exists on origin"
elif is_merged_dead "$name"; then
why="merged into ${MERGED_INTO}, whose tip already contains its every commit"
why_summary="merged into \`${MERGED_INTO}\`"
else
continue
fi
is_locked "$dir" && { spared_any=1; continue; }
dir_gb=$(usage_gb "$dir")
evict_dir "$dir" || { spared_any=1; continue; }
echo "::warning::pruned dead-branch cache ${name} (${dir_gb} GB) — no matching branch on origin"
summary_line "- pruned dead-branch cache \`${name}\` (${dir_gb} GB) — branch no longer exists on origin"
echo "::warning::pruned dead-branch cache ${name} (${dir_gb} GB) — ${why}"
summary_line "- pruned dead-branch cache \`${name}\` (${dir_gb} GB) — ${why_summary}"
pruned_any=1
done
if [ "$pruned_any" = "0" ]; then
@@ -329,16 +540,45 @@ else
fi
echo
echo "=== pass 2/3: disk pressure (threshold: free < ${MIN_FREE_PCT}%) ==="
echo "=== pass 2/3: disk pressure ==="
read -r TOTAL_KB FREE_KB <<< "$(read_df "$ROOT")"
THRESHOLD_KB=$(( TOTAL_KB * MIN_FREE_PCT / 100 ))
PCT_KB=$(( TOTAL_KB * MIN_FREE_PCT / 100 ))
if [ "$FREE_KB" -ge "$THRESHOLD_KB" ]; then
echo "cache: $(basename "$OWN_DIR") $(usage_gb "$OWN_DIR") GB | $(report_df host "$FREE_KB" "$TOTAL_KB")"
# The two floors, and which of them governs. They are kept apart all the way
# down rather than collapsed here, because falling short of them means
# different things: the derived one predicts that the very next step cannot
# finish, and the percentage one is a hygiene target for the volume.
REQUIRED_KB="$CLONE_KB"
GOVERNS="the clone this run is about to make"
if [ "$PCT_KB" -gt "$REQUIRED_KB" ]; then
REQUIRED_KB="$PCT_KB"
GOVERNS="the ${MIN_FREE_PCT}% floor"
fi
echo "required: $(gb "$REQUIRED_KB") GB free — ${GOVERNS} (clone $(gb "$CLONE_KB") GB, floor $(gb "$PCT_KB") GB)"
own_report() {
if [ -d "$OWN_DIR" ]; then
echo "cache: $(basename "$OWN_DIR") $(usage_gb "$OWN_DIR") GB | $(report_df host "$1" "$2")"
else
# Ordinary now that this pass runs ahead of the seed: on a branch's first
# run of the day the directory does not exist yet, and reporting 0.0 GB
# for it would read as an emptied cache.
echo "cache: $(basename "$OWN_DIR") not seeded yet | $(report_df host "$1" "$2")"
fi
}
if [ "$FREE_KB" -ge "$REQUIRED_KB" ]; then
own_report "$FREE_KB" "$TOTAL_KB"
exit 0
fi
echo "::warning::$(report_df disk "$FREE_KB" "$TOTAL_KB") < ${MIN_FREE_PCT}% threshold"
echo "::warning::$(report_df disk "$FREE_KB" "$TOTAL_KB") < the $(gb "$REQUIRED_KB") GB this run requires"
# What survived the pass, and why, in the order the pass considered them. Read
# only by the failure report at the bottom: a run that cannot fit its clone is
# a run whose log has to answer "then what is all that space?" without anyone
# having to reconstruct the pass by hand.
KEPT=()
# A plain loop over a pre-materialised, pre-sorted list rather than a live
# `find | while` pipeline, so `rm -rf` inside the loop cannot make a running
@@ -346,21 +586,56 @@ echo "::warning::$(report_df disk "$FREE_KB" "$TOTAL_KB") < ${MIN_FREE_PCT}% thr
mapfile -t LRU < <(list_by_lru)
for dir in "${LRU[@]}"; do
[ -d "$dir" ] || continue
is_protected "$dir" && continue
if reason=$(protected_reason "$dir"); then
KEPT+=("$(basename "$dir") — ${reason}")
continue
fi
read -r TOTAL_KB FREE_KB <<< "$(read_df "$ROOT")"
[ "$FREE_KB" -ge "$THRESHOLD_KB" ] && break
is_locked "$dir" && continue
[ "$FREE_KB" -ge "$REQUIRED_KB" ] && break
if is_locked "$dir"; then
KEPT+=("$(basename "$dir") — held open by a running job")
continue
fi
dir_gb=$(usage_gb "$dir")
evict_dir "$dir" || continue
if ! evict_dir "$dir"; then
KEPT+=("$(basename "$dir") — claimed by a job while its eviction was in flight")
continue
fi
echo "::warning::evicted $(basename "$dir") (${dir_gb} GB, LRU under disk pressure)"
summary_line "- evicted \`$(basename "$dir")\` (${dir_gb} GB, LRU under disk pressure)"
done
read -r TOTAL_KB FREE_KB <<< "$(read_df "$ROOT")"
if [ "$FREE_KB" -lt "$THRESHOLD_KB" ]; then
# Falling short of the DERIVED requirement is a failure, not a warning. The
# seed step is next, it will clone that source, and it will run out of disk
# part-way through unsharing the clone's mutable paths — reported against a
# staging path, with nothing in the message about which cache was holding the
# space. Failing here says that instead.
#
# There is nothing to self-clear on this path and it is not skipped in error:
# a non-zero requirement means the seed is about to CLONE, which means this
# run has no own target dir to wipe (seed_clone_source returns nothing when it
# does), so pass 3 has no candidate. See the header.
if [ "$FREE_KB" -lt "$CLONE_KB" ]; then
echo "::error::prune: $(gb "$FREE_KB") GB free after evicting every eligible cache, but seeding from $(basename "$SEED_SRC") needs $(gb "$CLONE_KB") GB — short by $(gb "$(( CLONE_KB - FREE_KB ))") GB"
summary_line "- **out of disk**: seeding from \`$(basename "$SEED_SRC")\` needs $(gb "$CLONE_KB") GB, $(gb "$FREE_KB") GB free"
echo "prune: kept, and why:"
for entry in ${KEPT[@]+"${KEPT[@]}"}; do echo " ${entry}"; done
[ "${#KEPT[@]}" -gt 0 ] || echo " (nothing — the volume holds no cache directories at all)"
exit 1
fi
if [ "$FREE_KB" -lt "$PCT_KB" ]; then
OWN_GB=$(usage_gb "$OWN_DIR")
echo "::warning::still under threshold after evicting every eligible sibling; clearing own $(basename "$OWN_DIR") (was ${OWN_GB} GB) — this run pays a cold rebuild"
summary_line "- **self-clear**: \`$(basename "$OWN_DIR")\` (was ${OWN_GB} GB) wiped — this run pays a cold rebuild"
# Recreated empty rather than left absent, and that is what keeps this path
# out of the requirement above: the seed reuses an own target dir that
# exists, whatever is in it, so a self-cleared run clones nothing and needs
# no headroom. Leaving it absent would send that run to the base snapshot
# instead, needing a clone this pass has just spent its last eligible bytes
# not sizing for.
rm -rf "$OWN_DIR"
mkdir -p "$OWN_DIR"
else
+4 -2
View File
@@ -53,9 +53,11 @@ pass_count=0
fail() { echo "ASSERTION FAILED: $*" >&2; exit 1; }
ok() { pass_count=$((pass_count + 1)); echo "PASS: $*"; }
# Cargo and rustc REPLACE an artifact (write elsewhere, rename over the path)
# rustc REPLACES an `.rlib`/`.rmeta` (writes elsewhere, renames over the path)
# rather than truncating it in place, which is exactly why a snapshot may
# share artifact inodes with the live target dir it was cloned from. The
# share those inodes with the live target dir it was cloned from. Linker
# outputs are the exception and are real-copied instead — see
# `unshare_mutable_paths` in cache-lib.sh. The
# fixtures here have to model that faithfully — a plain `>` redirect truncates
# in place and would write straight through the shared inode into the
# snapshot and every consumer, which is a property of the test fixture, not of
+180
View File
@@ -0,0 +1,180 @@
#!/usr/bin/env bash
# Regression test for release-v1.sh against a scratch bare origin: v1 reaches
# main's tip once it has been gated, never lands on an ungated commit, and
# never moves backwards when two writers race (gitdan-actions#27).
#
# run_sweep mirrors release-sweep.yaml's step order -- check, gate only when
# needed, push the gated tip leased on the v1 the check read -- with the gate
# stood in for by a command, so a failing gate is a failing sweep.
set -euo pipefail
script_dir=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)
release="$script_dir/release-v1.sh"
scratch=$(mktemp -d)
trap 'rm -rf "$scratch"' EXIT
pass_count=0
fail() { echo "ASSERTION FAILED: $*" >&2; exit 1; }
ok() { pass_count=$((pass_count + 1)); echo "PASS: $*"; }
export GIT_AUTHOR_NAME=t GIT_AUTHOR_EMAIL=t@t GIT_COMMITTER_NAME=t GIT_COMMITTER_EMAIL=t@t
unset GITHUB_OUTPUT
# A fresh origin with main at one commit and v1 on it; dev pushes to main,
# ci and ci2 are the runners' clones.
fresh() {
rm -rf "$scratch/w"; mkdir -p "$scratch/w"
git init -q --bare "$scratch/w/origin.git"
git clone -q "$scratch/w/origin.git" "$scratch/w/dev" 2>/dev/null
commit_to_main >/dev/null
git -C "$scratch/w/dev" push -q origin HEAD:refs/tags/v1
git clone -q "$scratch/w/origin.git" "$scratch/w/ci"
git clone -q "$scratch/w/origin.git" "$scratch/w/ci2"
}
commit_to_main() {
git -C "$scratch/w/dev" commit -q --allow-empty -m "c$RANDOM"
git -C "$scratch/w/dev" push -q origin HEAD:refs/heads/main
git -C "$scratch/w/dev" rev-parse HEAD
}
origin_v1() { git -C "$scratch/w/origin.git" rev-parse -q --verify 'refs/tags/v1^{commit}' || true; }
origin_tip() { git -C "$scratch/w/origin.git" rev-parse refs/heads/main; }
in_ci() { (cd "$scratch/w/${CLONE:-ci}" && bash "$release" "$@"); }
field() { sed -n "s/^$1=//p"; }
run_sweep() {
local gate="$1" out tip v1
out=$(in_ci sweep-check)
[ "$(field needed <<<"$out")" = true ] || return 0
tip=$(field tip <<<"$out"); v1=$(field v1 <<<"$out")
"$gate" || return 1
in_ci push "$tip" "$v1"
}
echo "=== 1. sweep with v1 at the tip is a no-op ==="
fresh
before=$(origin_v1)
out=$(in_ci sweep-check)
[ "$(field needed <<<"$out")" = false ] || fail "a current v1 was reported as lagging"
run_sweep false || fail "a current v1 ran the gate"
[ "$(origin_v1)" = "$before" ] || fail "a no-op sweep moved v1"
ok "v1 == tip: needed=false, gate not run, v1 unchanged"
echo
echo "=== 2. sweep with v1 ahead of the tip is a no-op ==="
fresh
ahead=$(git -C "$scratch/w/dev" commit-tree -p HEAD -m ahead 'HEAD^{tree}')
git -C "$scratch/w/dev" push -q -f origin "$ahead:refs/tags/v1"
out=$(in_ci sweep-check)
[ "$(field needed <<<"$out")" = false ] || fail "a v1 descending from the tip was reported as lagging"
ok "v1 descends from tip: needed=false"
echo
echo "=== 3. sweep with v1 behind and a passing gate tags the tip ==="
fresh
tip=$(commit_to_main)
run_sweep true || fail "a passing sweep failed"
[ "$(origin_v1)" = "$tip" ] || fail "v1 is $(origin_v1), not the gated tip $tip"
ok "v1 behind, gate green: v1 -> tip"
echo
echo "=== 4. sweep with v1 behind and a failing gate goes red and tags nothing ==="
fresh
before=$(origin_v1)
commit_to_main >/dev/null
if run_sweep false; then fail "a sweep over a failing gate succeeded"; fi
[ "$(origin_v1)" = "$before" ] || fail "a failing gate still moved v1"
ok "v1 behind, gate red: sweep red, v1 unchanged"
echo
echo "=== 5. a lost lease to a newer writer is a clean skip, never a step back ==="
fresh
t1=$(commit_to_main)
out=$(in_ci sweep-check); v1_read=$(field v1 <<<"$out")
t2=$(commit_to_main)
CLONE=ci2 in_ci merge "$t2" >/dev/null
[ "$(origin_v1)" = "$t2" ] || fail "the merge job did not release its own tip"
in_ci push "$t1" "$v1_read" || fail "a lease lost to a newer v1 went red"
[ "$(origin_v1)" = "$t2" ] || fail "v1 went backwards from $t2 to $(origin_v1)"
ok "older writer lost the lease: exit 0, v1 stays at the newer $t2"
git -C "$scratch/w/ci" push -q -f origin "$t1:refs/tags/v1"
[ "$(origin_v1)" = "$t1" ] || fail "control: an unleased push did not step v1 back"
ok "control: the same push without the lease steps v1 back to $t1"
echo
echo "=== 6. a lease lost to an older writer retries and lands the newer commit ==="
fresh
t1=$(commit_to_main)
t2=$(commit_to_main)
out=$(in_ci sweep-check); v1_read=$(field v1 <<<"$out")
git -C "$scratch/w/dev" push -q -f origin "$t1:refs/tags/v1"
in_ci push "$t2" "$v1_read" || fail "a lease lost to an older v1 went red"
[ "$(origin_v1)" = "$t2" ] || fail "v1 is $(origin_v1), not $t2"
ok "v1 moved to an ancestor under us: retried, v1 -> $t2"
echo
echo "=== 7. a lease lost to an unrelated commit goes red ==="
fresh
tip=$(commit_to_main)
out=$(in_ci sweep-check); v1_read=$(field v1 <<<"$out")
stray=$(git -C "$scratch/w/dev" commit-tree -m stray 'HEAD^{tree}')
git -C "$scratch/w/dev" push -q -f origin "$stray:refs/tags/v1"
if in_ci push "$tip" "$v1_read" 2>/dev/null; then fail "a v1 moved sideways was accepted"; fi
[ "$(origin_v1)" = "$stray" ] || fail "the stray v1 was overwritten"
ok "v1 moved to a commit neither ahead nor behind: red, v1 untouched"
echo
echo "=== 8. a push rejected for another reason goes red ==="
fresh
tip=$(commit_to_main)
mkdir -p "$scratch/w/origin.git/hooks"
printf '#!/bin/sh\nexit 1\n' > "$scratch/w/origin.git/hooks/pre-receive"
chmod +x "$scratch/w/origin.git/hooks/pre-receive"
before=$(origin_v1)
if in_ci merge "$tip" 2>"$scratch/err"; then fail "a rejected push reported success"; fi
[ "$(origin_v1)" = "$before" ] || fail "v1 moved despite the rejection"
grep -q 'not a lost lease' "$scratch/err" || fail "the rejection was not diagnosed as one: $(cat "$scratch/err")"
ok "server rejection with v1 unmoved: red, not a lost lease"
echo
echo "=== 9. the merge job defers on a moved tip; the next sweep catches up ==="
# The stranded trace: C1's job runs after C2 merged and defers — a run that
# deferred and left no newer run behind it — and merges stop.
fresh
before=$(origin_v1)
c1=$(commit_to_main)
c2=$(commit_to_main)
in_ci merge "$c1" >/dev/null || fail "the deferring merge job went red"
[ "$(origin_v1)" = "$before" ] || fail "the merge job released a commit that was not the tip"
run_sweep true || fail "the catch-up sweep failed"
[ "$(origin_v1)" = "$c2" ] || fail "v1 is $(origin_v1), not the tip $c2"
ok "C1 deferred, C2 never ran: the sweep moved v1 to $c2"
echo
echo "=== 10. the merge job releases its own tip, and creates a missing v1 ==="
fresh
tip=$(commit_to_main)
in_ci merge "$tip" >/dev/null
[ "$(origin_v1)" = "$tip" ] || fail "the merge job did not release the tip"
git -C "$scratch/w/dev" push -q origin :refs/tags/v1
tip=$(commit_to_main)
in_ci merge "$tip" >/dev/null
[ "$(origin_v1)" = "$tip" ] || fail "the merge job did not create an absent v1"
[ "$(origin_tip)" = "$tip" ] || fail "main moved"
ok "tip == gated sha: released, including onto an absent v1"
echo
echo "=== 11. a v1 hand-placed on an unrelated commit is never silently overwritten ==="
# Unlike #7, nothing races here -- v1 already sits on the stray commit before
# the very first push attempt, so force-with-lease sees exactly the value it
# expects and would otherwise succeed outright.
fresh
stray=$(git -C "$scratch/w/dev" commit-tree -m stray 'HEAD^{tree}')
git -C "$scratch/w/dev" push -q -f origin "$stray:refs/tags/v1"
tip=$(commit_to_main)
if in_ci merge "$tip" 2>"$scratch/err"; then fail "an unrelated hand-placed v1 was overwritten"; fi
[ "$(origin_v1)" = "$stray" ] || fail "v1 moved off the hand-placed $stray"
grep -q "$stray" "$scratch/err" || fail "the error did not name the stray v1: $(cat "$scratch/err")"
grep -q "$tip" "$scratch/err" || fail "the error did not name the gated sha: $(cat "$scratch/err")"
ok "hand-placed v1, unrelated to tip: red on the first push, v1 untouched"
echo
echo "release-v1-selftest: all $pass_count assertions passed"
+109
View File
@@ -0,0 +1,109 @@
#!/usr/bin/env bash
# Moves the floating `v1` tag forward to a gated commit on `main`, and never
# backwards. Run from a clone whose `origin` is this repository.
#
# release-v1.sh merge <gated-sha> merge-triggered job: release <gated-sha>
# if it is still main's tip
# release-v1.sh sweep-check scheduled sweep: report whether v1 lags
# main (tip=, v1=, needed= to
# $GITHUB_OUTPUT, or stdout without one)
# release-v1.sh push <gated-sha> <v1-as-read>
# scheduled sweep, after gating the tip
#
# Every push is leased on the v1 value the caller reasoned about. A lost lease
# means another writer moved v1 first: that is a clean skip once v1 is at or
# ahead of <gated-sha>, a retry against the new value while v1 is still behind
# it, and a failure otherwise.
set -euo pipefail
MAX_ATTEMPTS=3
fetch_main() {
git fetch -q origin +refs/heads/main:refs/remotes/origin/main
git rev-parse refs/remotes/origin/main
}
# Prints origin's v1 commit, or nothing when origin has no v1.
fetch_v1() {
if [ -z "$(git ls-remote origin refs/tags/v1)" ]; then
git update-ref -d refs/release-v1/seen 2>/dev/null || true
return 0
fi
git fetch -q origin +refs/tags/v1:refs/release-v1/seen
git rev-parse 'refs/release-v1/seen^{commit}'
}
# True when v1 already covers <sha>: at it, or a descendant of it.
covers() {
local sha="$1" v1="$2"
[ -n "$v1" ] && git merge-base --is-ancestor "$sha" "$v1"
}
push_leased() {
local sha="$1" expect="$2" now attempt
# force-with-lease only compares the ref's current value, not ancestry, so
# an unrelated v1 -- neither behind <sha> nor covering it -- would
# otherwise be silently overwritten on the very first push.
if [ -n "$expect" ] && ! covers "$sha" "$expect" && ! git merge-base --is-ancestor "$expect" "$sha"; then
echo "ERROR: v1 ($expect) is neither an ancestor of $sha nor at/ahead of it -- refusing to overwrite an unrelated v1" >&2
return 1
fi
for ((attempt = 1; attempt <= MAX_ATTEMPTS; attempt++)); do
if git push -q --force-with-lease="refs/tags/v1:$expect" origin "$sha:refs/tags/v1"; then
echo "v1 moved ${expect:-<absent>} -> $sha"
return 0
fi
now=$(fetch_v1)
if [ "$now" = "$expect" ]; then
echo "ERROR: push of v1 -> $sha rejected while v1 was still ${expect:-<absent>} -- not a lost lease" >&2
return 1
fi
if covers "$sha" "$now"; then
echo "lost the lease: another writer moved v1 to $now, at or ahead of $sha -- nothing to do"
return 0
fi
if [ -n "$now" ] && ! git merge-base --is-ancestor "$now" "$sha"; then
echo "ERROR: v1 moved to $now, which is neither behind nor ahead of $sha" >&2
return 1
fi
echo "lost the lease: v1 moved to ${now:-<absent>}, still behind $sha -- retrying"
expect="$now"
done
echo "ERROR: lost the lease on v1 $MAX_ATTEMPTS times running" >&2
return 1
}
cmd="${1:?usage: release-v1.sh merge <sha> | sweep-check | push <sha> <v1-as-read>}"
shift
case "$cmd" in
merge)
SHA="${1:?usage: release-v1.sh merge <gated-sha>}"
TIP=$(fetch_main)
if [ "$TIP" != "$SHA" ]; then
echo "main's tip ($TIP) is past this run's gated commit ($SHA) -- deferring; the sweep releases the tip"
exit 0
fi
V1=$(fetch_v1)
if covers "$SHA" "$V1"; then
echo "v1 ($V1) already at or ahead of $SHA -- nothing to do"
exit 0
fi
push_leased "$SHA" "$V1"
;;
sweep-check)
TIP=$(fetch_main)
V1=$(fetch_v1)
if covers "$TIP" "$V1"; then NEEDED=false; else NEEDED=true; fi
echo "main=$TIP v1=${V1:-<absent>} release-needed=$NEEDED"
printf 'tip=%s\nv1=%s\nneeded=%s\n' "$TIP" "$V1" "$NEEDED" >> "${GITHUB_OUTPUT:-/dev/stdout}"
;;
push)
push_leased "${1:?usage: release-v1.sh push <gated-sha> <v1-as-read>}" "${2-}"
;;
*)
echo "release-v1.sh: unknown command '$cmd'" >&2
exit 2
;;
esac
+2 -1
View File
@@ -64,7 +64,8 @@
# two refs' fingerprints ever share a directory and this script never has to
# arbitrate freshness across refs — only within one ref's own history, which
# is exactly what it is built to do soundly. On a nightly toolchain,
# CARGO_UNSTABLE_CHECKSUM_FRESHNESS is a complementary, stronger guarantee
# CARGO_UNSTABLE_CHECKSUM_FRESHNESS plus CARGO_BUILD_FINGERPRINT=content (both,
# since cargo PR #17382 on 2026-08-22) is a complementary, stronger guarantee
# (content-addressed rather than mtime-based freshness); this script is not
# made redundant by it, because directory-form `rerun-if-changed` build-script
# watches are not covered by it and stable historical mtimes stay cheap
+5 -4
View File
@@ -61,10 +61,11 @@ fi
# snapshot already exists.
# 3. an explicit fallback directory — migration off a pre-existing flat
# cache, so the first run under this scheme isn't a needless cold build.
CANDIDATES=()
[ -n "$BASE_KEY" ] && CANDIDATES+=("$(snapshot_dir_for "$ROOT" "$BASE_KEY"):base snapshot")
CANDIDATES+=("$(snapshot_dir_for "$ROOT" "$OWN_KEY"):own snapshot")
[ -n "$FALLBACK" ] && CANDIDATES+=("${FALLBACK}:fallback dir")
#
# The list itself lives in cache-lib.sh because prune-cache.sh reads it too:
# it runs ahead of this step and has to free enough disk for the clone below,
# which means resolving the same source this loop will pick.
mapfile -t CANDIDATES < <(seed_source_candidates "$ROOT" "$OWN_KEY" "$BASE_KEY" "$FALLBACK")
for entry in "${CANDIDATES[@]}"; do
src="${entry%%:*}"; label="${entry#*:}"
+1 -1
View File
@@ -13,7 +13,7 @@ script_dir=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)
FAST=0
[ "${1:-}" = "--fast" ] && FAST=1
FIXTURE_TESTS=(cache-root-selftest.sh seed-target-dir-selftest.sh publish-snapshot-selftest.sh prune-cache-selftest.sh)
FIXTURE_TESTS=(cache-root-selftest.sh seed-target-dir-selftest.sh publish-snapshot-selftest.sh prune-cache-selftest.sh release-v1-selftest.sh)
CARGO_TESTS=(hardlink-clone-selftest.sh restore-mtimes-selftest.sh)
TESTS=("${FIXTURE_TESTS[@]}")