68 Commits
Author SHA1 Message Date
claude 680ef8049c Merge pull request 'docs: clear PR #28's docs-drift catalog entries' (#31) from chore/docs-catalog-2026-09-23 into main
CI / shellcheck + selftests (push) Successful in 1m29s
CI / move v1 to main (push) Successful in 4s
v1
2026-09-23 03:02:39 +00:00
claudeandClaude Opus 5.5 22dafe45e2 docs: clear PR #28's docs-drift catalog entries
CI / shellcheck + selftests (pull_request) Successful in 1m23s
CI / move v1 to main (pull_request) Skipped
Marks the token grant as verified now that #28's own merge ran
release-tag and moved v1 to that commit (confirmed against the CI
status API and the v1 tag on origin), and rewords the scenario-9
stranding framing to "a run that deferred and left no newer run
behind it" rather than a concurrency-group cancellation ci.yaml no
longer allows.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 21:53:32 -05:00
claude 0284fcd54b Merge pull request 'ci: release v1 automatically on a green merge to main' (#28) from chore/v1-release-gate into main
CI / shellcheck + selftests (push) Successful in 1m45s
CI / move v1 to main (push) Successful in 4s
2026-09-22 23:12:59 +00:00
claudeandClaude Opus 5.5 77cc5917b6 fix(release): refuse to overwrite an unrelated hand-placed v1
CI / shellcheck + selftests (pull_request) Successful in 1m47s
CI / move v1 to main (pull_request) Skipped
release-v1.sh's push_leased() only detected a lost lease after a push
was *rejected* -- but force-with-lease compares the remote ref's raw
value against the caller's expected value, not ancestry. If v1 already
sat on a hand-placed, unrelated commit when push_leased() was first
called (no race, nobody moves it mid-call), the very first push found
the ref exactly where it expected, succeeded outright, and silently
overwrote the unrelated v1 with <sha> -- skipping every ancestry check
the function has, since those only run after a rejection.

Fix: before the first push attempt, check whether the caller's
`expect` is neither an ancestor of `sha` (the ordinary stale-v1 case)
nor already covering it (nothing to do) -- and go red naming both SHAs
if so. `expect` is always a peeled commit (fetch_v1() reads
`refs/tags/v1^{commit}`), so this doesn't add a second failure mode
for an annotated v1; that tag form's existing "not a lost lease"
behavior on the first rejected push is untouched.

Surfaced by PR #28's final review. New selftest scenario 11 in
release-v1-selftest.sh, red-proven against the unfixed script (v1 was
silently moved off the stray commit); green after the fix, with the
full 7-suite gate (shellcheck + selftest.sh) passing.

Ride-alongs from the same review:
- ci.yaml:120-123 claimed the README's Versioning section documented
  what's verified about the release token's write access; it said
  nothing. Added an accurate sentence there (the grant is unobserved
  until the first merge, capped by repo/owner token-permission maxima
  and unreadable tag protections) and pointed the comment at it.
- Deleted two comment-as-decision-history paragraphs per
  comments-are-not-exposition: ci.yaml's "no job-level concurrency"
  rationale (kept one line of intent) and
  prune-cache-selftest.sh scenario 9's account of how an
  existence-only assertion used to pass with the pass-2 guard removed
  (kept a one-line statement of what it checks).
- README's release-v1-selftest.sh table row now names the new
  hand-placed-v1 scenario.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 16:14:41 -05:00
claudeandClaude Opus 5.5 ca0ee132d9 fix(ci): add a scheduled v1 sweep and lease every v1 push
CI / shellcheck + selftests (pull_request) Successful in 1m46s
CI / move v1 to main (pull_request) Skipped
The release guard in 17d87b0 was safe but not live. Gitea 1.27.2 calls
CancelPreviousJobsByJobConcurrency whenever a job's `needs` resolve
(services/actions/clear_tasks.go:91, models/actions/run_job.go:641), so
a job's place in the `release-tag-v1` group followed when its own
selftest finished, not merge order. A newer merge C2 finishing selftest
first queued behind the older C1, C1 cancelled it, saw tip = C2, and
deferred: nobody pushed, and if merges then stopped v1 stayed stale
indefinitely behind a Skipped and a Cancelled job. The "always catches
up once merges pause" claim in ci.yaml and README was false.

What now holds:

- release-sweep.yaml runs on `schedule` every 15 minutes, in its own
  workflow and concurrency group, so nothing in ci.yaml can cancel it.
  When v1 already covers main's tip it stops after a checkout and one
  merge-base. Otherwise it checks out the tip, runs the same shellcheck
  and selftest.sh as ci.yaml's selftest job, and tags the tip only if
  they pass; a failing main therefore turns the sweep red on every tick
  while v1 lags, which is #27's AC1 loud-failure half. It reads the tip
  itself because a scheduled run's github.sha is the CommitSHA recorded
  when the schedule was registered on the last push to main
  (services/actions/notifier_helper.go:569-580,
  services/actions/schedule_tasks.go:126-141), and ref is the default
  branch: schedules are registered only from it
  (notifier_helper.go:120, :531, :603-604). event_name is "schedule"
  (context.go:71 reads TriggerEvent, set at schedule_tasks.go:136).
  Cron is 5-field robfig in UTC (models/actions/schedule_spec.go:38-41).

- 15 minutes, not 10: the sweep is the fallback, not the release path,
  and every tick is a run on gitdan-ci's shared slots and a row in the
  Actions list. 96 no-op runs a day of a few seconds each is the cost;
  the lag bound it buys is one interval plus one selftest run.

- Both writers go through scripts/release-v1.sh and push with
  --force-with-lease=refs/tags/v1:<v1 as read>, so v1 cannot move
  backwards when the sweep and a merge job race. A lost lease re-reads
  v1: at or ahead of this run's gated commit is a clean skip (the other
  writer released something at least as new); still behind it is a
  retry leased on the new value, up to three attempts, since the other
  writer may have tagged an older commit and giving up there would leave
  v1 short of a commit this run did gate; anything else goes red. A
  rejection with v1 unmoved is diagnosed as a non-lease failure and goes
  red at once.

- release-tag loses its job-level concurrency group. The lease already
  gives the ordering the group was there for, and the group was what
  cancelled the one job that could have released the newest merge.
  Without it each merge's job runs, and the one whose commit is still
  the tip when it checks releases it.

The shell moves out of ci.yaml into scripts/release-v1.sh so shellcheck
and selftest.sh cover it. release-v1-selftest.sh runs it against a
scratch bare origin: sweep no-op at and ahead of the tip, tag on a
green gate, no tag and a failing sweep on a red one, the stranded trace
above followed by a catching-up sweep, and each lost-lease outcome, with
a control showing an unleased push does step v1 back. Red-proved by
seven mutations of release-v1.sh, each failing a named assertion: plain
--force, accepting any lost lease, a merge job that never defers,
ancestry reduced to equality, no non-lease diagnosis, a sweep that never
needs to run, and a retry that does not re-lease.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 15:01:05 -05:00
claudeandClaude Sonnet 5 17d87b0647 fix(ci): only ever push a commit this run actually gated; fix vacuous scenario-9 guard
## v1 could advance onto an ungated commit

`needs: selftest` gates this run's own commit, but the push targeted
origin/main's freshly-fetched tip with nothing comparing the two.
Trace: M1 merges green; M2 merges while M1's selftest is still
running; M1's release-tag job fetches tip = M2 and pushes v1 = M2,
whose own selftest may be queued, running, or red. If M2 is red, its
own job is skipped, so v1 sits on a red commit across every consuming
project until the next green merge -- with nothing red pointing at
the release itself. ci.yaml:107-109 and README.md:640-641 both
asserted this couldn't happen; ea48c03's own diff established the
precondition (tip "can be minutes stale... behind its own selftest
job") without closing it.

Fix: skip the push unless origin/main's tip IS this run's own
github.sha, checked before the existing v1-monotonicity check
(dc1e631) rather than replacing it -- the two compose (tip-mismatch
first, since it's the coarser reason to defer; ancestor-check second,
for a duplicate run whose own commit is still current). Reverted the
push target from the fetched tip back to `${{ github.sha }}`, now
that the guard makes them provably equal whenever the push fires.

## Convergence trace: does the newest commit's job still run?

Read gitea's source further at the pinned v1.27.2 tag:
PrepareToStartJobWithConcurrency (services/actions/clear_tasks.go)
calls CancelPreviousJobsByJobConcurrency on every job entering the
group, unconditionally cancelling whatever was previously
Waiting/Blocked there -- so at most one job sits queued in the group
at a time; each new arrival supersedes it. Because job-level
concurrency is only evaluated once `needs: selftest` is satisfied
(job_emitter.go re-evaluates readiness there), "arrival order" tracks
each commit's own selftest-completion time, not raw merge order -- an
older commit with a slower selftest can enter the group after a
newer one and cancel its queued slot.

That cancelled job is gone for good; it will never push. But the
commit that's genuinely current at the moment merges stop arriving is
always the one still queued when the running job finishes, because
every subsequent arrival (from every subsequent merge, not just the
"newest" one at any single instant) keeps re-superseding the queue.
So v1 always eventually catches up -- "one merge later" in the common
case, "at the next merge, whenever that happens" in the adversarial
case where a stale survivor runs, finds itself no longer current,
defers, and nothing is left queued. It cannot get stuck forever short
of the repository never receiving another merge, because every future
push re-attempts the same check against whatever's current by then.
Stated this plainly in the comment and README rather than repeating
the false "never" guarantee in softer words.

Re-derived the truth table against the new guard in a scratch
origin+clone, six cases: own commit == tip, no v1 (push); v1 already
== own commit (skip, duplicate run); tip moved past own gated commit
because a newer merge landed (skip, defers); the newer commit's own
run once nothing further has landed (push); v1 already ahead of a
now-stale gated commit (skip, tip-mismatch catches it first); tip ==
own commit but v1 independently ahead via a local-only descendant,
isolating the second (ancestor) check on its own (skip). All six
resolved as intended.

## Scenario 9's pass-3 guard was vacuous

scripts/prune-cache-selftest.sh:230 grepped 'self-clear' against
$scratch/log. summary_line() (cache-lib.sh) writes only to
$GITHUB_STEP_SUMMARY; the self-clear branch (prune-cache.sh:632)
writes 'self-clear' there and 'clearing own' to stdout (:631) --
'self-clear' never appears in $scratch/log at all, so the branch was
unreachable and the `ok` unconditional. Scenario 8 already greps the
right string against the right file (assert_log "clearing own" ...);
scenario 9 now does the same, staying on $scratch/log where it
already was -- the string was wrong, not the file.

Red-proved by capturing a real prune-cache.sh log where self-clear
genuinely fired (scenario 9's own fixture with MIN_FREE_PCT
temporarily raised to 100, in a scratch copy, reverted after) and
running both patterns against it: `grep -q 'self-clear'` -> no match
(the old check's vacuous pass, confirmed); `grep -q 'clearing own'`
-> match (the fix's correct fail). Matches the reviewer's own
measurement exactly. The real prune-cache-selftest.sh was untouched
during this experiment; only the grep string changed in the actual
commit.

bash scripts/selftest.sh: all 6 suites green (67 prune-cache
assertions, unchanged in count -- the fix corrects what scenario 9's
existing check compares, not what it asserts). shellcheck -x
--source-path=scripts scripts/*.sh: clean, as before not covering the
inline ci.yaml shell.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 14:01:28 -05:00
claudeandClaude Sonnet 5 ea48c03aeb fix(ci): target origin/main's live tip, not the run's own trigger commit
CI / shellcheck + selftests (pull_request) Successful in 1m35s
CI / move v1 to main (pull_request) Skipped
The wake-one-cancel-the-rest mechanism behind the concurrency group
(CancelPreviousJobsByJobConcurrency, models/actions/run_job.go, at
the v1.27.2 tag this instance runs) picks its survivor from
models/actions/run_job_list.go's query, which carries no `ORDER BY`
-- so under three-way contention on a shared 2-slot runner, the
*newest* commit's job can be the one cancelled while an older sibling
survives and, correctly from its own vantage, advances v1 forward
from a stale view. No push regresses v1 (the ancestor check from
dc1e631 already prevented that), but the newest merge goes silently
unreleased behind a cancelled job that reads as benign, not red --
exactly the failure #27 exists to end.

Every job now resolves `origin/main`'s tip fresh, right before the
push, instead of using `${{ github.sha }}`. Re-fetched explicitly
rather than trusted from the checkout step, which can be minutes
stale by this point behind its own selftest job. Every execution that
reaches the push step now converges on the same target regardless of
which job the concurrency group lets through, so which one wins the
wake no longer matters -- the survivor pushes where any of them
would have.

That doesn't make the push safe on its own: two jobs can still read
main at genuinely different moments if it advances between their two
fetches, so whichever read the tip earlier must not overwrite the
other's already-pushed, newer one. The ancestor check from dc1e631 is
kept for exactly this -- its target changed (origin/main's live tip,
not this job's own trigger commit) but its job didn't.

What each guard now protects against, after this change:
- concurrency group: stops two jobs from pushing at the same time --
  wasted work now that a cancelled job costs nothing, not a
  correctness backstop by itself.
- ancestor check: stops a job whose own fetch of the tip is stale
  relative to another job's already-pushed, fresher one from
  regressing v1.

Rewrote both the job-level comments and README's Versioning section,
which described "force-moves v1 to that commit" and said nothing
about the concurrency group or the skip-as-success case.

Re-derived the truth table against the new target (a live origin/main
tip, not a fixed commit) in a scratch origin+clone: no v1 yet (push),
v1 exactly at the tip (skip), the tip moved past v1 because a newer
merge landed (push, to the new tip -- not stuck on any prior commit),
v1 already ahead of the tip (skip, defensive), unrelated histories
(push, defensive). All five resolved as intended; ran the exact
condition and fetch sequence the workflow step uses, not a
simulation of it.

bash scripts/selftest.sh: all 6 suites green (67 prune-cache
assertions, unchanged). shellcheck -x --source-path=scripts
scripts/*.sh: clean -- as before, this does not cover the inline
`run:` shell in ci.yaml.

What remains unverifiable short of a real merge is unchanged from
dc1e631: the workflow step's actual execution inside a real Actions
run, and whether the built-in token has write access at all. Neither
this commit nor the one before it can exercise those.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 12:49:05 -05:00
claudeandClaude Sonnet 5 dc1e6317c6 fix(ci): make the v1 push monotonic, not just mutually exclusive
The concurrency group added in af1233f only excludes two release-tag
jobs that are simultaneously Running/Waiting/Blocked
(services/actions/clear_tasks.go:64-80,
models/actions/run_job.go:641-660 at the v1.27.2 tag this instance
runs) -- it has no notion of commit order between jobs that never
overlap. On a 2-slot runner shared across 4 repos, with a
multi-minute selftest gating each release-tag job, two merges landing
close together routinely finish in the opposite order from the pushes
that triggered them: if the newer commit's job completes and exits
first, the older commit's job later finds no live holder in the
group, is not blocked, and force-pushes v1 backward to itself. The
concurrency comment's "never regress it to an older one" and af1233f's
commit message both asserted the opposite -- true of the simultaneous
case the guard covers, false of the staggered one it doesn't, so
authored-false rather than drift.

Added a merge-base check before the push: skip if v1 already points
at this commit or a descendant of it (`--is-ancestor` treats a commit
as its own ancestor, so "at" and "ahead" are the same branch). A v1
that doesn't exist yet, or shares no history with this commit, falls
through to the push -- the guard only ever skips, never fails. Needs
`fetch-depth: 0` on the checkout: actions/checkout's own description
for that value is "all history for all branches and tags", and its
source (dist/index.js: fetchDepth <= 0 selects
getRefSpecForAllHistory, which includes the tags refspec) confirms
tags are fetched as part of that, not gated behind the separate
fetch-tags input -- so refs/tags/v1 and the history behind it are both
guaranteed present locally without a second fetch step.

Rewrote both false claims: the concurrency comment now says what the
group actually bounds (simultaneous competing pushes, not completion
order), and states plainly that the guard below is what makes the
outcome order-independent.

Verified the check's five cases (no tag yet, tag ahead of this
commit, tag behind this commit, tag equal to this commit, unrelated
history) against a real scratch git repo, extracting the exact
condition used in the workflow step -- all five resolved as intended
(skip only when v1 is already at or ahead). The workflow step itself
cannot be exercised outside a real Actions run.

bash scripts/selftest.sh: all 6 suites green (67 prune-cache
assertions, unchanged). shellcheck -x --source-path=scripts
scripts/*.sh: clean -- note this does not cover the inline `run:`
shell in ci.yaml, which shellcheck was never wired to check in this
repo (verified against the CI job itself, which shellchecks only
scripts/*.sh).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 12:24:27 -05:00
claudeandClaude Sonnet 5 af1233f14a fix(ci): serialise release-tag across concurrent merges
The workflow-level concurrency group is keyed per-commit (github.sha,
ci.yaml:19-28) so unrelated pushes never block each other -- but that
also means two merges landing close together run two concurrent
release-tag jobs, each force-pushing its own commit to v1. If the
older commit's job finishes last, v1 regresses to a stale-but-green
commit and stays there until the next merge corrects it forward.
Bounded blast radius (never a red commit, self-heals on the next
merge), but a silently wrong v1 is the exact failure #27 exists to
end.

Added a job-level `concurrency:` on release-tag with a fixed group
name and `cancel-in-progress: false`. Confirmed this is additive to
the workflow-level group, not a replacement, by reading gitea's source
at the v1.27.2 tag this instance runs (`tea api version`): run-level
and job-level concurrency are separate model fields
(ActionRunAttempt.ConcurrencyGroup vs ActionRunJob.ConcurrencyGroup),
evaluated by separate functions (EvaluateRunConcurrencyFillModel vs
EvaluateJobConcurrencyFillModel, services/actions/concurrency.go) and
enforced by separate cancellation paths (CancelPreviousJobsByRunConcurrency
in models/actions/run.go:364 vs CancelPreviousJobsByJobConcurrency in
models/actions/run_job.go:641) -- job_emitter.go's checkRunConcurrency
checks both groups independently (services/actions/job_emitter.go:210-236).
A fixed, non-sha group name is what serialises release-tag across
commits without touching the per-sha grouping every other job still
relies on. `cancel-in-progress: false` matters because Gitea's default
queue behaviour (documented same as GitHub's: a new queued job
supersedes an older *queued* one in the same group, but not a
*running* one) already gives the newest commit's push priority once
the running job clears; cancelling the running job too would abandon
whichever merge is currently pushing mid-flight, the same defect
wearing different clothes.

Re-verified: YAML parses, shellcheck clean, `bash scripts/selftest.sh`
all 6 suites green (67 prune-cache assertions, unchanged by this
commit).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 12:07:01 -05:00
claudeandClaude Sonnet 5 21b444121d test(prune): red-prove $OWN_DIR protection under genuine disk pressure
Scenario 9 asserted only that target-$OWN existed, checked right after
a no-pressure run (3b) where nothing was ever a candidate for
eviction, and prune-cache.sh's self-clear step unconditionally
recreates an empty $OWN_DIR whenever the run ends under the percentage
floor regardless of what pass 2 did to it. Either way the existence
check passed whether or not the pass-2 guard (protected_reason,
prune-cache.sh:221) actually protected the directory. Moving that
guard's $OWN_DIR check into the pass-1-only predicate — the exact
mutation gitdan-actions#26 describes — left all 64 assertions green,
confirmed here before the fix.

Rewritten to run under a real, shrinking `df` (the scenario-17
pattern: a fake df that re-measures the fixture with `du` on every
call, so eviction genuinely lowers the reported pressure), with
MIN_FREE_PCT=0 and a clone-headroom floor sized so self-clear's
percentage check can never fire — only pass 2's guard decides the
outcome. The fixture gives $OWN_DIR real content and an older
timestamp than a sibling target dir, sized so the requirement is met
by evicting exactly one of them. With the guard removed, $OWN_DIR is
the one evicted (LRU-oldest, and self-clear is structurally disabled
by MIN_FREE_PCT=0 so there is nothing left to recreate it) — the
scenario now fails loudly on the same mutation that left it green
before. Restored and reverified green (67 assertions, up from 64) with
the guard intact.

bash scripts/selftest.sh: all 6 suites green. shellcheck -x
--source-path=scripts scripts/*.sh: clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 11:52:01 -05:00
claudeandClaude Sonnet 5 7f18cb2436 feat(ci): advance v1 automatically once the gate is green on main
Nothing moved v1 when main advanced, so a merged change was inert
until someone remembered to retag by hand — it happened on 2026-09-22
(PR #25 merged, v1 stayed on the previous release) and was only caught
because a person asked whether the tag had moved.

Option 1 from gitdan-actions#27 (automate it) over option 2 (fail loud
while it lags): a `release-tag` job, gated with `needs: selftest` so a
broken build never reaches it, force-moves v1 to the pushed commit
using the run's built-in GITHUB_TOKEN. If that token turns out not to
have write access, the push fails and the job goes red in the Actions
UI — a loud failure either way, not the silent one this replaces.
Whether the token actually has write access here is unverified short
of a real merge; that merge is the next step for this branch.

README's Versioning section documented the old manual step as a
deliberate decision "never something a merge does by itself" — that
claim is now false, so it's rewritten to describe the automated job
and keeps the manual command as the recovery path for when the job
can't push.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 11:51:51 -05:00
claude 21dffdb725 Merge pull request 'fix(prune): stop protecting publisher target dirs under disk pressure' (#25) from fix/prune-unprotect-target into main
CI / shellcheck + selftests (push) Successful in 1m49s
2026-09-22 16:07:52 +00:00
claudeandClaude Sonnet 5 dc473f0d0c fix(prune): stop protecting publisher target dirs under disk pressure
CI / shellcheck + selftests (pull_request) Successful in 1m48s
A publisher branch's target-<ref> was protected identically to its
snapshot-<ref>, so it was never a pressure-pass candidate however old
and however tight the disk. With one branch's target dir permanently
resident alongside its snapshot, a second branch had no room to seed,
and every PR that night needed a hand eviction between runs
(daniel/gitdan-actions#24).

A publisher's target dir is a convenience cache its own next run
reseeds from the snapshot, so losing it under pressure is cheap;
nothing downstream depends on it surviving. Only the snapshot stays
protected in the pressure and self-clear passes.

Unprotecting the target dir outright surfaced a second bug the fix
would otherwise have shipped: a branch's tip is trivially an ancestor
of itself, so once a protected ref's target dir was no longer skipped
before reaching the merged-branch check, pass 1 read it as "merged
into itself" and deleted it unconditionally on every run, independent
of disk pressure. is_protected_from_liveness keeps a protected ref's
target dir out of pass 1 alone, so it stays an ordinary pressure-pass
candidate without ever reaching that check. Scenario 3b in the
selftest red-proves this against the unprotect-only version of the
fix.

Also corrects the header's inode-sharing claim, measured false on the
live volume by daniel/zemyna#1073: publish-snapshot.sh unshares every
executable after its cp -al, and executables are most of the tree by
bytes, so a snapshot eviction is a real, large disk cost rather than
the near-free one the old text described — the target-before-snapshot
ordering still holds, now for the warm-start reason alone plus that
cost.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 10:42:08 -05:00
claude 0184df25a2 Merge pull request 'fix(prune): reclaim merged branches and size the volume for the clone' (#21) from fix/prune-merged-live-and-seed-headroom into main
CI / shellcheck + selftests (push) Successful in 1m24s
2026-09-07 05:41:12 +00:00
claudeandClaude Fable 5.1 38a6387936 docs(cache): recommend delete-on-merge, and say what ancestry still covers
CI / shellcheck + selftests (pull_request) Successful in 1m27s
`daniel/zemyna` enabled `default_delete_branch_after_merge` after this
branch was written, so the deleted-branch signal will fire there on
future merges. That makes the setting worth recommending — it is the
cheapest case for this scheme, decidable from `ls-remote` with no
checkout, no objects and no walk — and it does not make the ancestry
signal redundant.

Three things it leaves behind, now enumerated in the README's eviction
section rather than implied: every branch merged before the setting was
turned on, of which zemyna carried 48 and which nothing retroactively
deletes; every merge whose deletion the forge declines or is never asked
to make, since it is best-effort and silent and an API merge without the
flag never asks; and every repo that has not enabled it, which is the
default.

Two present-tense claims about one repo's configuration are reworded
into the conditions they were standing in for, in prune-cache.sh's
header and beside is_merged_dead, plus the two in the selftest that
asserted the forge keeps branches rather than describing the fixture.
No behaviour change; the suite is green and unchanged at 61 assertions.

Refs #20.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXMQCJ5Eg5f9G9cfYzyh4Z
2026-09-07 00:27:11 -05:00
claudeandClaude Fable 5.1 24f87a6b98 docs(cache): describe both liveness signals and the derived requirement
CI / shellcheck + selftests (pull_request) Skipped
The eviction section described one way for a branch to be dead and a
percentage threshold that no longer exists as a gate. Rewritten around
what the pass actually does: two dead-branch signals with the limits of
the ancestry one stated, a requirement measured off the clone's mutable
set, and a failure that names its shortfall.

The 10% figure is deleted rather than corrected — it was the default of
a gate, and min-free-percent is now an additional floor defaulting to 0,
so there is no percentage left to state. The inputs table, the prune and
liveness rows, the selftest coverage row and the fetch-depth comment in
the usage example all say what they now mean; fetch-depth: 0 has a
second reason to be required, since a shallow checkout cannot answer
ancestry.

Refs #20.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXMQCJ5Eg5f9G9cfYzyh4Z
2026-09-07 00:21:12 -05:00
claudeandClaude Fable 5.1 07ba53ca79 fix(prune): reclaim merged branches, and size the volume for the clone
Two assumptions in the eviction pass did not hold on this forge, and
between them a volume filled up three times in three days with nothing
reclaimed automatically. Both are replaced here; the pass also moves
ahead of the seed, which is the only order in which its work can help
the run performing it.

LIVENESS. Pass 1 evicted a cache only when its branch was gone from
origin. Gitea keeps a PR's branch after the merge unless the repo opts
into delete-on-merge, and zemyna does not — so ls-remote reports fifty
merged branches and the signal fires for none of them. A second signal
is added beside it: a branch still on origin whose tip is an ancestor of
a protected branch's tip holds no commit that branch does not, so its
cache will never be read again and goes in the same unconditional pass.

Ancestry is answered from the commits in the job's own checkout, so the
answer "cannot tell" exists and stays distinct from "not merged" at both
granularities. A shallow checkout withholds the signal entirely, since a
missing object is its normal case rather than evidence. A single branch
whose tip is not in the checkout is kept, with a warning naming it. A
squash or rebase merge leaves no ancestry and reads as live until the
branch is deleted. All three are missed reclamations, which cost disk;
the other direction costs a branch its cache mid-build.

HEADROOM. Passes 2 and 3 gated on a percentage of the volume, which
cannot express the failure they have to prevent: a clone runs out of
disk while unsharing its mutable paths, and how much that needs is a
property of the snapshot rather than of the disk. Staging failed at 34 G
free and passed at 74 G, so a 10% floor — 19 G here — never fired first.
The requirement is now measured per run off the very source the seed
will read, the pass evicts oldest-first until it is met and stops there,
and falling short of it fails with the shortfall and every directory it
kept, rather than letting the seed fail seconds later against a staging
path that names none of that. min-free-percent survives as an additional
floor, defaulting to 0, and falling short of that one is still a warning
and a self-clear.

ORDERING. The prune step ran after the seed, so each run freed space for
the next one. It now runs between resolve and seed. Two things that
makes newly reachable are closed: the source about to be cloned is
excluded from every pass by name, and a concurrent job's target dir
already carries its lock from the instant it appears under its final
name, so nothing is seen unlocked that is in use.

Red-proven: sixteen assertions across the four new scenarios fail
against the pre-fix scripts, including the zemyna layout evicting
nothing where it should evict exactly one directory, and the headroom
scenario exiting 0 where it should exit 1.

Refs #20.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXMQCJ5Eg5f9G9cfYzyh4Z
2026-09-07 00:21:12 -05:00
claudeandClaude Fable 5.1 a960c8f91b refactor(cache): name the mutable set once, and measure it
The set of paths a hardlink clone has to real-copy — dep-info,
build-script metadata, linked outputs, the pruned directories that hold
them — was spelled out inline in unshare_mutable_paths, in four find
invocations. Nothing else needed it, so one spelling was enough.

Something else needs it now: the prune has to know what a clone will
cost before it happens, and a sizer with its own copy of the predicates
would drift from the copier silently and in the dangerous direction — an
under-measured clone is one that starts and runs out of disk halfway
through unsharing. So the directory names become one array and the file
rules one dispatcher, applied by a callback per side, with each rule's
rationale moved to the rule rather than left at the old call site.

mutable_set_kb measures that set off a snapshot, skipping the subtrees
already measured whole so nothing is counted twice; clone_headroom_kb
scales it by a hand-written margin and floor for what the measurement
cannot see (cp -al materialising every directory for real, and the
unshare holding one subtree twice at its peak). Both residuals are named
where the function is, in both directions.

seed_source_candidates moves the seed's source-preference list into
cache-lib for the same reason: the prune ahead of it has to resolve the
same source the seed will clone, and two agreeing derivations are one
edit away from disagreeing.

No behaviour change — the copier applies the same rules to the same
tree, verified by the hardlink-clone suite's inode partition in both
directions.

Refs #20.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXMQCJ5Eg5f9G9cfYzyh4Z
2026-09-07 00:20:45 -05:00
claude 0b099ecf62 Merge pull request 'fix(ci): trigger on edited so un-drafting actually lifts the draft skip' (#18) from fix/ci-edited-trigger into main
CI / shellcheck + selftests (push) Successful in 1m24s
2026-09-02 18:54:43 +00:00
claudeandClaude Opus 5 5abd0a9968 docs(ci): name the timing fields this forge actually returns
CI / shellcheck + selftests (pull_request) Successful in 1m30s
The accepted-cost derivation cited `run_started_at` and `updated_at`, which are
GitHub's field names. This Gitea's runs payload has neither -- it returns
`started_at` and `completed_at`, and omits `run_started_at`, `updated_at` and
`created_at` entirely (confirmed by dumping the keys of a run object). The
figures are unaffected: the script that produced them fell through to the real
fields, so 83 s median over 77-106 s stands.

The derivation was named so a reader could re-take the measurement, and as
written it returned nothing when followed literally, which defeats the only
reason it was there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LjbhSqQf3pwnPA6MVaWcWL
2026-09-02 13:03:45 -05:00
claudeandClaude Opus 5 8217f53d4a docs(ci): measure the accepted cost instead of comparing it unmeasured
CI / shellcheck + selftests (pull_request) Successful in 1m21s
The accepted-cost paragraph claimed this repo's job made the `edited` trade
worse than in the sibling repos that shipped it first, reasoning from the
toolchain install, the Cargo-driving suites and `timeout-minutes: 20`. The
measurement inverts it: last twelve non-skipped runs here are 83 s median
(77-106), against ~118 s for daniel/gitdan and ~330 s for daniel/emowheel --
this is the cheapest of the three, and 20 minutes is a hang ceiling, not a
duration. This PR's own runs measured 87 s and 90 s.

The comparison is dropped rather than re-pointed; the absolute figure replaces
it, with the derivation named (`run_started_at` to `updated_at` off the Actions
API) so a reader can re-take it. The neighbouring concurrency claim was read
off this repo's own file and is unchanged -- it was the checked half of a
paragraph whose other half was not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LjbhSqQf3pwnPA6MVaWcWL
2026-09-02 12:55:24 -05:00
claudeandClaude Opus 5 5b6acd7b3b docs(ci): date the cancellation observation to 1.26.0, not the current version
CI / shellcheck + selftests (pull_request) Successful in 1m30s
The concurrency note said "this Gitea (1.26.0)" while `GET /version` now
returns 1.27.2 — the instance was upgraded on 2026-08-25/26 (daniel/gitdan's
runbook, gitdan#46). Swapping the number would have asserted the cancellation
was observed on 1.27.2, which nobody has checked: the runs it cites were seen
before the upgrade. The comment now dates the observation to 1.26.0, names the
current version, and says persistence is unverified — which is why every commit
gets its own group rather than trusting `cancel-in-progress`. That reasoning is
unchanged; only the implied currency was wrong.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LjbhSqQf3pwnPA6MVaWcWL
2026-09-02 12:49:56 -05:00
claudeandClaude Opus 5 b791896e03 fix(ci): trigger on edited so un-drafting actually lifts the draft skip
CI / shellcheck + selftests (pull_request) Successful in 1m27s
`ready_for_review` does not exist as a pull_request action on this Gitea, so
it never fired and the draft skip never lifted: a PR opened as `WIP:` carried
its skip decision to merge unless a later push happened to create a run. Draft
here is not a persisted column — it is the `WIP:` title prefix, derived by
`issue.IsWorkInProgress` — so un-drafting is a title edit, which fires a plain
`edited`.

Ported from daniel/emowheel commit 08da820 (see daniel/emowheel#71) and
daniel/gitdan (see daniel/gitdan#92), where the identical change landed first.

Accepted cost: Gitea populates no `changes` field for a title-or-body edit, so
the workflow cannot tell an un-drafting edit from an ordinary body PATCH, and
every body edit on a non-draft PR now starts a real run. That is a heavier
trade here than in the sibling repos -- this job installs two Rust toolchains
and drives a real Cargo, at `timeout-minutes: 20` -- but the `concurrency:`
block groups `pull_request` runs on `github.ref` with `cancel-in-progress:
true`, so a burst of edits collapses to one run. Still-draft PRs are unchanged.

Docs: the workflow comment is rewritten and shortened, and README's Development
section retires the empty-commit workaround it prescribed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LjbhSqQf3pwnPA6MVaWcWL
2026-09-02 12:34:24 -05:00
claude 14bea98626 Merge pull request 'fix(hardlink): name the mutable set directly, under both build-dir layouts' (#16) from fix/layout-v2-selection into main
CI / shellcheck + selftests (push) Successful in 1m46s
2026-08-27 20:03:40 +00:00
claude 4bb880b7a7 chore(ci): trigger CI after un-WIP
CI / shellcheck + selftests (pull_request) Successful in 1m21s
2026-08-27 14:52:11 -05:00
claude fb3c72aa28 docs(hardlink): withdraw the nlink claim, which asserted more than was measured
CI / shellcheck + selftests (pull_request) Skipped
The comments said hardlink count at link time was ruled out as the
discriminator between a rewritten executable and an intact one, citing a
lib+bin crate whose two test binaries were both `nlink == 1` and appeared to
behave differently. Re-checked on review: the intact one had not been rebuilt
at all — same content, same inode — so it demonstrated nothing, and forcing
both to rebuild rewrote both.

What was actually observed is narrower and now says so: every executable
measured intact had an uplift hardlink twin Cargo must re-create anyway, every
one measured rewritten had none, and whether the twin is the mechanism or a
correlate was not determined. The rule does not rest on the answer — exempting
twinned executables would recover none of the bytes this change newly copies.

This PR exists because a claim outlived its evidence; it should not ship one.
2026-08-27 14:50:53 -05:00
claude 0553a6b956 fix(hardlink): resolve an ambiguous out/ toward unsharing, and pin the linked-output scenario on a shape that exhibits it
CI / shellcheck + selftests (pull_request) Skipped
Two review findings on #16.

The `out/` discriminator keyed on Cargo's record of a build-script execution,
which Cargo writes only AFTER the script exits successfully. A build script
that populates OUT_DIR and then fails leaves a unit with no record at all, so
its OUT_DIR read as a compile unit's artifact directory and stayed shared —
a regression against the old `-name build` selection, which real-copied that
state by construction. Reproduced on cargo 1.93.1 stable.

An `out` directory now stays shared only when two independent signals agree:
it holds an `.rlib`/`.rmeta` of its own, and its unit carries no execution
record. Either one missing real-copies it. The cost is unchanged to the byte —
the newly-unshared directories hold only executables and `*.d`, both already
privately owned by the file rules.

The live linked-test-binary scenario was built on the lib+bin probe crate,
whose test binaries relink to a fresh inode — a shape gitdan-actions#17 records
as measured safe. Both halves passed green against the unfixed selection on
dep-info mutations the previous scenario already covers. It now builds a
bin-only crate with a unit test, reads only executables, and skips loudly with
a warning rather than passing quietly if the toolchain does not exhibit the
rewrite at all.

Also: drop a clause asserting the linker writes in place "whenever the path has
no other hard link", which this change's own evidence denies; move the
load-bearing comment block back above `unshare_mutable_paths`; correct a
superseded 99.998% figure; and record both cost rows in the README rather than
only the flattering whole-tree one.
2026-08-27 14:39:29 -05:00
claude e5b26a9368 fix(hardlink): name the mutable set directly, under both build-dir layouts
CI / shellcheck + selftests (pull_request) Skipped
`unshare_mutable_paths` selected `.fingerprint` and `build` directories. Under
Cargo's build-dir layout v2 the first clause matches nothing and the second
matches the whole tree, because v2 regroups artifacts under `build/` alongside
the metadata. Measured on one scratch crate: 39.3% of the tree real-copied
under v1, 99.996% under v2.

The selection now names the mutable set rather than the container it used to
live in: fingerprint directories under either spelling, layout v2's `run/`
directories, layout v1's loose build-script run metadata, and the `out`
directories that are a build script's OUT_DIR rather than a compile unit's
artifact directory. The two are told apart structurally, by Cargo's record of
the build-script execution sitting beside the OUT_DIR and nowhere else.

Verifying that turned up a second, layout-independent hazard: a linked
executable is written through whatever inode is already at its path, so a
`cargo test --no-run` inside a `cp -al` clone rewrites the source's own test
binary. Reproduced on cargo 1.93.1 stable, 1.96.0-nightly, 1.98.0-nightly and
1.100.0-nightly, under both layouts. Every executable is now real-copied;
`.rlib`, `.rmeta` and `incremental/` are what stay shared.

hardlink-clone-selftest.sh gains two file-only layout fixtures that pin the
partition in both directions without a compiler, and a live scenario that
relinks a test binary.
2026-08-27 14:12:01 -05:00
claude 4114996954 Merge pull request 'fix(hardlink): content freshness moved switches, it was not withdrawn' (#15) from chore/checksum-freshness into main
CI / shellcheck + selftests (push) Successful in 1m19s
2026-08-27 03:20:03 +00:00
claude fa3cef53e0 docs(ci): the nightly does enable content freshness, as of 2026-08-26
CI / shellcheck + selftests (pull_request) Successful in 1m36s
The README's CI section still described the nightly toolchain step as
buying nothing: "which no nightly currently enables, so it is skipped and
the step is kept only against the day upstream restores it". That
sentence predates this branch and states as fact the exact reading the
rest of the PR retracts.

Three things in this PR falsify it. The corrected `env:` block at
README.md:110-117 records that since cargo PR #17382 (2026-08-22) the
`-Z` gate only unlocks the feature and `build.fingerprint` selects it, so
setting both turns it on. The corrected workflow comment in
.gitea/workflows/ci.yaml says "as of 2026-08-26 it does". And this
branch's own CI run printed `=== checksum-freshness mode: on ===` and
`hardlink-clone-selftest: 4 assertions passed` on 1.100.0-nightly
(787af2b8c 2026-08-25) — the scenario is not skipped, it runs.

The paragraph also carried no date, which is the failure mode every other
block this PR touched was rewritten to prevent. The replacement is dated
and names the toolchain, matching the corrected blocks elsewhere.

It deliberately stops short of "the scenario always runs": the suite
still settles the question by experiment on every run and still skips
loudly when it cannot measure, so the step is not unconditionally
exercised. Saying otherwise would trade one overclaim for its mirror.

Docs-only; no behaviour change.
2026-08-26 20:02:02 -05:00
claude 554310186f fix(hardlink): content freshness moved switches, it was not withdrawn
CI / shellcheck + selftests (pull_request) Successful in 1m17s
`unshare_mutable_paths`' comment recorded that upstream had stopped
rewriting `dep-<target>` in place, on a measurement taken against
1.100.0-nightly. It had not. Two unrelated cargo changes landed within
four days of each other and between them moved the switch that turns the
behaviour on and the path it writes to:

  - cargo PR #17382 (2026-08-22) demoted `-Z checksum-freshness` to a
    gate and gave `build.fingerprint` the choice, defaulting to `mtime`.
    Setting only the gate is accepted and does nothing, which is exactly
    the result that was read as a withdrawal.
  - build-dir layout v2 (cargo PR #17354, stable 1.100.0 on 2026-11-12,
    nightly default since 1.99) moved the file from
    `.fingerprint/<unit>/dep-*` to `build/<pkg>/<hash>/fingerprint/dep-*`.

Measured 2026-08-26 on 1.100.0-nightly (e8cb624d5): same toolchain, same
clone procedure, one env var apart — with the gate alone a `cp -al` clone
mutates only the build/ and *.d families; add
`CARGO_BUILD_FINGERPRINT=content` and the source's dep-info file is
mutated through the shared inode again. The hazard is intact.

So the suite now exports both switches, and its strongest scenario runs
again on current nightlies — verified passing against both layouts. Its
control note learned the v2 path too: it looked for the v1 path only, and
so printed "does NOT rewrite ... in place" three lines beneath a listing
that showed the rewrite.

Every claim these comments make is now dated and cited, because the
defect being fixed is a comment that cited one measurement and silently
stopped reproducing.

Layout v2 also drags the artifacts under `build/`, which collapses this
function's real-copy set from 21.3% of a target dir to 99.998%. That is a
live cost, not a correctness problem, and it is filed as gitdan-actions#14
rather than fixed here.

Part of daniel/gitdan#62.
2026-08-26 19:48:07 -05:00
daniel 50e430f4de Merge pull request 'fix(cache): separate build directories for same-ref jobs, and a CI gate' (#13) from fix/lock-contention into main
CI / shellcheck + selftests (push) Successful in 1m22s
Reviewed-on: #13
Reviewed-by: claude-reviewer <3113+claude-reviewer@noreply.gitdan.com>
2026-08-26 21:38:35 +00:00
claude e94c46f2ec test(hardlink): move the probe's INACTIVE answer off bash's error codes
CI / shellcheck + selftests (pull_request) Successful in 1m21s
Review finding on #13, and the last change to this function. A typo'd variable
name on a line that HAS its `|| exit 2` guard — `$CONTNET_B` for `$CONTENT_B`
— fails in EXPANSION, before the command runs, so neither the guard nor the
ERR trap ever sees it. Under `set -u` that exits 1, which was the code meaning
"measured, and the answer is mtime". Not a live defect: the probe references
only set variables today, and the 1 gitdan-ci reports is a genuine measurement.

The defect is that the answer codes and the error codes overlapped at all.
Three rounds on this function each found a narrower way for a shell-generated
1 to be read as a measurement — an unguarded command, then a guarded line
whose guard could not fire — and each was closed by narrowing the failure
surface, which is a game with no last move.

INACTIVE is now 3, and anything that is not 0 or 3 is NOT MEASURED. Bash
generates 1, 2, 126, 127 and 128+n for its own errors and never 3, so "not an
answer" is decided by a property of the shell rather than by an enumeration of
the ways a step can go wrong. A step added later without a guard, or with a
guard that cannot fire, lands on NOT MEASURED by construction.

The guards and the trap stay, with their job restated: reporting rather than
correctness. They put a failed step on 2 with both probe logs printed instead
of on some incidental status — nicer to debug, and the same destination either
way. The `set +e` bracket at the call site keeps its own note, since a `set -e`
that cannot fire is the kind of thing a reader assumes works.

The not-measured message now quotes the actual exit code, so the two ways to
reach it are distinguishable in a log without reading this file.

Red-proven with the reviewer's own mutation, A/B: with INACTIVE at 1 the typo
reports `mode: off — resolves freshness by mtime`, a false measurement; at 3 it
reports `mode: unmeasured — the probe exited 1`. All four paths re-verified —
unmodified 4 assertions; env stripped inside the probe, measured INACTIVE with
the mtime reason; a broken probe build, unmeasured via a guard; an unguarded
`false`, unmeasured via the trap.
2026-08-26 14:40:27 -05:00
claude 5b46986cd7 test(hardlink): make the errexit invariant real, and stop gpg deciding a gate
CI / shellcheck + selftests (pull_request) Successful in 1m23s
Two review findings on #13.

THE STATED INVARIANT WAS NOT IN FORCE. The header claimed a `set -e` abort
could never be mistaken for the `1` that means "measured, and the answer is
mtime". It could not fire at all: `checksum_freshness_probe || probe_rc=$?`
runs the left side with errexit suppressed, and that suppression propagates
into the subshell. An unguarded step there fell through to `exit 0` and
reported ACTIVE — a fourth outcome the header said was impossible. The guard
audit was complete, so nothing was broken; the wrong mechanism had the credit,
in a comment inviting the next editor to add an unguarded step and rely on it.

The suggested fix — `set -e` as the first line of the subshell — does not work,
and measured on bash 5.3 it makes things worse rather than not-better:

  set -e inside a subshell called via `||` or `if`   still falls through: rc=0
  set -e inside, called with errexit off at the site aborts, but with the
                                                     FAILING COMMAND's status —
                                                     `false` gives 1, which is
                                                     exactly the value that
                                                     means "mtime"

So both halves are needed and neither is decoration: `set +e` around the call
site, so the subshell can arm its own errexit at all, and `trap 'exit 2' ERR`
inside it, so an abort lands on "not measured" instead of on an answer. The
explicit `|| exit 2` guards stay as the first line of defence. All of that is
now written down where the false claim was.

Red-proven with one unguarded `false` between the `touch` and the second
build, on a box where freshness is live: without the fix, mode `on` and 4
assertions — the reviewer's fourth outcome, reproduced; with it, mode
`unmeasured`, the warning, 3 assertions; unmodified, mode `on` and 4.

Recorded in the same block, since the next person to hit a `2` will ask: why it
skips rather than fails. The gated scenario is the only thing here that depends
on freshness mode, and failing would turn a statement about one machine into a
red gate reading "the hardlink scheme is broken" across three consuming repos.
What would change the answer is `2` becoming the everyday CI outcome, and run
2583 measures that it is not — gitdan-ci reports `1`.

GPG. The two suites that commit in scratch repos inherited the developer's
GLOBAL `commit.gpgsign`, so whether this gate passes could depend on their gpg
agent — seen as a red `restore-mtimes-selftest.sh` caused by a full disk
breaking gpg, in a suite with nothing to say about either. Pinned off locally,
in the throwaway repos only: `git config commit.gpgsign false` beside the
identity the scratch repo already sets, and `-c commit.gpgsign=false` on
prune-cache-selftest's three commits, matching its existing `-c` style.
Red-proven under a GIT_CONFIG_GLOBAL with gpgsign on and a nonexistent gpg
program: without it `fatal: failed to write commit object`, with it both
suites pass.
2026-08-26 14:28:22 -05:00
claude 5a90d5c829 test(hardlink): let the freshness probe report that it could not measure
CI / shellcheck + selftests (pull_request) Successful in 1m21s
Review finding on #13. `checksum_freshness_active()` returned non-zero for any
reason — including its own `cargo build` failing — and the caller's `else`
branch then announced "this nightly accepts -Z checksum-freshness but still
resolves freshness by mtime" regardless. On a machine where content freshness
IS live, breaking only the probe's build produced that sentence, which is
false, and silently dropped the scenario, which is lost coverage. Reachable,
not theoretical: it needs a half-installed toolchain, which is exactly the
state a probe exists to notice.

The whole point of the preceding commit was to settle this by experiment
rather than by asking, and an experiment that cannot tell a negative result
from a failed measurement is not settling it. "This toolchain resolves
freshness by mtime" is a statement about Cargo; "a probe build failed" is a
statement about this machine. They are not interchangeable.

Three outcomes now, carried in an exit code:

  0  measured ACTIVE     both builds ran, the backdated rebuild recompiled
  1  measured INACTIVE   both builds ran, the backdated rebuild was Fresh
  2  NOT MEASURED        a probe step failed; nothing was learned

Every step that could fail for a reason other than the experiment's own
outcome exits 2 explicitly, so a `set -e` abort cannot be mistaken for the 1
that means "measured, and the answer is mtime". Outcome 2 emits a ::warning::
saying in those words that this is a failure to measure and not a finding
about Cargo, and prints the tail of both probe logs. The scenario still skips
for 1 and 2 alike — everything else the suite asserts is independent of
freshness mode — but the reason is now carried to the skip line, so the two
are distinguishable at a glance.

Verified all three states on a machine where freshness is live: unmodified,
mode `on` and 4 assertions; probe build broken with a bogus flag, mode
`unmeasured` with the warning and 3 assertions, and no claim about mtime;
CARGO_UNSTABLE_CHECKSUM_FRESHNESS stripped inside the probe — CI's condition —
mode `off` naming mtime, and 3 assertions.
2026-08-26 13:34:42 -05:00
claude eb3f0bd09b docs(ci): say what the nightly step actually buys today
CI / shellcheck + selftests (pull_request) Successful in 1m24s
The step's comment claimed nightly was not optional. CI's own first green run
falsified that: 1.100.0-nightly accepts -Z checksum-freshness and resolves
freshness by mtime regardless, so the suite's probe reports it and skips the
scenario either way.

The step stays — twenty seconds, and the coverage returns by itself the day
upstream restores the behaviour — but the comment now says that rather than
the opposite. The open question about what upstream actually did is
daniel/gitdan#62.
2026-08-26 12:57:15 -05:00
claude 6a80423bd0 test(hardlink): settle checksum freshness by experiment, not by asking
CI / shellcheck + selftests (pull_request) Successful in 1m17s
Third toolchain probe in this branch, and the first one that asks the
question the suite actually depends on. The two before it each let the suite
assert a property the toolchain did not have:

  cargo +nightly -V                    answers "did a proxy called with
                                       +nightly exit 0". `-V` short-circuits
                                       before `-Z` is parsed at all.
  cargo +nightly -Z checksum-freshness answers "is the flag still accepted".
   locate-project                      1.100.0-nightly (2026-08-25) accepts it
                                       and resolves freshness by mtime anyway,
                                       which is how CI reached the final
                                       scenario and failed there.

The final scenario depends on exactly one property: that changed content with
an OLDER mtime rebuilds. Under mtime freshness the correct answer is Fresh, so
under mtime freshness that scenario asserts a bug — which is what CI reported.

So the probe performs that experiment, on its own crate and its own target
dir, with no clone anywhere near it. That separation is also what keeps it a
control rather than a restatement: the probe establishes that the toolchain
rebuilds on content, the scenario establishes that a hardlink clone did not
take that away.

Verified both ways locally: with a real content-freshness nightly, mode on and
4 assertions; with the env var stripped inside the probe — the runner's
condition, faithfully — the probe says so by name, mode goes off, the scenario
is skipped loudly and the remaining 3 assertions pass.
2026-08-26 12:54:24 -05:00
claude 5e8e773f78 test(hardlink): report the mutation families, don't assert one of them
CI / shellcheck + selftests (pull_request) Failing after 1m21s
CI's nightly is 1.100.0-nightly (2026-08-25); the machine this suite was
written on had 1.96.0-nightly (2026-02-24). On the newer one the control's
`cp -al` clone mutates only the build/ and *.d families — upstream appears to
have stopped rewriting `.fingerprint/*/dep-*` in place under checksum
freshness — so the suite went red on the *absence* of a hazard.

That is the wrong shape for a gate. The control's job is to prove the hazard
exists at all, which a non-empty mutated set already does; naming one family
as mandatory makes the suite red whenever upstream stops doing something we
never wanted it to do, and red in the CONTROL, where a failure reads as "the
hazard is gone" rather than "upstream changed". It is now a note either way.

Nothing is given up. The fix scenario asserts the source is byte-identical
after a full rebuild in the clone, which covers every family the running Cargo
has, named or not — and the checksum-freshness scenario after it tests the
stale-reuse hazard directly. The dep-* line only ever documented which family
was in play.

`unshare_mutable_paths` keeps unsharing dep-* regardless, and its measurement
block now records both observations with their versions: 22 MB of a 6.9 GB
tree against a failure mode that is a wrong answer rather than a slow one.
2026-08-26 12:50:42 -05:00
claude 3f97d3d1e7 fix(selftest): make the two compiler-backed suites survive a CI runner
CI / shellcheck + selftests (pull_request) Failing after 1m36s
The new gate's first run went red on both suites that drive a real Cargo.
Neither failure was a defect in what they test; both were assumptions about
the machine, which had only ever been a dev box with a nightly installed and
no colour forcing. Fixed here rather than waived — turning a gate on is what
obliges fixing what it finds.

COLOUR. Every assertion in restore-mtimes-selftest.sh, and one in
hardlink-clone-selftest.sh, reads cargo's own words out of a build log
(`Compiling libdep`, `Fresh probe`). gitdan-ci's runner image forces colour, so
cargo wrote `Compiling\e[0m libdep` into the log and `grep -q "Compiling
libdep"` stopped matching. The suite then reported the opposite of what
happened — the failure printed the log, and the log plainly said `Compiling
libdep`. Both suites now pin CARGO_TERM_COLOR=never, which is the format their
assertions are written against.

Red-proven: with the export removed and CARGO_TERM_COLOR=always,
restore-mtimes-selftest reproduces the CI failure verbatim ("expected to find:
Compiling libdep"); with it, 14/14 pass under the same forced colour.

THE NIGHTLY PROBE ASKED THE WRONG QUESTION. `cargo +nightly -V` answers "did a
cargo proxy called with +nightly exit 0", which is not "will this build have
checksum freshness": `-V` short-circuits before `-Z` is validated at all. So
the probe said "on" on a runner where the flag was not in effect, and the
suite asserted the checksum-freshness mutation family against a build that
never had it — failing in the CONTROL, where a failure reads as "the hazard is
gone" rather than "the toolchain is wrong". It now probes the capability:
`cargo +nightly -Z checksum-freshness locate-project`, the narrowest command
that actually parses the flag. It rejects the stable channel and an unknown
flag name alike, needs no network and builds nothing.

Red-proven: the old channel probe, pointed at a cargo without the flag,
reproduces the CI failure exactly.

AND THE OFF PATH DID NOT WORK EITHER. The suite's header claimed that without
checksum freshness "the test still covers the build/ and *.d families". It did
not: the final scenario backdates the source to 2001 and asserts the rebuild
is not Fresh, which is checksum-freshness-only reasoning. Under Cargo's
ordinary mtime freshness a 2001 source IS older than the artifact and Fresh is
the correct answer, so the scenario asserted a bug. It is now gated on the
mode and skipped loudly, like the control's dep-* assertion already was: 5
assertions with a nightly, 3 without.

CI installs a nightly as well as stable (nightly first, so stable stays the
default) — that hazard is the one this whole scheme exists to close, and a CI
that skips it is checking the cheap half.
2026-08-26 12:46:03 -05:00
claude fb7a788c90 feat(cache): give same-ref jobs separate build directories via cache-lineage
CI / shellcheck + selftests (pull_request) Failing after 1m19s
A cache key names a REF. What a target directory holds is the product of a ref
and a build configuration, and emowheel builds the same ref twice on every
push — once for the host, once for wasm32, in two jobs that start together.
Keyed on the ref alone, `cargo-cache@v1` handed both the same
CARGO_TARGET_DIR, and Cargo's target-directory lock is exclusive: the second
job sat on "Blocking waiting for file lock on build directory" for the length
of the first while holding a runner capacity slot, so a third repo's queued
job waited behind a job doing nothing.

`cache-lineage` is that second dimension. It names ONE directory level under
the cache root:

    <cache-root>/target-<key>              no lineage (unchanged)
    <cache-root>/<lineage>/target-<key>    a lineage

Nesting, not a suffix on the key, and that is the whole design decision.
`prune-cache.sh`'s liveness pass classifies a directory by recomputing
`target-<cache_key(branch)>` for every branch on origin and evicting whatever
does not match — a `target-<key>-wasm32` matches nothing, so it would be
classified dead and evicted unconditionally on every run. daniel/gitdan's
host-level arbiter reads the same shape (BRANCH_DIR_RE); a suffixed name falls
out of that too, so those caches would never be reclaim candidates and a whole
lineage would go missing from the shared disk budget. Nesting leaves both
matchers reading exactly the names they already read, one level down — which
is a layout that arbiter already walks (CI_CACHE_MAX_DEPTH is 2, and its own
suite pins the depth-2 case).

Every interacting part, checked rather than assumed:

- SEED: `seed-target-dir.sh` takes the root as an argument, so a PR branch in
  a lineage layers over THAT lineage's base snapshot. Asserted.
- PUBLISH: `publish-snapshot.sh` derives both ends of the swap from the root.
  The publish action now takes the root from the `CARGO_CACHE_ROOT` the
  consume step exported, and CHECKS its own inputs against it — a publish step
  left at the default while its consume step nested would otherwise republish
  a different lineage's live target dir over that lineage's snapshot, on every
  push, silently. `mode: release-lock` is exempt: it releases a lock on
  `$CARGO_TARGET_DIR` and never touches a root.
- WATERMARK: per target dir, so it follows the lineage. Unchanged.
- PRUNE and LIVENESS: scoped to the root they are given, so a pass in one
  lineage neither evicts nor sees a sibling's caches, or the flat layout's.
  Liveness keeps resolving real branch names, which is what a key suffix would
  have broken.
- ci_cache_reclaim: verified by dry-run against a fixture in this layout —
  all six nested and flat dirs collected as candidates, protection resolved
  correctly on the nested ones, and a `.stage-` stranded inside the lineage
  found by the leftover sweep.

Refused lineage names are refused at resolve time, each rejection naming the
reader that imposes it: a path separator (the arbiter's depth budget), a
Cargo profile name (its no-descend list), a `target-`/`snapshot-` prefix (this
repo's own prune globs), a hex suffix (its per-branch-dir shape), a dot prefix
(the leftover-naming contract). None of these fails visibly on its own — each
produces a working directory that some pass silently stops seeing.

Setting no lineage resolves to the cache root byte for byte, so lublub, zemyna
and emowheel's `ci` job keep the exact directories they have on the volume.

New suite `cache-root-selftest.sh` (19 assertions), red-proven against three
deliberate breakages: a `cache_root_for` that ignores the lineage, a disabled
validator, and a `verify` that never rejects a mismatch.
2026-08-26 12:32:21 -05:00
claude f76789358d ci(gitea): gate scripts/ with shellcheck and the selftests
This repo had eight scripts, six selftest suites and nothing that ran any of
them. It is a composite action consumed at `@v1` — a moving tag — by three
repos' CI, so a bad edit here reaches zemyna, emowheel and lublub at once and
is discovered by whichever of them builds next.

One job: `shellcheck -x --source-path=scripts scripts/*.sh`, then
`bash scripts/selftest.sh`. `-x` follows the `. cache-lib.sh` every script
sources, which is where most of the logic being checked lives; without it
shellcheck reports SC1091 and analyses each file with a hole in it.

The full suite, not `--fast`: `hardlink-clone-selftest.sh` and
`restore-mtimes-selftest.sh` are the two that drive a real Cargo rather than a
fixture, and selftest.sh's own header says `--fast` is for iterating, not for
signing off a change. So the job installs a stable toolchain. The scratch
workspaces they build use path dependencies only, so nothing reaches
crates.io, and the job references no credentials at all.

Modelled on daniel/gitdan's workflow, including the two clauses of the
draft-skip guard (`github.event_name != 'pull_request' ||
!github.event.pull_request.draft` — the first is what stops the second from
also skipping pushes to `main`, where there is no `pull_request` context) and
the per-commit concurrency group for `push`.

No `container.volumes:` entry, deliberately: the suites want a cold target dir
every run, since "did this rebuild?" is exactly what restore-mtimes-selftest
asks. This job takes no share of the shared CI cache disk budget.

Turning the gate on surfaced three pre-existing findings, all fixed here
rather than suppressed or waived:

- `restore-mtimes-selftest.sh` SC2038: `find | xargs touch` -> `-print0 | xargs -0`.
- `hardlink-clone-selftest.sh` SC2295: `${f#$base_fix/}` -> `${f#"$base_fix"/}`.
- `cache-lib.sh` SC2016: a per-line disable with the rationale. The single
  quotes are load-bearing — the program is for the inner shell, where `$f` is
  its loop variable, `$rc` its accumulator and `$$` its pid — so this is the
  documented-false-positive case, not a stopgap.
2026-08-26 12:31:50 -05:00
claude 1aca90b461 test(seed): pin reader_lock_acquire in hardlink_clone_into (#12)
Co-authored-by: Claude <claude@gitdan.com>
Co-committed-by: Claude <claude@gitdan.com>
2026-08-24 23:05:18 +00:00
claude df6f1b91fb docs(cache): producer half of the depended-upon names contract (#11)
Co-authored-by: Claude <claude@gitdan.com>
Co-committed-by: Claude <claude@gitdan.com>
2026-08-24 23:03:29 +00:00
claude a9e9190e5a Merge pull request 'test(seed): pin the two unguarded terms of the torn-clone condition' (#9) from test/pin-clone-guards into main 2026-08-24 19:49:20 +00:00
claude 68da3e69a3 Merge pull request 'docs(cache): write down the producer half of the leftover naming contract' (#8) from docs/contract-and-xrefs into main 2026-08-24 19:29:53 +00:00
claude 31b4113a26 docs(readme): drop an ordering claim scenario 9 does not make
The previous wording said the unshare pass aborts the clone "before the copy's
own exit status is ever consulted", which describes neither version. Unmutated,
the status is consulted immediately after the copy and fires first, so the
unshare pass is never reached; mutated, there is no check left to consult at
either point. As written a reader could take it for a claim that
unshare_mutable_paths runs before the torn-clone condition inside
hardlink_clone_into, which is the kind of ordering this file is otherwise
careful to state exactly (see the reader-marker ordering proof it sits under).

Names the mutation instead: deleting the exit-status check does not change the
outcome, because the unshare pass aborts the clone in its place. The mechanism
is unchanged and still holds — `cp -al` over a mode-000 source creates the
destination preserving mode 000 before failing, and `find | xargs` over that
returns 1 under pipefail, which _unshare_files propagates.
2026-08-24 14:28:08 -05:00
claude eb7878b822 docs(readme): correct the scenario census and the #5 citation
Review of #9 found the methodology section falsified by that PR and owned by
nobody — the scoping that fenced it off was wrong, #8 never touches these
paragraphs.

Three fixes:

* The PATH-stub paragraph named 8a and 8b as the seed suite's stubs. It is now
  four scenarios on two commands: 8a/8b/8c stub `cp` at the clone, 10 stubs it
  one level down at the per-file unshare, and 8d stubs `stat` — a mechanism
  the paragraph did not mention at all, and the only way to make an identity
  that could not be READ the sole witness.

* The assert-which-guard-fired paragraph cited issue #5 as a live example of a
  surviving mutation. #5 is the issue this PR closes, so a reader following
  that citation landed on "removing it leaves every suite green", which is no
  longer true. Scenario 9's own sentence stands — the matrix confirms it
  survives every mutant — so it now says WHY it survives (an unreadable source
  leaves the staging dir at mode 000, and the unshare pass aborts the clone
  before the copy's exit status is consulted) instead of citing a closed
  issue.

* Added the mutual-masking hazard the sweep turned up, since it is the general
  lesson rather than a fact about two particular terms: two guards that can
  each catch the same fault make each other unnecessary, so no fixture built
  around that fault pins either one.

Also names scenario 8d for what it is in its own comment — a regression guard
on a defensive term, not a reproduction of a reachable state. Every route to
the state it constructs is closed off (a rotation hands the witness to 8a, a
genuinely absent source hands it to 8c), which is the reason it is worth
pinning rather than a reason to doubt it.
2026-08-24 14:02:10 -05:00
claude bd60b430e0 test(seed): pin the two unguarded terms of the torn-clone condition
hardlink_clone_into's torn-clone detection is a four-term condition, and a
mutation sweep found two of the four unpinned: removing either
`[ "$cp_rc" -eq 0 ]` (issue #5) or `[ "$i_before" != missing ]` left all five
suites green. They were unpinned for the same reason — they mask each other.
A source that vanishes mid-clone reads as `missing` at both ends AND fails
`cp -al`, so with both terms present either one catches it and neither is
individually necessary.

Isolating them needs a state each term alone can see:

  8c  cp reports failure over a tree that is in fact whole. Neither inference
      sees anything — 0 entries short, one unchanged inode — so the exit
      status is the only witness. Forced with the PATH stub 8a/8b already
      use, on the consumer's own top-level clone.
  8d  both identity reads fail while the copy succeeds. `_dir_inode` folds
      every stat failure into the string `missing`, so two failed reads
      compare equal TO EACH OTHER; without the sentinel term the tree is
      published on the strength of two errors. Stubs `stat` narrowly — only
      the `%i` reads of this clone's own source — because taking the source
      away would fail `cp -al` too and pin 8c's property over again.

The sweep also found `_unshare_files`'s xargs status unpinned, which is the
guard that stops a staging tree whose dep-info files still point at the
SOURCE's inodes from being renamed into place — not a tear, so all four
clone checks pass it, and exactly the silent cross-branch stale-reuse the
scheme exists to prevent. Scenario 10 pins it by refusing the dep-info
unshares and asserting the clone discards rather than publishes.

assert_tear gains an `unreadable` identity expectation, and its `same` case
now demands a READ identity rather than two equal strings — `missing` equals
`missing`, which is the exact confusion 8d exists to pin.

Each of the five pinning scenarios was run in isolation against each
mutation; the result is a clean diagonal, so every scenario fails only for
its own term.

Closes #5
2026-08-24 13:24:53 -05:00
claude e3869c5920 docs(publish): address review nits on the contract block
Four corrections from the review of #8, all in the files this PR already
touches:

- A local signal on the line that creates .publish-new-. The other four shapes
  each got a note at their producing line, which is the whole premise of #7 —
  someone renaming TMP_DST reads its own comment block and would never see the
  contract note 45 lines up at OLD.

- '.publish-old- is milder' understated it. Milder is true; bounded is not. A
  key whose branch is merged, deleted or renamed is never published again, so
  its rotated generation stays until something outside this repo takes it —
  which is the case gitdan#30 itself makes, two lines away.

- The drift check now says the prefix constants ci-cache-reclaim.sh DECLARES,
  not the ones it enumerates. Those are different numbers: collect_entries()
  globs .stage- and .evicting- only, because .reading- is read and never
  swept. Only the declared reading makes the five-against-five count work, and
  a countability check that needs a coin flip to count is not one.

- publish-snapshot.sh's own header had two stale names ten lines above the
  stale pointer this PR fixes: step 1 staged at .stage-<tag> (that is
  hardlink_clone_into's inner path; the staged snapshot is .publish-new-<tag>)
  and step 2 named .publish-old-<tag> without the key. Pre-existing and
  outside both ACs, but #6's thesis is that a plausible-looking wrong name is
  the worst kind, and these are in the file the PR is about.

Comments and docs only. With comments and blank lines stripped, all three
scripts hash identically to origin/main.
2026-08-24 13:06:49 -05:00
claude 0118c28f01 docs(cache): state the .publish-* gap as a ticket, not a present state
The contract block asserted that neither side reclaims .publish-new- 'today'.
True when written and about to stop being true: gitdan#30 tracks adding both
.publish-* prefixes to the arbiter's enumeration, and a sibling track is
landing it this round. A comment that dates itself against a merge in flight
is worse than no comment.

Rewords all four sites (cache-lib.sh, publish-snapshot.sh, and README's table
row and prose) to reference gitdan#30 and keep the mechanism that made the
shape worth catching — .publish-new- is tagged per job per run exactly as
.stage- is — rather than the arbiter's momentary contents. The rule itself is
unchanged; it is the durable part, and it is what found this.

Adds the counting check while there: the five names here and the prefixes
ci-cache-reclaim.sh enumerates are meant to be the same length, so a mismatch
is the cheapest signal that one side gained a shape without telling the other.
2026-08-24 12:56:18 -05:00