Author SHA1 Message Date
claude e94c46f2ec test(hardlink): move the probe's INACTIVE answer off bash's error codes
CI / shellcheck + selftests (pull_request) Successful in 1m21s
Review finding on #13, and the last change to this function. A typo'd variable
name on a line that HAS its `|| exit 2` guard — `$CONTNET_B` for `$CONTENT_B`
— fails in EXPANSION, before the command runs, so neither the guard nor the
ERR trap ever sees it. Under `set -u` that exits 1, which was the code meaning
"measured, and the answer is mtime". Not a live defect: the probe references
only set variables today, and the 1 gitdan-ci reports is a genuine measurement.

The defect is that the answer codes and the error codes overlapped at all.
Three rounds on this function each found a narrower way for a shell-generated
1 to be read as a measurement — an unguarded command, then a guarded line
whose guard could not fire — and each was closed by narrowing the failure
surface, which is a game with no last move.

INACTIVE is now 3, and anything that is not 0 or 3 is NOT MEASURED. Bash
generates 1, 2, 126, 127 and 128+n for its own errors and never 3, so "not an
answer" is decided by a property of the shell rather than by an enumeration of
the ways a step can go wrong. A step added later without a guard, or with a
guard that cannot fire, lands on NOT MEASURED by construction.

The guards and the trap stay, with their job restated: reporting rather than
correctness. They put a failed step on 2 with both probe logs printed instead
of on some incidental status — nicer to debug, and the same destination either
way. The `set +e` bracket at the call site keeps its own note, since a `set -e`
that cannot fire is the kind of thing a reader assumes works.

The not-measured message now quotes the actual exit code, so the two ways to
reach it are distinguishable in a log without reading this file.

Red-proven with the reviewer's own mutation, A/B: with INACTIVE at 1 the typo
reports `mode: off — resolves freshness by mtime`, a false measurement; at 3 it
reports `mode: unmeasured — the probe exited 1`. All four paths re-verified —
unmodified 4 assertions; env stripped inside the probe, measured INACTIVE with
the mtime reason; a broken probe build, unmeasured via a guard; an unguarded
`false`, unmeasured via the trap.
2026-08-26 14:40:27 -05:00
claude 5b46986cd7 test(hardlink): make the errexit invariant real, and stop gpg deciding a gate
CI / shellcheck + selftests (pull_request) Successful in 1m23s
Two review findings on #13.

THE STATED INVARIANT WAS NOT IN FORCE. The header claimed a `set -e` abort
could never be mistaken for the `1` that means "measured, and the answer is
mtime". It could not fire at all: `checksum_freshness_probe || probe_rc=$?`
runs the left side with errexit suppressed, and that suppression propagates
into the subshell. An unguarded step there fell through to `exit 0` and
reported ACTIVE — a fourth outcome the header said was impossible. The guard
audit was complete, so nothing was broken; the wrong mechanism had the credit,
in a comment inviting the next editor to add an unguarded step and rely on it.

The suggested fix — `set -e` as the first line of the subshell — does not work,
and measured on bash 5.3 it makes things worse rather than not-better:

  set -e inside a subshell called via `||` or `if`   still falls through: rc=0
  set -e inside, called with errexit off at the site aborts, but with the
                                                     FAILING COMMAND's status —
                                                     `false` gives 1, which is
                                                     exactly the value that
                                                     means "mtime"

So both halves are needed and neither is decoration: `set +e` around the call
site, so the subshell can arm its own errexit at all, and `trap 'exit 2' ERR`
inside it, so an abort lands on "not measured" instead of on an answer. The
explicit `|| exit 2` guards stay as the first line of defence. All of that is
now written down where the false claim was.

Red-proven with one unguarded `false` between the `touch` and the second
build, on a box where freshness is live: without the fix, mode `on` and 4
assertions — the reviewer's fourth outcome, reproduced; with it, mode
`unmeasured`, the warning, 3 assertions; unmodified, mode `on` and 4.

Recorded in the same block, since the next person to hit a `2` will ask: why it
skips rather than fails. The gated scenario is the only thing here that depends
on freshness mode, and failing would turn a statement about one machine into a
red gate reading "the hardlink scheme is broken" across three consuming repos.
What would change the answer is `2` becoming the everyday CI outcome, and run
2583 measures that it is not — gitdan-ci reports `1`.

GPG. The two suites that commit in scratch repos inherited the developer's
GLOBAL `commit.gpgsign`, so whether this gate passes could depend on their gpg
agent — seen as a red `restore-mtimes-selftest.sh` caused by a full disk
breaking gpg, in a suite with nothing to say about either. Pinned off locally,
in the throwaway repos only: `git config commit.gpgsign false` beside the
identity the scratch repo already sets, and `-c commit.gpgsign=false` on
prune-cache-selftest's three commits, matching its existing `-c` style.
Red-proven under a GIT_CONFIG_GLOBAL with gpgsign on and a nonexistent gpg
program: without it `fatal: failed to write commit object`, with it both
suites pass.
2026-08-26 14:28:22 -05:00
claude 5a90d5c829 test(hardlink): let the freshness probe report that it could not measure
CI / shellcheck + selftests (pull_request) Successful in 1m21s
Review finding on #13. `checksum_freshness_active()` returned non-zero for any
reason — including its own `cargo build` failing — and the caller's `else`
branch then announced "this nightly accepts -Z checksum-freshness but still
resolves freshness by mtime" regardless. On a machine where content freshness
IS live, breaking only the probe's build produced that sentence, which is
false, and silently dropped the scenario, which is lost coverage. Reachable,
not theoretical: it needs a half-installed toolchain, which is exactly the
state a probe exists to notice.

The whole point of the preceding commit was to settle this by experiment
rather than by asking, and an experiment that cannot tell a negative result
from a failed measurement is not settling it. "This toolchain resolves
freshness by mtime" is a statement about Cargo; "a probe build failed" is a
statement about this machine. They are not interchangeable.

Three outcomes now, carried in an exit code:

  0  measured ACTIVE     both builds ran, the backdated rebuild recompiled
  1  measured INACTIVE   both builds ran, the backdated rebuild was Fresh
  2  NOT MEASURED        a probe step failed; nothing was learned

Every step that could fail for a reason other than the experiment's own
outcome exits 2 explicitly, so a `set -e` abort cannot be mistaken for the 1
that means "measured, and the answer is mtime". Outcome 2 emits a ::warning::
saying in those words that this is a failure to measure and not a finding
about Cargo, and prints the tail of both probe logs. The scenario still skips
for 1 and 2 alike — everything else the suite asserts is independent of
freshness mode — but the reason is now carried to the skip line, so the two
are distinguishable at a glance.

Verified all three states on a machine where freshness is live: unmodified,
mode `on` and 4 assertions; probe build broken with a bogus flag, mode
`unmeasured` with the warning and 3 assertions, and no claim about mtime;
CARGO_UNSTABLE_CHECKSUM_FRESHNESS stripped inside the probe — CI's condition —
mode `off` naming mtime, and 3 assertions.
2026-08-26 13:34:42 -05:00
claude eb3f0bd09b docs(ci): say what the nightly step actually buys today
CI / shellcheck + selftests (pull_request) Successful in 1m24s
The step's comment claimed nightly was not optional. CI's own first green run
falsified that: 1.100.0-nightly accepts -Z checksum-freshness and resolves
freshness by mtime regardless, so the suite's probe reports it and skips the
scenario either way.

The step stays — twenty seconds, and the coverage returns by itself the day
upstream restores the behaviour — but the comment now says that rather than
the opposite. The open question about what upstream actually did is
daniel/gitdan#62.
2026-08-26 12:57:15 -05:00
claude 6a80423bd0 test(hardlink): settle checksum freshness by experiment, not by asking
CI / shellcheck + selftests (pull_request) Successful in 1m17s
Third toolchain probe in this branch, and the first one that asks the
question the suite actually depends on. The two before it each let the suite
assert a property the toolchain did not have:

  cargo +nightly -V                    answers "did a proxy called with
                                       +nightly exit 0". `-V` short-circuits
                                       before `-Z` is parsed at all.
  cargo +nightly -Z checksum-freshness answers "is the flag still accepted".
   locate-project                      1.100.0-nightly (2026-08-25) accepts it
                                       and resolves freshness by mtime anyway,
                                       which is how CI reached the final
                                       scenario and failed there.

The final scenario depends on exactly one property: that changed content with
an OLDER mtime rebuilds. Under mtime freshness the correct answer is Fresh, so
under mtime freshness that scenario asserts a bug — which is what CI reported.

So the probe performs that experiment, on its own crate and its own target
dir, with no clone anywhere near it. That separation is also what keeps it a
control rather than a restatement: the probe establishes that the toolchain
rebuilds on content, the scenario establishes that a hardlink clone did not
take that away.

Verified both ways locally: with a real content-freshness nightly, mode on and
4 assertions; with the env var stripped inside the probe — the runner's
condition, faithfully — the probe says so by name, mode goes off, the scenario
is skipped loudly and the remaining 3 assertions pass.
2026-08-26 12:54:24 -05:00
claude 5e8e773f78 test(hardlink): report the mutation families, don't assert one of them
CI / shellcheck + selftests (pull_request) Failing after 1m21s
CI's nightly is 1.100.0-nightly (2026-08-25); the machine this suite was
written on had 1.96.0-nightly (2026-02-24). On the newer one the control's
`cp -al` clone mutates only the build/ and *.d families — upstream appears to
have stopped rewriting `.fingerprint/*/dep-*` in place under checksum
freshness — so the suite went red on the *absence* of a hazard.

That is the wrong shape for a gate. The control's job is to prove the hazard
exists at all, which a non-empty mutated set already does; naming one family
as mandatory makes the suite red whenever upstream stops doing something we
never wanted it to do, and red in the CONTROL, where a failure reads as "the
hazard is gone" rather than "upstream changed". It is now a note either way.

Nothing is given up. The fix scenario asserts the source is byte-identical
after a full rebuild in the clone, which covers every family the running Cargo
has, named or not — and the checksum-freshness scenario after it tests the
stale-reuse hazard directly. The dep-* line only ever documented which family
was in play.

`unshare_mutable_paths` keeps unsharing dep-* regardless, and its measurement
block now records both observations with their versions: 22 MB of a 6.9 GB
tree against a failure mode that is a wrong answer rather than a slow one.
2026-08-26 12:50:42 -05:00
claude 3f97d3d1e7 fix(selftest): make the two compiler-backed suites survive a CI runner
CI / shellcheck + selftests (pull_request) Failing after 1m36s
The new gate's first run went red on both suites that drive a real Cargo.
Neither failure was a defect in what they test; both were assumptions about
the machine, which had only ever been a dev box with a nightly installed and
no colour forcing. Fixed here rather than waived — turning a gate on is what
obliges fixing what it finds.

COLOUR. Every assertion in restore-mtimes-selftest.sh, and one in
hardlink-clone-selftest.sh, reads cargo's own words out of a build log
(`Compiling libdep`, `Fresh probe`). gitdan-ci's runner image forces colour, so
cargo wrote `Compiling\e[0m libdep` into the log and `grep -q "Compiling
libdep"` stopped matching. The suite then reported the opposite of what
happened — the failure printed the log, and the log plainly said `Compiling
libdep`. Both suites now pin CARGO_TERM_COLOR=never, which is the format their
assertions are written against.

Red-proven: with the export removed and CARGO_TERM_COLOR=always,
restore-mtimes-selftest reproduces the CI failure verbatim ("expected to find:
Compiling libdep"); with it, 14/14 pass under the same forced colour.

THE NIGHTLY PROBE ASKED THE WRONG QUESTION. `cargo +nightly -V` answers "did a
cargo proxy called with +nightly exit 0", which is not "will this build have
checksum freshness": `-V` short-circuits before `-Z` is validated at all. So
the probe said "on" on a runner where the flag was not in effect, and the
suite asserted the checksum-freshness mutation family against a build that
never had it — failing in the CONTROL, where a failure reads as "the hazard is
gone" rather than "the toolchain is wrong". It now probes the capability:
`cargo +nightly -Z checksum-freshness locate-project`, the narrowest command
that actually parses the flag. It rejects the stable channel and an unknown
flag name alike, needs no network and builds nothing.

Red-proven: the old channel probe, pointed at a cargo without the flag,
reproduces the CI failure exactly.

AND THE OFF PATH DID NOT WORK EITHER. The suite's header claimed that without
checksum freshness "the test still covers the build/ and *.d families". It did
not: the final scenario backdates the source to 2001 and asserts the rebuild
is not Fresh, which is checksum-freshness-only reasoning. Under Cargo's
ordinary mtime freshness a 2001 source IS older than the artifact and Fresh is
the correct answer, so the scenario asserted a bug. It is now gated on the
mode and skipped loudly, like the control's dep-* assertion already was: 5
assertions with a nightly, 3 without.

CI installs a nightly as well as stable (nightly first, so stable stays the
default) — that hazard is the one this whole scheme exists to close, and a CI
that skips it is checking the cheap half.
2026-08-26 12:46:03 -05:00
claude fb7a788c90 feat(cache): give same-ref jobs separate build directories via cache-lineage
CI / shellcheck + selftests (pull_request) Failing after 1m19s
A cache key names a REF. What a target directory holds is the product of a ref
and a build configuration, and emowheel builds the same ref twice on every
push — once for the host, once for wasm32, in two jobs that start together.
Keyed on the ref alone, `cargo-cache@v1` handed both the same
CARGO_TARGET_DIR, and Cargo's target-directory lock is exclusive: the second
job sat on "Blocking waiting for file lock on build directory" for the length
of the first while holding a runner capacity slot, so a third repo's queued
job waited behind a job doing nothing.

`cache-lineage` is that second dimension. It names ONE directory level under
the cache root:

    <cache-root>/target-<key>              no lineage (unchanged)
    <cache-root>/<lineage>/target-<key>    a lineage

Nesting, not a suffix on the key, and that is the whole design decision.
`prune-cache.sh`'s liveness pass classifies a directory by recomputing
`target-<cache_key(branch)>` for every branch on origin and evicting whatever
does not match — a `target-<key>-wasm32` matches nothing, so it would be
classified dead and evicted unconditionally on every run. daniel/gitdan's
host-level arbiter reads the same shape (BRANCH_DIR_RE); a suffixed name falls
out of that too, so those caches would never be reclaim candidates and a whole
lineage would go missing from the shared disk budget. Nesting leaves both
matchers reading exactly the names they already read, one level down — which
is a layout that arbiter already walks (CI_CACHE_MAX_DEPTH is 2, and its own
suite pins the depth-2 case).

Every interacting part, checked rather than assumed:

- SEED: `seed-target-dir.sh` takes the root as an argument, so a PR branch in
  a lineage layers over THAT lineage's base snapshot. Asserted.
- PUBLISH: `publish-snapshot.sh` derives both ends of the swap from the root.
  The publish action now takes the root from the `CARGO_CACHE_ROOT` the
  consume step exported, and CHECKS its own inputs against it — a publish step
  left at the default while its consume step nested would otherwise republish
  a different lineage's live target dir over that lineage's snapshot, on every
  push, silently. `mode: release-lock` is exempt: it releases a lock on
  `$CARGO_TARGET_DIR` and never touches a root.
- WATERMARK: per target dir, so it follows the lineage. Unchanged.
- PRUNE and LIVENESS: scoped to the root they are given, so a pass in one
  lineage neither evicts nor sees a sibling's caches, or the flat layout's.
  Liveness keeps resolving real branch names, which is what a key suffix would
  have broken.
- ci_cache_reclaim: verified by dry-run against a fixture in this layout —
  all six nested and flat dirs collected as candidates, protection resolved
  correctly on the nested ones, and a `.stage-` stranded inside the lineage
  found by the leftover sweep.

Refused lineage names are refused at resolve time, each rejection naming the
reader that imposes it: a path separator (the arbiter's depth budget), a
Cargo profile name (its no-descend list), a `target-`/`snapshot-` prefix (this
repo's own prune globs), a hex suffix (its per-branch-dir shape), a dot prefix
(the leftover-naming contract). None of these fails visibly on its own — each
produces a working directory that some pass silently stops seeing.

Setting no lineage resolves to the cache root byte for byte, so lublub, zemyna
and emowheel's `ci` job keep the exact directories they have on the volume.

New suite `cache-root-selftest.sh` (19 assertions), red-proven against three
deliberate breakages: a `cache_root_for` that ignores the lineage, a disabled
validator, and a `verify` that never rejects a mismatch.
2026-08-26 12:32:21 -05:00
claude f76789358d ci(gitea): gate scripts/ with shellcheck and the selftests
This repo had eight scripts, six selftest suites and nothing that ran any of
them. It is a composite action consumed at `@v1` — a moving tag — by three
repos' CI, so a bad edit here reaches zemyna, emowheel and lublub at once and
is discovered by whichever of them builds next.

One job: `shellcheck -x --source-path=scripts scripts/*.sh`, then
`bash scripts/selftest.sh`. `-x` follows the `. cache-lib.sh` every script
sources, which is where most of the logic being checked lives; without it
shellcheck reports SC1091 and analyses each file with a hole in it.

The full suite, not `--fast`: `hardlink-clone-selftest.sh` and
`restore-mtimes-selftest.sh` are the two that drive a real Cargo rather than a
fixture, and selftest.sh's own header says `--fast` is for iterating, not for
signing off a change. So the job installs a stable toolchain. The scratch
workspaces they build use path dependencies only, so nothing reaches
crates.io, and the job references no credentials at all.

Modelled on daniel/gitdan's workflow, including the two clauses of the
draft-skip guard (`github.event_name != 'pull_request' ||
!github.event.pull_request.draft` — the first is what stops the second from
also skipping pushes to `main`, where there is no `pull_request` context) and
the per-commit concurrency group for `push`.

No `container.volumes:` entry, deliberately: the suites want a cold target dir
every run, since "did this rebuild?" is exactly what restore-mtimes-selftest
asks. This job takes no share of the shared CI cache disk budget.

Turning the gate on surfaced three pre-existing findings, all fixed here
rather than suppressed or waived:

- `restore-mtimes-selftest.sh` SC2038: `find | xargs touch` -> `-print0 | xargs -0`.
- `hardlink-clone-selftest.sh` SC2295: `${f#$base_fix/}` -> `${f#"$base_fix"/}`.
- `cache-lib.sh` SC2016: a per-line disable with the rationale. The single
  quotes are load-bearing — the program is for the inner shell, where `$f` is
  its loop variable, `$rc` its accumulator and `$$` its pid — so this is the
  documented-false-positive case, not a stopgap.
2026-08-26 12:31:50 -05:00
claude 1aca90b461 test(seed): pin reader_lock_acquire in hardlink_clone_into (#12)
Co-authored-by: Claude <claude@gitdan.com>
Co-committed-by: Claude <claude@gitdan.com>
2026-08-24 23:05:18 +00:00
claude df6f1b91fb docs(cache): producer half of the depended-upon names contract (#11)
Co-authored-by: Claude <claude@gitdan.com>
Co-committed-by: Claude <claude@gitdan.com>
2026-08-24 23:03:29 +00:00
12 changed files with 966 additions and 52 deletions
+98
View File
@@ -0,0 +1,98 @@
name: CI
on:
push:
branches: [main]
pull_request:
branches: [main]
# Spelled out only to keep `ready_for_review` in the list — naming any type
# replaces the whole default set, so the other three have to be restated.
# It is inert on this instance (draft state here is the `WIP:` title
# prefix, so un-drafting is a title edit and raises no
# `ready_for_review` action) and costs nothing.
#
# The consequence, which is the part that bites: the `if:` guard below is
# evaluated when a run is CREATED, and un-drafting creates no run. A PR
# opened as a draft keeps its skip decision until something else produces
# one. Push an empty commit after un-WIP'ing.
types: [opened, synchronize, reopened, ready_for_review]
# gitdan-ci runs four repos' CI on two capacity slots, and the compiler-backed
# suites below are multi-minute. A superseded run costs a slot in front of
# somebody's build, so drop it.
#
# `push` groups on `github.sha` rather than `github.ref`: a constant per-branch
# group is what let this Gitea (1.26.0) cancel two of daniel/gitdan's merge
# runs outright while `cancel-in-progress` was gated away from `push` entirely
# — see the long note in that repo's ci.yaml for the evidence. Giving every
# commit its own group leaves that behaviour nothing to act on.
concurrency:
group: ${{ github.workflow }}-${{ github.event_name == 'pull_request' && github.ref || github.sha }}
cancel-in-progress: true
jobs:
selftest:
name: shellcheck + selftests
# Two clauses, both load-bearing. The second skips draft PRs: Gitea sets
# draft:true when the title starts with `WIP:`, so work-in-progress pushes
# cost the shared runner nothing until the PR is un-WIP'd. The first is
# what keeps that from also skipping pushes to `main` — a `push` event has
# no `pull_request` context, so `github.event.pull_request.draft` is empty
# there and the negation alone would be unreliable. Never drop it.
if: ${{ github.event_name != 'pull_request' || !github.event.pull_request.draft }}
runs-on: ubuntu-latest
timeout-minutes: 20
# Deliberately no `container.volumes:` entry, unlike every repo that
# CONSUMES this action. The suites below build throwaway workspaces under
# `mktemp -d` and want a cold target dir every time — a persistent cache
# would make "did this run rebuild?" unanswerable, which is the question
# restore-mtimes-selftest.sh exists to ask. So this job takes no share of
# the shared CI cache disk budget.
steps:
- name: Checkout sources
uses: actions/checkout@v4
- name: Install shellcheck
uses: taiki-e/install-action@v2
with:
tool: shellcheck
# `hardlink-clone-selftest.sh` and `restore-mtimes-selftest.sh` drive a
# real Cargo against a real scratch workspace — they are the only things
# here that verify the hardlink-aliasing and mtime-freshness behaviour
# against the compiler rather than against a fixture, and selftest.sh's
# own header says `--fast` is for iterating, not for signing off a
# change. So CI installs a toolchain and runs the full set.
#
# The scratch workspaces use path dependencies only, so nothing here
# reaches crates.io.
# Nightly first, stable second, so stable ends up the default and
# nightly is only reachable through an explicit `+nightly`.
#
# `hardlink-clone-selftest.sh`'s last scenario needs a Cargo that
# resolves freshness by CONTENT — the mode where the dep-info file
# carries per-source checksums, which is the mutation that turns a
# hardlink clone into silent stale-artifact reuse rather than a slow
# build. As of 1.100.0-nightly (2026-08-25) no nightly provides it:
# `-Z checksum-freshness` is still accepted and freshness is still
# resolved by mtime, so the suite's probe reports that by name and skips
# the scenario. This step therefore buys nothing today and is kept
# anyway — it costs about twenty seconds, and the day upstream restores
# the behaviour the coverage comes back with no edit here. See daniel/gitdan#62.
- name: Install Rust nightly
uses: dtolnay/rust-toolchain@nightly
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
# Every script, including the suites themselves. `-x` follows the
# `. cache-lib.sh` each one sources, which is where most of the logic
# being checked actually lives; without it shellcheck reports SC1091 and
# analyses each file with a hole in it.
- name: shellcheck
run: shellcheck -x --source-path=scripts scripts/*.sh
# One command, not six: selftest.sh is the entry point a developer runs,
# so a suite added there is gated here without a matching edit in this
# file.
- name: Selftests
run: bash scripts/selftest.sh
+140 -15
View File
@@ -317,6 +317,39 @@ spelling of some new dot-prefixed name — a name it reads for a decision but
never reclaims — that is exactly this category, and it needs the matching never reclaims — that is exactly this category, and it needs the matching
`DEPEND_*` entry on gitdan's side before it ships, not after. `DEPEND_*` entry on gitdan's side before it ships, not after.
### The directory LAYOUT is part of that contract as well
Names are one half; where they sit is the other. gitdan's arbiter walks a
volume's `_data` tree to `CI_CACHE_MAX_DEPTH`, which is **2** — deliberately
tight, because a deeper walk starts meeting Cargo's own
`incremental/<crate>-<hash>` directories, which match the same name shape it
uses to recognise a cache dir and must never be evicted individually. So:
```
_data/target-<key> depth 1 — no lineage
_data/<lineage>/target-<key> depth 2 — a lineage
_data/<a>/<b>/target-<key> depth 3 — INVISIBLE to the arbiter
```
That budget is the whole reason `cache-lineage` is one path component and not
a path. Nesting deeper is not an error anywhere: the caches work, the
in-workflow prune pass keeps managing them, and the one script whose job is
the shared disk budget across every repo simply never sees them again.
It is also why the fix for daniel/gitdan#60 nests rather than suffixing the
cache key. A `target-<key>-<lineage>` name would be read as dead by
`prune-cache.sh`'s liveness pass — which classifies by recomputing
`target-<cache_key(branch)>` for every branch on origin — and evicted
unconditionally on every run; and it falls out of the arbiter's own
`BRANCH_DIR_RE` too, so the same directories would never be candidates there
either. Nesting leaves both matchers reading exactly the names they already
read, one level down.
One known rough edge, on gitdan's side and cosmetic: that script logs an
eviction as `<volume>/<basename>`, so a nested `target-<key>` and a flat one
of the same key are indistinguishable in its output. It evicts the right
directory; the line just doesn't say which.
--- ---
## Inputs ## Inputs
@@ -326,6 +359,7 @@ never reclaims — that is exactly this category, and it needs the matching
| input | default | meaning | | input | default | meaning |
|---|---|---| |---|---|---|
| `cache-root` | `/cache` | mount point of the persistent volume inside the job container | | `cache-root` | `/cache` | mount point of the persistent volume inside the job container |
| `cache-lineage` | *(empty)* | one directory level under `cache-root`, for a second job building the same ref for a different target or profile — see [Multiple jobs in one workflow](#multiple-jobs-in-one-workflow) |
| `protected-branches` | `dev main` | refs that publish snapshots and are never evicted | | `protected-branches` | `dev main` | refs that publish snapshots and are never evicted |
| `min-free-percent` | `10` | prune when free space drops below this | | `min-free-percent` | `10` | prune when free space drops below this |
| `restore-mtimes` | `true` | restore tracked-file mtimes from git history | | `restore-mtimes` | `true` | restore tracked-file mtimes from git history |
@@ -334,7 +368,7 @@ never reclaims — that is exactly this category, and it needs the matching
| `own-ref` | *(auto)* | override; defaults to `github.head_ref`, else `github.ref_name` | | `own-ref` | *(auto)* | override; defaults to `github.head_ref`, else `github.ref_name` |
| `base-ref` | *(auto)* | override; defaults to `github.base_ref` (empty on push) | | `base-ref` | *(auto)* | override; defaults to `github.base_ref` (empty on push) |
| `seed-fallback-dir` | *(empty)* | absolute path to seed from when no snapshot exists — for migrating off an existing flat cache | | `seed-fallback-dir` | *(empty)* | absolute path to seed from when no snapshot exists — for migrating off an existing flat cache |
| `watermark-file` | `.ci-watermark-<job>-sha` | must differ per job when two jobs share one cache key | | `watermark-file` | `.ci-watermark-<job>-sha` | must differ per job when two jobs share one target directory; the default already does |
| `lock-id` | `<job>-<run_id>` | identifies this job's cache lock | | `lock-id` | `<job>-<run_id>` | identifies this job's cache lock |
| `stale-lock-seconds` | `7200` | age past which another job's lock is treated as abandoned | | `stale-lock-seconds` | `7200` | age past which another job's lock is treated as abandoned |
@@ -350,6 +384,7 @@ Exports to the job environment: `CARGO_TARGET_DIR`, `CARGO_CACHE_ROOT`,
| input | default | meaning | | input | default | meaning |
|---|---|---| |---|---|---|
| `cache-root` | `/cache` | must match the consume action | | `cache-root` | `/cache` | must match the consume action |
| `cache-lineage` | *(empty)* | must match the consume action; a mismatch fails the step rather than publishing the wrong tree |
| `protected-branches` | `dev main` | refs that publish snapshots | | `protected-branches` | `dev main` | refs that publish snapshots |
| `mode` | `publish` | `publish`, or `release-lock` for the `if: always()` step | | `mode` | `publish` | `publish`, or `release-lock` for the `if: always()` step |
| `own-ref` | *(auto)* | override; defaults to `github.head_ref`, else `github.ref_name` | | `own-ref` | *(auto)* | override; defaults to `github.head_ref`, else `github.ref_name` |
@@ -366,13 +401,70 @@ of a merge-preview build, which is not what `dev` is.
## Multiple jobs in one workflow ## Multiple jobs in one workflow
Jobs sharing a cache key (a `ci` job and a `wasm` job on the same branch, say) Two jobs building the same ref — a `ci` job and a `wasm` job, say — are two
each need their **own** watermark file. A shared one breaks the moment two consumers of one cache key, and the cache key alone is not enough to keep them
jobs run in sequence within one trigger: job A advances the watermark to HEAD, apart.
and job B then reads that just-advanced value, computes an empty diff, and
loses the merge protection entirely. The default (`.ci-watermark-<job>-sha`) **Give each its own lineage.** A cache key names a *ref*; what a target
already gives each job its own; only override `watermark-file` if you also directory holds is the product of a ref and a build configuration. Left to the
override `lock-id`, and then keep both distinct per job. key alone, both jobs export the same `CARGO_TARGET_DIR`, and Cargo's
build-directory lock is exclusive — so on a runner with more than one slot the
second job sits on `Blocking waiting for file lock on build directory` for the
length of the first, occupying a capacity slot while doing nothing
(daniel/gitdan#60). `cache-lineage` is that second dimension:
```yaml
- name: Restore the Cargo cache
uses: https://gitdan.com/daniel/gitdan-actions/cargo-cache@v1
with:
cache-lineage: wasm32 # the `ci` job sets none
# ... build steps ...
- name: Record watermark, publish cache snapshot
uses: https://gitdan.com/daniel/gitdan-actions/cargo-cache-publish@v1
with:
cache-lineage: wasm32 # the SAME value, or the step fails
```
A lineage nests one directory level under the cache root
(`<cache-root>/<lineage>/target-<key>`), so each lineage gets its own target
dirs, its own snapshots, and its own prune pass. Everything else works as it
already did, one level down: a PR branch in a lineage layers over **that
lineage's** base snapshot, the publisher branch publishes into it, and a prune
pass run inside it never sees a sibling lineage's caches.
Setting no lineage resolves to the cache root unchanged, byte for byte, so a
workflow that does not use one keeps the exact directories it already has on
the volume.
**Both actions need the same value.** `cargo-cache-publish` derives both ends
of the snapshot swap from its own `cache-root`, so a publish step left at the
default while its consume step nested would republish a *different* lineage's
live target dir over that lineage's snapshot, on every push, with nothing in
the log to say so. The publish action therefore compares its own inputs
against the `CARGO_CACHE_ROOT` the consume step exported and fails the step on
a mismatch. (The `mode: release-lock` call is exempt: it releases a lock on
`$CARGO_TARGET_DIR` and never touches a cache root, so it takes no lineage.)
**Some lineage names are refused.** A lineage is one path component, drawn
from `[A-Za-z0-9._-]`, and several otherwise-reasonable names are rejected at
resolve time because a *reader elsewhere* would stop seeing the caches
underneath them: a Cargo profile name (`debug`, `release`, `doc`, …) is one
gitdan's arbiter never descends into, a `target-`/`snapshot-` prefix makes the
lineage directory itself an eviction candidate for this repo's own prune pass,
and a hex-suffixed name is read by that arbiter as a per-branch cache dir in
its own right. `validate_cache_lineage()` in `scripts/cache-lib.sh` states each
rejection with the reader that imposes it.
**Watermarks are still per job.** Two jobs in one lineage — or one job before
lineages were introduced — each need their **own** watermark file. A shared one
breaks the moment two jobs run in sequence within one trigger: job A advances
the watermark to HEAD, and job B then reads that just-advanced value, computes
an empty diff, and loses the merge protection entirely. The default
(`.ci-watermark-<job>-sha`) already gives each job its own; only override
`watermark-file` if you also override `lock-id`, and then keep both distinct
per job.
--- ---
@@ -445,14 +537,33 @@ not automatic.
## Development ## Development
```bash ```bash
bash scripts/selftest.sh # everything (~1 min; needs cargo) shellcheck -x --source-path=scripts scripts/*.sh
bash scripts/selftest.sh # everything (needs cargo)
bash scripts/selftest.sh --fast # fixture-only suites, no compiler bash scripts/selftest.sh --fast # fixture-only suites, no compiler
``` ```
Both run in CI — `.gitea/workflows/ci.yaml`, one job, on pushes to `main` and
on PRs that were non-draft when the run was created. It installs shellcheck
and both a stable and a nightly Rust toolchain (nightly so
`hardlink-clone-selftest.sh` can run its content-freshness scenario — which no
nightly currently enables, so it is skipped and the step is kept only against
the day upstream restores it) and references no credentials; the scratch
workspaces the compiler-backed suites build use path dependencies only, so
nothing reaches crates.io. It runs the full suite rather than `--fast`,
because the two compiler-backed suites are the ones that check this scheme
against real Cargo instead of against a fixture. Draft (`WIP:`-titled) PRs
skip it, and un-drafting does **not** un-skip them — the guard is evaluated
when a run is created and un-drafting creates none, so push an empty commit
after un-WIP'ing.
This repo is consumed by three other repos' CI at `@v1`, a moving tag, so a
change here reaches all of them at once. That is what the gate is for.
| suite | covers | | suite | covers |
|---|---| |---|---|
| `hardlink-clone-selftest.sh` | that a build in a clone cannot mutate its source — with a control proving a raw `cp -al` does. Needs a real compiler. | | `cache-root-selftest.sh` | that a lineage nests one level and nothing else moves: no lineage resolves byte-for-byte to the cache root, two lineages on one cache key get disjoint target dirs, seed/publish/prune all stay inside their own lineage, a PR layers over its own lineage's base snapshot — **and one rejection per lineage name a reader elsewhere would stop seeing**, plus the publish-side mismatch guard |
| `seed-target-dir-selftest.sh` | seed-source preference, lock-file stripping, two jobs racing on one cache key, **and one scenario per check a hardlink clone is validated against**: a source rotated wholesale, a subtree silently lost from the walk, a copy that reports failure over a tree both other checks read as whole, and a source identity that resolved at neither end — plus a staging tree that could not be privately owned being discarded rather than published | | `hardlink-clone-selftest.sh` | that a build in a clone cannot mutate its source — with a control proving a raw `cp -al` does. Needs a real compiler, **and a nightly that actually resolves freshness by content for its last scenario**: the source's-next-build check reasons about content rather than mtime, so under mtime freshness it would assert a bug. Whether the toolchain does is settled by experiment on a throwaway crate, not by asking it — 1.100.0-nightly accepts `-Z checksum-freshness` and rebuilds on mtime anyway. The experiment reports **three** outcomes, not two: active, measured-inactive, and *not measured*. Its answer codes are `0` and `3`, deliberately clear of every status bash generates for its own errors — so nothing that goes wrong inside the probe, including an expansion failure no guard can catch, can be read as an answer. The scenario is skipped for the last two alike, but a failure to measure is never reported as a measurement. The control also reports which mutation families the running Cargo exhibits — a note, not an assertion, since that set moves upstream. |
| `seed-target-dir-selftest.sh` | seed-source preference, lock-file stripping, two jobs racing on one cache key, **and one scenario per check a hardlink clone is validated against**: a source rotated wholesale, a subtree silently lost from the walk, a copy that reports failure over a tree both other checks read as whole, and a source identity that resolved at neither end — plus a staging tree that could not be privately owned being discarded rather than published, and the publisher's log showing it waited on the consumer's own reader-lock marker before reclaiming a rotated snapshot |
| `publish-snapshot-selftest.sh` | the atomic swap, that a live consumer survives a republish, and the publisher's side of the rotation race: deferred reclamation under a live reader, and its sweep once the reader is gone | | `publish-snapshot-selftest.sh` | the atomic swap, that a live consumer survives a republish, and the publisher's side of the rotation race: deferred reclamation under a live reader, and its sweep once the reader is gone |
| `prune-cache-selftest.sh` | liveness, protection, locking, eviction order, self-clear, **and that a cache a job claims *inside* the check-to-unlink window survives it** — against a real scratch `origin` | | `prune-cache-selftest.sh` | liveness, protection, locking, eviction order, self-clear, **and that a cache a job claims *inside* the check-to-unlink window survives it** — against a real scratch `origin` |
| `restore-mtimes-selftest.sh` | the merge hazard and the watermark that closes it, including the two-jobs-one-namespace case. Needs a real compiler. | | `restore-mtimes-selftest.sh` | the merge hazard and the watermark that closes it, including the two-jobs-one-namespace case. Needs a real compiler. |
@@ -472,7 +583,7 @@ key and asserts only what must be true whichever of them wins the rename.
which places the interference inside the window rather than hoping it lands which places the interference inside the window rather than hoping it lands
there. `prune-cache-selftest.sh` scenario 12 stubs `du`, so the pass's own there. `prune-cache-selftest.sh` scenario 12 stubs `du`, so the pass's own
measurement publishes a reader marker strictly between its check and its measurement publishes a reader marker strictly between its check and its
unlink. `seed-target-dir-selftest.sh` uses the shape four times over, on two unlink. `seed-target-dir-selftest.sh` uses the shape five times over, on two
different commands. Scenarios 8a, 8b and 8c stub `cp`, so the consumer's own different commands. Scenarios 8a, 8b and 8c stub `cp`, so the consumer's own
clone is what rotates the snapshot underneath it, loses a subtree of its own clone is what rotates the snapshot underneath it, loses a subtree of its own
source, or reports a failure over a tree that is in fact whole — each strictly source, or reports a failure over a tree that is in fact whole — each strictly
@@ -483,9 +594,14 @@ could not be READ rather than on one that changed, and the only way to make
that the sole witness is to fail the identity reads while the copy between that the sole witness is to fail the identity reads while the copy between
them succeeds. 8a also starts a second real process — the actual them succeeds. 8a also starts a second real process — the actual
`publish-snapshot.sh` — but the stub is what fixes where its swap lands; the `publish-snapshot.sh` — but the stub is what fixes where its swap lands; the
concurrency is incidental to the determinism. Every stub asserts that it concurrency is incidental to the determinism. Scenario 11 reuses 8a's exact
fired, because a scenario whose interference silently did not happen passes stub and the same forced rotation, but reads a different witness: not the
for the wrong reason. consumer's own checks, but a line in the *publisher's* log reporting that it
waited on the reader marker this consumer's clone wrote — `reader_lock_acquire`
is exercised by 8a already, but nothing asserts it actually fired until this
scenario reads that line back. Every stub asserts that it fired, because a
scenario whose interference silently did not happen passes for the wrong
reason.
**A synthetic stand-in for the other side, where that artefact *is* the **A synthetic stand-in for the other side, where that artefact *is* the
contract.** `publish-snapshot-selftest.sh` scenarios 6 to 8 hold a contract.** `publish-snapshot-selftest.sh` scenarios 6 to 8 hold a
@@ -517,6 +633,15 @@ how both `[ "$cp_rc" -eq 0 ]` and `[ "$i_before" != missing ]` sat unpinned
(issue #5) while looking well covered. Isolating a check means constructing (issue #5) while looking well covered. Isolating a check means constructing
the state only it can see, not the state that trips several at once. the state only it can see, not the state that trips several at once.
Scenario 11 applies the same discipline to a witness outside the clone
entirely: not which of several checks inside `hardlink_clone_into` caught a
fault, but whether `reader_lock_acquire`'s marker was observed by anything
outside it at all. The consumer's own log and exit status are silent either
way — a run with the marker deleted still succeeds — so what is asserted is
one line in the *publisher's* log reporting that it waited. Deleting
`reader_lock_acquire` (issue #10) leaves that line unwritten without failing
anything else in the suite.
The action YAML holds no logic beyond wiring; everything testable lives in The action YAML holds no logic beyond wiring; everything testable lives in
`scripts/`. A composite action needs `shell: bash` on every `run:` step, and `scripts/`. A composite action needs `shell: bash` on every `run:` step, and
the actions reach their shared scripts through the actions reach their shared scripts through
+27 -1
View File
@@ -10,6 +10,15 @@ inputs:
description: 'Mount point of the persistent cache volume. Must match the consume action.' description: 'Mount point of the persistent cache volume. Must match the consume action.'
required: false required: false
default: '/cache' default: '/cache'
cache-lineage:
description: >-
Must match the cargo-cache step in this job. Checked rather than
assumed: this action derives both ends of the snapshot swap from its own
cache-root, so a publish step left at the default while its consume step
nested would republish a DIFFERENT lineage's live target dir over that
lineage's snapshot, on every push, silently. A mismatch fails the step.
required: false
default: ''
protected-branches: protected-branches:
description: >- description: >-
Space-separated refs that publish snapshots. A run whose own ref is not Space-separated refs that publish snapshots. A run whose own ref is not
@@ -67,6 +76,11 @@ runs:
# Everything here reads the environment the consume action exported, so a # Everything here reads the environment the consume action exported, so a
# workflow that forgets to run cargo-cache first fails loudly here rather # workflow that forgets to run cargo-cache first fails loudly here rather
# than silently publishing a snapshot of the wrong directory. # than silently publishing a snapshot of the wrong directory.
#
# The cache root is taken from that same environment for the same reason,
# and this action's own cache-root/cache-lineage inputs are checked against
# it rather than used. The two have always had to agree; until a lineage
# existed they always did, because nobody overrode the default.
- id: resolve - id: resolve
shell: bash shell: bash
env: env:
@@ -92,6 +106,18 @@ runs:
: "${CARGO_CACHE_KEY:?cargo-cache-publish: CARGO_CACHE_KEY not set by the cargo-cache action}" : "${CARGO_CACHE_KEY:?cargo-cache-publish: CARGO_CACHE_KEY not set by the cargo-cache action}"
echo "active=yes" >> "$GITHUB_OUTPUT" echo "active=yes" >> "$GITHUB_OUTPUT"
# Only in publish mode. The `release-lock` call is an `if: always()`
# step that consuming workflows invoke with `mode:` and nothing else —
# it releases $CARGO_CACHE_LOCK_ID on $CARGO_TARGET_DIR and never
# touches a cache root at all, so holding its default inputs to the
# consume step's would fail the cleanup step of every job that sets a
# lineage, for a value it does not use.
if [ "${{ inputs.mode }}" = "publish" ]; then
: "${CARGO_CACHE_ROOT:?cargo-cache-publish: CARGO_CACHE_ROOT not set by the cargo-cache action}"
bash "${CARGO_CACHE_SCRIPTS}/cache-root.sh" verify \
"${{ inputs.cache-root }}" "${{ inputs.cache-lineage }}" "$CARGO_CACHE_ROOT"
fi
OWN_REF="${{ inputs.own-ref }}" OWN_REF="${{ inputs.own-ref }}"
[ -n "$OWN_REF" ] || OWN_REF="${{ github.head_ref || github.ref_name }}" [ -n "$OWN_REF" ] || OWN_REF="${{ github.head_ref || github.ref_name }}"
@@ -131,7 +157,7 @@ runs:
run: | run: |
set -euo pipefail set -euo pipefail
bash "${CARGO_CACHE_SCRIPTS}/publish-snapshot.sh" \ bash "${CARGO_CACHE_SCRIPTS}/publish-snapshot.sh" \
"$CARGO_CACHE_KEY" "${{ inputs.cache-root }}" \ "$CARGO_CACHE_KEY" "$CARGO_CACHE_ROOT" \
"${{ github.job }}-${{ github.run_id }}-$$" "${{ github.job }}-${{ github.run_id }}-$$"
# Released in both modes. In `publish` mode this is the normal end-of-job # Released in both modes. In `publish` mode this is the normal end-of-job
+38 -6
View File
@@ -10,6 +10,18 @@ inputs:
description: 'Mount point of the persistent cache volume inside the job container.' description: 'Mount point of the persistent cache volume inside the job container.'
required: false required: false
default: '/cache' default: '/cache'
cache-lineage:
description: >-
Distinguishes two jobs that build the SAME ref for different targets or
profiles (a host build and a wasm32 build, say) and would otherwise
resolve to one CARGO_TARGET_DIR and serialise on Cargo's exclusive
build-directory lock. Names one directory level under cache-root:
<cache-root>/<lineage>/target-<key>. Must be a single path component;
several names are refused outright because a reader elsewhere would stop
seeing the caches under them (see validate_cache_lineage in
scripts/cache-lib.sh). Pass the same value to cargo-cache-publish.
required: false
default: ''
protected-branches: protected-branches:
description: >- description: >-
Space-separated refs that publish snapshots and are never evicted. Space-separated refs that publish snapshots and are never evicted.
@@ -54,8 +66,9 @@ inputs:
watermark-file: watermark-file:
description: >- description: >-
Name of this job's build-watermark file inside the target dir. MUST be Name of this job's build-watermark file inside the target dir. MUST be
distinct per job when two jobs share one cache key. Defaults to distinct per job when two jobs share one target directory — which two
.ci-watermark-<job>-sha. jobs no longer need to do; cache-lineage gives them separate ones.
Defaults to .ci-watermark-<job>-sha, already distinct per job.
required: false required: false
default: '' default: ''
lock-id: lock-id:
@@ -74,6 +87,12 @@ outputs:
cache-key: cache-key:
description: 'Sanitized cache key for this run''s own ref.' description: 'Sanitized cache key for this run''s own ref.'
value: ${{ steps.resolve.outputs.cache-key }} value: ${{ steps.resolve.outputs.cache-key }}
cache-root:
description: >-
Resolved cache root — cache-root, plus the lineage directory when one is
set. Every directory this action reads or writes is under it. Also
exported as CARGO_CACHE_ROOT.
value: ${{ steps.resolve.outputs.cache-root }}
seeded-from: seeded-from:
description: 'Where the target dir came from: own | base-snapshot | own-snapshot | fallback-dir | concurrent-peer | cold.' description: 'Where the target dir came from: own | base-snapshot | own-snapshot | fallback-dir | concurrent-peer | cold.'
value: ${{ steps.seed.outputs.seeded-from }} value: ${{ steps.seed.outputs.seeded-from }}
@@ -95,6 +114,16 @@ runs:
# `base_ref` is populated only for pull_request events. A push run has # `base_ref` is populated only for pull_request events. A push run has
# nothing to layer over: its own ref IS the reference branch. It # nothing to layer over: its own ref IS the reference branch. It
# publishes, it does not consume. # publishes, it does not consume.
#
# The cache root is resolved first because every path below hangs off it.
# A lineage nests one directory level (`<root>/<lineage>/target-<key>`),
# which is what lets two jobs on ONE ref hold two build directories and so
# not serialise on Cargo's exclusive lock. It is resolved through
# cache-root.sh rather than interpolated here so the name is validated —
# several otherwise-reasonable lineage names put their whole subtree out of
# reach of a pass that has to see it. Everything downstream reads the
# resolved value, and it is exported as CARGO_CACHE_ROOT so the publish
# action can check it agrees with its own inputs.
- id: resolve - id: resolve
shell: bash shell: bash
run: | run: |
@@ -102,6 +131,8 @@ runs:
SCRIPTS=$(cd "${{ github.action_path }}/.." && pwd)/scripts SCRIPTS=$(cd "${{ github.action_path }}/.." && pwd)/scripts
[ -d "$SCRIPTS" ] || { echo "::error::cargo-cache: scripts/ not found at $SCRIPTS"; exit 1; } [ -d "$SCRIPTS" ] || { echo "::error::cargo-cache: scripts/ not found at $SCRIPTS"; exit 1; }
echo "CARGO_CACHE_SCRIPTS=${SCRIPTS}" >> "$GITHUB_ENV" echo "CARGO_CACHE_SCRIPTS=${SCRIPTS}" >> "$GITHUB_ENV"
CACHE_ROOT=$(bash "${SCRIPTS}/cache-root.sh" resolve \
"${{ inputs.cache-root }}" "${{ inputs.cache-lineage }}")
OWN_REF="${{ inputs.own-ref }}" OWN_REF="${{ inputs.own-ref }}"
[ -n "$OWN_REF" ] || OWN_REF="${{ github.head_ref || github.ref_name }}" [ -n "$OWN_REF" ] || OWN_REF="${{ github.head_ref || github.ref_name }}"
BASE_REF="${{ inputs.base-ref }}" BASE_REF="${{ inputs.base-ref }}"
@@ -116,7 +147,7 @@ runs:
echo "cache: own ref '${OWN_REF}' -> ${OWN_KEY} (no base ref — this ref publishes, it does not consume)" echo "cache: own ref '${OWN_REF}' -> ${OWN_KEY} (no base ref — this ref publishes, it does not consume)"
fi fi
TARGET_DIR="${{ inputs.cache-root }}/target-${OWN_KEY}" TARGET_DIR="${CACHE_ROOT}/target-${OWN_KEY}"
WATERMARK="${{ inputs.watermark-file }}" WATERMARK="${{ inputs.watermark-file }}"
[ -n "$WATERMARK" ] || WATERMARK=".ci-watermark-${{ github.job }}-sha" [ -n "$WATERMARK" ] || WATERMARK=".ci-watermark-${{ github.job }}-sha"
LOCK_ID="${{ inputs.lock-id }}" LOCK_ID="${{ inputs.lock-id }}"
@@ -128,10 +159,11 @@ runs:
echo "base-key=${BASE_KEY}" echo "base-key=${BASE_KEY}"
echo "lock-id=${LOCK_ID}" echo "lock-id=${LOCK_ID}"
echo "watermark-file=${WATERMARK}" echo "watermark-file=${WATERMARK}"
echo "cache-root=${CACHE_ROOT}"
} >> "$GITHUB_OUTPUT" } >> "$GITHUB_OUTPUT"
{ {
echo "CARGO_TARGET_DIR=${TARGET_DIR}" echo "CARGO_TARGET_DIR=${TARGET_DIR}"
echo "CARGO_CACHE_ROOT=${{ inputs.cache-root }}" echo "CARGO_CACHE_ROOT=${CACHE_ROOT}"
echo "CARGO_CACHE_KEY=${OWN_KEY}" echo "CARGO_CACHE_KEY=${OWN_KEY}"
echo "CARGO_CACHE_LOCK_ID=${LOCK_ID}" echo "CARGO_CACHE_LOCK_ID=${LOCK_ID}"
echo "CI_WATERMARK_FILE=${WATERMARK}" echo "CI_WATERMARK_FILE=${WATERMARK}"
@@ -155,7 +187,7 @@ runs:
bash "${SCRIPTS}/seed-target-dir.sh" \ bash "${SCRIPTS}/seed-target-dir.sh" \
"${{ steps.resolve.outputs.cache-key }}" \ "${{ steps.resolve.outputs.cache-key }}" \
"${{ steps.resolve.outputs.base-key }}" \ "${{ steps.resolve.outputs.base-key }}" \
"${{ inputs.cache-root }}" \ "${{ steps.resolve.outputs.cache-root }}" \
"${{ github.job }}-${{ github.run_id }}-$$" \ "${{ github.job }}-${{ github.run_id }}-$$" \
"${{ inputs.seed-fallback-dir }}" \ "${{ inputs.seed-fallback-dir }}" \
"${{ steps.resolve.outputs.lock-id }}" "${{ steps.resolve.outputs.lock-id }}"
@@ -206,7 +238,7 @@ runs:
set -euo pipefail set -euo pipefail
SCRIPTS=$(cd "${{ github.action_path }}/.." && pwd)/scripts SCRIPTS=$(cd "${{ github.action_path }}/.." && pwd)/scripts
bash "${SCRIPTS}/prune-cache.sh" \ bash "${SCRIPTS}/prune-cache.sh" \
"${{ inputs.cache-root }}" \ "${{ steps.resolve.outputs.cache-root }}" \
"${{ steps.resolve.outputs.target-dir }}" \ "${{ steps.resolve.outputs.target-dir }}" \
"${{ inputs.protected-branches }}" \ "${{ inputs.protected-branches }}" \
"${{ inputs.min-free-percent }}" "${{ inputs.min-free-percent }}"
+114 -1
View File
@@ -114,6 +114,106 @@ cache_key() {
target_dir_for() { printf '%s/target-%s' "$1" "$2"; } target_dir_for() { printf '%s/target-%s' "$1" "$2"; }
snapshot_dir_for() { printf '%s/snapshot-%s' "$1" "$2"; } snapshot_dir_for() { printf '%s/snapshot-%s' "$1" "$2"; }
# ---------------------------------------------------------------------------
# Cache lineages
# ---------------------------------------------------------------------------
#
# A cache key names a REF. What a target directory holds is the product of a
# ref and a BUILD CONFIGURATION, and the two are not the same thing: emowheel
# builds the same ref twice on every push, once for the host and once for
# wasm32, in two jobs that run concurrently. Keyed on the ref alone both
# resolve to one CARGO_TARGET_DIR, and Cargo's build-directory lock is
# exclusive — so the second job sits on `Blocking waiting for file lock on
# build directory` for the length of the first, holding a runner capacity slot
# while doing nothing (daniel/gitdan#60).
#
# A lineage is that second dimension, and it is expressed as ONE DIRECTORY
# LEVEL above the per-ref directories rather than as a suffix on the key:
#
# <cache-root>/target-<key> no lineage (the flat layout)
# <cache-root>/<lineage>/target-<key> a lineage
#
# Nesting rather than suffixing is what keeps every existing reader correct
# without teaching any of them a new name shape. `prune-cache.sh` resolves
# liveness by recomputing `target-<cache_key(branch)>` for every branch on
# origin and evicting whatever does not match — a suffixed `target-<key>-wasm32`
# matches nothing, so it would be classified dead and unconditionally evicted
# on every single run. The host-level arbiter in daniel/gitdan reads the same
# shape (its BRANCH_DIR_RE), and a suffixed name falls out of it too: not
# evicted there, but never a candidate either, so a whole lineage becomes
# invisible to the global disk budget. Nesting leaves both matchers reading
# exactly the names they already read, one directory deeper.
#
# ONE LEVEL, AND NOT TWO. The arbiter walks a volume to CI_CACHE_MAX_DEPTH,
# which is 2 — `_data/target-<key>` and `_data/<lineage>/target-<key>`. It is
# kept tight there on purpose (a deeper walk starts meeting Cargo's own
# `incremental/<crate>-<hash>` directories, which match the same name shape and
# must never be evicted individually), so a lineage is a single path component
# and validate_cache_lineage refuses one containing a slash.
# Directory names daniel/gitdan's ci-cache-reclaim.sh refuses to descend into
# (its CI_CACHE_NODESCEND_NAMES). A lineage named one of these puts its whole
# subtree outside the global arbiter's reach: the caches accumulate and the one
# script whose job is the shared disk budget cannot see them.
CACHE_LINEAGE_RESERVED_NAMES="debug release deps incremental build .fingerprint tmp examples doc"
# The shape that same script reads as a per-branch cache directory (its
# BRANCH_DIR_RE). A lineage matching it is taken for a cache dir in its own
# right — never descended into, and an eviction candidate whole, which is the
# entire lineage rather than one ref's share of it.
CACHE_LINEAGE_BRANCH_DIR_RE='^.+-[0-9a-f]{7,40}$'
# Every rejection below names the reader that imposes it, because that is the
# only way the constraint survives: none of these is a filesystem limit, and a
# name that trips one produces no error anywhere — it produces a lineage that
# silently stops being pruned, or silently stops being reclaimed.
validate_cache_lineage() {
local lineage="$1" reserved
[ -n "$lineage" ] || return 0
case "$lineage" in
*/*)
echo "::error::cache lineage '${lineage}' must be a single path component: the host-level arbiter walks a cache volume to depth 2, so <cache-root>/<lineage>/target-<key> is as deep as a cache directory may sit and still be reclaimable" >&2
return 1
;;
.*)
echo "::error::cache lineage '${lineage}' must not start with a dot: every dot-prefixed entry under a cache root belongs to the leftover-naming contract (see the top of this file), and a lineage is not garbage to be reclaimed" >&2
return 1
;;
target-* | snapshot-*)
echo "::error::cache lineage '${lineage}' must not start with 'target-' or 'snapshot-': prune-cache.sh globs both prefixes at the cache root, so the lineage directory itself would become an eviction candidate" >&2
return 1
;;
*[!A-Za-z0-9._-]*)
echo "::error::cache lineage '${lineage}' may contain only [A-Za-z0-9._-] — the same charset cache_key() sanitises a ref down to" >&2
return 1
;;
esac
for reserved in $CACHE_LINEAGE_RESERVED_NAMES; do
if [ "$lineage" = "$reserved" ]; then
echo "::error::cache lineage '${lineage}' is one of the Cargo directory names daniel/gitdan's ci-cache-reclaim.sh never descends into (CI_CACHE_NODESCEND_NAMES) — every cache under it would be invisible to the host-level disk budget" >&2
return 1
fi
done
if [[ $lineage =~ $CACHE_LINEAGE_BRANCH_DIR_RE ]]; then
echo "::error::cache lineage '${lineage}' ends in a hex suffix, which is the shape daniel/gitdan's ci-cache-reclaim.sh reads as a per-branch cache directory — it would treat the lineage directory as one cache and evict the whole thing" >&2
return 1
fi
return 0
}
# The cache root a lineage's directories actually live under. An empty lineage
# resolves to the cache root unchanged, byte for byte: that is what makes this
# a no-op for every consumer that does not set one, rather than a migration.
cache_root_for() {
local root="$1" lineage="${2:-}"
validate_cache_lineage "$lineage" || return 1
printf '%s%s' "$root" "${lineage:+/$lineage}"
}
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
# Disk accounting # Disk accounting
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
@@ -210,6 +310,9 @@ _unshare_files() {
# The inner shell propagates a failure of any individual copy-and-rename out # The inner shell propagates a failure of any individual copy-and-rename out
# through xargs (which exits 123 if any invocation exits 1-125), so a # through xargs (which exits 123 if any invocation exits 1-125), so a
# partially-unshared tree is reported rather than silently accepted. # partially-unshared tree is reported rather than silently accepted.
# shellcheck disable=SC2016 # the quoted program is for the INNER shell: $f
# is its loop variable, $rc its accumulator and $$ its pid. Expanding any of
# them here is what the single quotes exist to prevent.
find "$@" -links +1 -print0 2>/dev/null | find "$@" -links +1 -print0 2>/dev/null |
xargs -0 -r -n 64 bash -c 'rc=0; for f; do cp -p -- "$f" "$f.unshare.$$" && mv -f -- "$f.unshare.$$" "$f" || rc=1; done; exit $rc' _ xargs -0 -r -n 64 bash -c 'rc=0; for f; do cp -p -- "$f" "$f.unshare.$$" && mv -f -- "$f.unshare.$$" "$f" || rc=1; done; exit $rc' _
} }
@@ -229,7 +332,17 @@ _unshare_files() {
# <profile>/.fingerprint/<unit>/dep-<target> (only under # <profile>/.fingerprint/<unit>/dep-<target> (only under
# CARGO_UNSTABLE_CHECKSUM_FRESHNESS, # CARGO_UNSTABLE_CHECKSUM_FRESHNESS,
# where this file carries the # where this file carries the
# per-source blake3 checksums) # per-source blake3 checksums.
# NOT reproduced on
# 1.100.0-nightly (2026-08-25),
# measured by this repo's own CI
# — upstream appears to have
# stopped writing it in place.
# Kept in the unshared set
# anyway: it costs 22 MB of a
# 6.9 GB tree, and the failure
# it guards is a wrong answer,
# not a slow one.)
# <profile>/build/<pkg>/output, root-output (Cargo build-script metadata) # <profile>/build/<pkg>/output, root-output (Cargo build-script metadata)
# <profile>/build/<pkg>/out/** (whatever the build script # <profile>/build/<pkg>/out/** (whatever the build script
# writes into OUT_DIR — build # writes into OUT_DIR — build
+194
View File
@@ -0,0 +1,194 @@
#!/usr/bin/env bash
# Regression test for cache-root.sh and the lineage rules in cache-lib.sh —
# the fix for daniel/gitdan#60, where two jobs building the same ref for
# different targets resolved to one CARGO_TARGET_DIR and serialised on Cargo's
# exclusive build-directory lock.
#
# 1. NO LINEAGE CHANGES NOTHING — the effective root is the cache root byte
# for byte, so every consumer that does not set a lineage keeps the exact
# directories it already has on the volume. This is the whole of the
# migration story for lublub, zemyna and emowheel's `ci` job, so it is
# asserted rather than assumed.
# 2. TWO LINEAGES, ONE REF, TWO TARGET DIRS — the bug itself. The two jobs
# keep one cache key (they are the same ref) and still get directories
# that are neither equal nor nested one inside the other, which is what
# Cargo's per-directory lock needs in order not to serialise them.
# 3. THE WHOLE PIPELINE MOVES TOGETHER — seed, publish and prune all operate
# inside the lineage root. A pass run in one lineage must not evict, or
# even see, a sibling lineage's caches or the flat layout's.
# 4. BASE SEEDING IS PER LINEAGE — a PR branch layers over ITS OWN lineage's
# base snapshot, not over whatever the flat root happens to hold. This is
# the property that keeps a warm start for both jobs rather than one.
# 5. A NAME NO READER CAN HANDLE IS REFUSED AT RESOLVE TIME — one assertion
# per constraint, each named for the reader that imposes it. None of these
# is a filesystem limit: every one of them produces a working directory
# that some pass silently stops seeing, which is the failure mode this
# whole scheme exists to avoid rather than to relocate.
# 6. A PUBLISH THAT DISAGREES WITH ITS CONSUME STEP FAILS LOUDLY — the
# footgun the lineage input introduces. cargo-cache-publish derives both
# ends of the snapshot swap from its own `cache-root`, so a publish step
# left at the default while its consume step nested would republish the
# OTHER lineage's live target dir over that lineage's snapshot, silently.
set -euo pipefail
script_dir=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)
. "$script_dir/cache-lib.sh"
scratch=$(mktemp -d)
trap 'rm -rf "$scratch"' EXIT
root="$scratch/cache"; mkdir -p "$root"
pass_count=0
fail() { echo "ASSERTION FAILED: $*" >&2; exit 1; }
ok() { pass_count=$((pass_count + 1)); echo "PASS: $*"; }
resolve() { bash "$script_dir/cache-root.sh" resolve "$@"; }
# A tree that looks enough like a Cargo target dir for the pipeline scripts,
# with a marker naming which lineage produced it — scenario 4 turns on reading
# that marker back out of a seeded directory.
make_tree() {
local d="$1" marker="$2"
mkdir -p "$d/debug/deps" "$d/debug/.fingerprint/x"
echo "$marker" > "$d/debug/deps/libx.rlib"
echo "$marker" > "$d/lineage-marker"
echo "$marker" > "$d/debug/.fingerprint/x/dep-lib-x"
}
echo "=== 1. no lineage changes nothing ==="
[ "$(resolve /cache)" = /cache ] || fail "an omitted lineage changed the root"
[ "$(resolve /cache '')" = /cache ] || fail "an empty lineage changed the root"
ok "no lineage resolves to the cache root unchanged"
KEY=$(cache_key dev)
[ "$(target_dir_for "$(resolve /cache '')" "$KEY")" = "/cache/target-$KEY" ] \
|| fail "the flat target-dir name moved"
[ "$(snapshot_dir_for "$(resolve /cache '')" "$KEY")" = "/cache/snapshot-$KEY" ] \
|| fail "the flat snapshot name moved"
ok "the flat layout's directory names are untouched"
echo
echo "=== 2. two lineages, one ref, two target dirs ==="
HOST_ROOT=$(resolve /cache)
WASM_ROOT=$(resolve /cache wasm32)
[ "$WASM_ROOT" = /cache/wasm32 ] || fail "lineage root resolved to '$WASM_ROOT'"
HOST_DIR=$(target_dir_for "$HOST_ROOT" "$KEY")
WASM_DIR=$(target_dir_for "$WASM_ROOT" "$KEY")
[ "$HOST_DIR" != "$WASM_DIR" ] || fail "both lineages resolved to $HOST_DIR"
case "$WASM_DIR" in "$HOST_DIR"/*) fail "the wasm target dir sits inside the host one" ;; esac
case "$HOST_DIR" in "$WASM_DIR"/*) fail "the host target dir sits inside the wasm one" ;; esac
ok "one cache key ($KEY), two disjoint target dirs: $HOST_DIR and $WASM_DIR"
echo
echo "=== 3. the whole pipeline moves together ==="
DEAD=$(cache_key feat/dead)
lin="$root/wasm32"
mkdir -p "$lin"
make_tree "$root/target-$DEAD" flat
make_tree "$root/native/target-$DEAD" native
make_tree "$lin/target-$DEAD" wasm32
bash "$script_dir/seed-target-dir.sh" "$KEY" "" "$lin" tag-3 > "$scratch/seed3.log" 2>&1 \
|| { cat "$scratch/seed3.log"; fail "seed inside a lineage root failed"; }
[ -d "$lin/target-$KEY" ] || fail "seed did not create $lin/target-$KEY"
[ -d "$root/target-$KEY" ] && fail "seed created a directory at the flat root as well"
ok "seed creates its directory under the lineage root and nowhere else"
make_tree "$lin/target-$KEY" wasm32
bash "$script_dir/publish-snapshot.sh" "$KEY" "$lin" tag-3 > "$scratch/pub3.log" 2>&1 \
|| { cat "$scratch/pub3.log"; fail "publish inside a lineage root failed"; }
[ -d "$lin/snapshot-$KEY" ] || fail "publish did not create $lin/snapshot-$KEY"
[ -d "$root/snapshot-$KEY" ] && fail "publish created a snapshot at the flat root as well"
ok "publish writes its snapshot under the lineage root and nowhere else"
# Free space far below the threshold, so pass 2 evicts every eligible
# directory it can see. What it can see is the point of the scenario.
CACHE_LIVENESS=false CACHE_DF_OVERRIDE="1000000 1000" \
bash "$script_dir/prune-cache.sh" "$lin" "$lin/target-$KEY" 'dev main' 10 \
> "$scratch/prune3.log" 2>&1 || { cat "$scratch/prune3.log"; fail "prune inside a lineage root failed"; }
[ -d "$lin/target-$DEAD" ] && { cat "$scratch/prune3.log"; fail "prune left its own lineage's evictable cache in place"; }
[ -d "$root/target-$DEAD" ] || fail "prune reached out of its lineage and evicted the flat root's cache"
[ -d "$root/native/target-$DEAD" ] || fail "prune reached into a sibling lineage and evicted its cache"
ok "prune under disk pressure evicts inside its own lineage only"
echo
echo "=== 4. base seeding is per lineage ==="
BASE=$(cache_key dev)
PR=$(cache_key feat/pr)
rm -rf "$lin" "$root/snapshot-$BASE"
mkdir -p "$lin"
make_tree "$root/snapshot-$BASE" flat-base
make_tree "$lin/snapshot-$BASE" wasm32-base
bash "$script_dir/seed-target-dir.sh" "$PR" "$BASE" "$lin" tag-4 > "$scratch/seed4.log" 2>&1 \
|| { cat "$scratch/seed4.log"; fail "seeding a PR branch inside a lineage failed"; }
[ -d "$lin/target-$PR" ] || fail "the PR branch's lineage target dir was not created"
got=$(cat "$lin/target-$PR/lineage-marker")
[ "$got" = wasm32-base ] || fail "the PR branch layered over '$got', not its own lineage's base snapshot"
grep -q 'base snapshot' "$scratch/seed4.log" || { cat "$scratch/seed4.log"; fail "seed did not report a base-snapshot clone"; }
ok "a PR branch layers over its own lineage's base snapshot ($got)"
echo
echo "=== 5. a name no reader can handle is refused ==="
reject() {
local lineage="$1" want="$2" desc="$3" out
if out=$(resolve /cache "$lineage" 2>&1); then
fail "lineage '$lineage' was accepted (resolved to '$out') — $desc"
fi
case "$out" in
*"$want"*) ;;
*) fail "lineage '$lineage' was rejected without naming '$want': $out" ;;
esac
ok "rejected '$lineage' — $desc"
}
reject 'a/b' 'single path component' "the arbiter walks a volume to depth 2"
reject '.hidden' 'dot' "every dot-prefixed name under a cache root is leftover-contract territory"
reject 'target-x' 'prune-cache.sh' "prune-cache.sh globs target-* at the cache root"
reject 'snapshot-x' 'prune-cache.sh' "prune-cache.sh globs snapshot-* at the cache root"
reject 'wasm 32' 'A-Za-z0-9._-' "a cache key is sanitised to that charset and a lineage sits beside one"
reject 'release' 'CI_CACHE_NODESCEND_NAMES' "the arbiter never descends into a Cargo profile name"
reject 'doc' 'CI_CACHE_NODESCEND_NAMES' "same, for the docs profile directory"
reject 'lineage-deadbeef' 'per-branch cache directory' "the arbiter reads a hex-suffixed name as one cache dir"
for good in wasm32 web android host wasm32.release lineage_2; do
out=$(resolve /cache "$good") || fail "lineage '$good' was rejected: $out"
[ "$out" = "/cache/$good" ] || fail "lineage '$good' resolved to '$out'"
done
ok "ordinary lineage names still resolve"
echo
echo "=== 6. a publish that disagrees with its consume step fails loudly ==="
verify() { bash "$script_dir/cache-root.sh" verify "$@"; }
verify /cache wasm32 /cache/wasm32 > /dev/null 2>&1 \
|| fail "verify rejected a publish step that agrees with its consume step"
verify /cache '' /cache > /dev/null 2>&1 \
|| fail "verify rejected an unmigrated consumer's matching default pair"
ok "verify accepts a publish step whose inputs match what the consume step exported"
# The exact shape of the mistake: the consume step nested, the publish step
# kept the default. Left unchecked this republishes the host lineage's live
# target dir over the host lineage's snapshot, from the wasm job.
if out=$(verify /cache '' /cache/wasm32 2>&1); then
fail "verify accepted a publish step that resolved to /cache while the job exported /cache/wasm32"
fi
case "$out" in
*'/cache/wasm32'*) ;;
*) fail "the mismatch error does not quote what the consume step exported: $out" ;;
esac
case "$out" in
*'SAME cache-root and cache-lineage'*) ;;
*) fail "the mismatch error does not say what to do about it: $out" ;;
esac
ok "verify rejects a publish step that forgot the lineage, and says so"
if verify /cache 'a/b' /cache/a/b > /dev/null 2>&1; then
fail "verify accepted an invalid lineage as long as both sides agreed on it"
fi
ok "verify validates the lineage as well as comparing it"
echo
echo "cache-root-selftest: ${pass_count} assertions passed"
+61
View File
@@ -0,0 +1,61 @@
#!/usr/bin/env bash
# Resolves — and cross-checks — the cache root a job's directories live under.
#
# cache-root.sh resolve <cache-root> [lineage]
# cache-root.sh verify <cache-root> <lineage> <exported-root>
#
# `resolve` prints the effective root: the cache root unchanged when no lineage
# is given, or `<cache-root>/<lineage>` when one is. Invalid lineage names are
# rejected here rather than downstream — see validate_cache_lineage() in
# cache-lib.sh, where every rejection names the reader that imposes it.
#
# `verify` is the publish side's guard. cargo-cache-publish resolves the same
# two inputs the consume action was given and compares the result against the
# CARGO_CACHE_ROOT the consume step exported into the job environment. The two
# actions have always had to agree — `cache-root`'s description in the publish
# action says "must match the consume action" — and until a lineage existed
# they always did, because nobody overrode the default. A disagreement is not
# a harmless no-op: publish-snapshot.sh takes the root as an argument and
# derives BOTH ends of the swap from it, so a publish step that kept the
# default while its consume step nested would read `<root>/target-<key>` — the
# OTHER lineage's live target dir — and republish it over `<root>/snapshot-<key>`,
# which is that lineage's snapshot. Two jobs would then be publishing one
# snapshot from one tree on every push, and nothing in either action would say
# so. Hence: fail the job, loudly, rather than resolve the ambiguity in
# either direction.
#
# A thin CLI over cache-lib.sh, kept as its own entry point for the same
# reason branch-cache-key.sh is: an out-of-band job that needs to find a
# lineage's directories should resolve the path the way the action does
# instead of reimplementing the rule.
set -euo pipefail
. "$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/cache-lib.sh"
MODE="${1:-}"
case "$MODE" in
resolve)
[ $# -ge 2 ] && [ $# -le 3 ] || {
echo "::error::cache-root.sh resolve: expected <cache-root> [lineage]" >&2
exit 1
}
[ -n "$2" ] || { echo "::error::cache-root.sh: cache-root must not be empty" >&2; exit 1; }
cache_root_for "$2" "${3:-}"
;;
verify)
[ $# -eq 4 ] || {
echo "::error::cache-root.sh verify: expected <cache-root> <lineage> <exported-root>" >&2
exit 1
}
[ -n "$2" ] || { echo "::error::cache-root.sh: cache-root must not be empty" >&2; exit 1; }
expected=$(cache_root_for "$2" "$3")
if [ "$expected" != "$4" ]; then
echo "::error::cache-root.sh: this step resolves its cache root to '${expected}' (cache-root '$2', cache-lineage '$3') but the cargo-cache step in this job exported '$4'. Pass the SAME cache-root and cache-lineage to both actions." >&2
exit 1
fi
echo "cache root: ${expected} (agrees with the cargo-cache step in this job)"
;;
*)
echo "::error::cache-root.sh: unknown mode '${MODE}' (expected resolve or verify)" >&2
exit 1
;;
esac
+185 -24
View File
@@ -73,22 +73,154 @@ mkcrate "$crate_dir"
cd "$crate_dir" cd "$crate_dir"
export CARGO_INCREMENTAL=0 export CARGO_INCREMENTAL=0
# Every assertion below reads cargo's own words out of a build log
# (`Compiling libdep`, `Fresh probe`). A CI image that forces colour splices an
# ANSI reset between the status word and the crate name, at which point every
# one of those greps silently stops matching and the suite reports the
# opposite of what happened — observed on gitdan-ci's runner image, where
# scenario 2 failed while the log it printed plainly showed `Compiling libdep`.
# Pin the format the assertions are written against.
export CARGO_TERM_COLOR=never
# Checksum freshness is where the worst failure lives (the dep-* file carries # Checksum freshness is where the worst failure lives (the dep-* file carries
# per-source checksums and is rewritten in place). Only available on nightly; # per-source checksums and is rewritten in place). Only available on nightly;
# without it the test still covers the build/ and *.d families. # without it the test still covers the build/ and *.d families.
CHECKSUM_MODE="off"
if cargo +nightly -V >/dev/null 2>&1; then
export CARGO_UNSTABLE_CHECKSUM_FRESHNESS=true
CARGO_BIN=(cargo +nightly)
CHECKSUM_MODE="on"
else
CARGO_BIN=(cargo)
fi
echo "=== checksum-freshness mode: ${CHECKSUM_MODE} ==="
CONTENT_A='pub fn f() -> u32 { 1 }' CONTENT_A='pub fn f() -> u32 { 1 }'
CONTENT_B='pub fn f() -> u32 { 22222 } pub fn g() -> u32 { 7 }' CONTENT_B='pub fn f() -> u32 { 22222 } pub fn g() -> u32 { 7 }'
# Probe the BEHAVIOUR, not the channel and not the flag. Two weaker probes
# were tried against gitdan-ci's runner and each let the suite assert a
# property the toolchain did not have:
#
# `cargo +nightly -V` — answers "did a proxy called with
# +nightly exit 0". `-V`
# short-circuits before `-Z` is
# even parsed.
# `cargo +nightly -Z checksum-freshness — answers "is this flag still
# locate-project` accepted". 1.100.0-nightly
# (2026-08-25) accepts it and does
# not resolve freshness by content
# anyway.
#
# The scenario at the end of this file depends on one thing and it is neither
# of those: that changed content with an OLDER mtime rebuilds. Under mtime
# freshness the correct answer is Fresh, so under mtime freshness that
# scenario asserts a bug. So the probe simply performs that experiment, on its
# own crate and its own target dir, with no clone anywhere near it — which is
# also what makes it a control rather than a restatement of the scenario: the
# probe establishes that the toolchain rebuilds on content, the scenario
# establishes that a hardlink clone did not take that away.
# THREE OUTCOMES, NOT TWO. An experiment that cannot tell a negative result
# from a failed measurement is not settling the question, and the two are not
# interchangeable here: "this toolchain resolves freshness by mtime" is a
# statement about Cargo, while "a probe build failed" is a statement about this
# machine. Collapsing them — which an earlier cut of this did, by returning
# non-zero for both — makes a half-installed toolchain print a confident and
# wrong explanation and quietly drop a scenario. The scenario still has to be
# skipped in either case; what must not happen is the log claiming to know why.
#
# 0 content freshness measured ACTIVE — both builds ran, the backdated
# rebuild recompiled
# 3 measured INACTIVE — both builds ran, the backdated
# rebuild reported Fresh
# anything else NOT MEASURED — nothing was learned about the
# toolchain
#
# THE ANSWER CODES ARE 0 AND 3, AND THE GAP IS THE MECHANISM. Bash produces 1
# for an ordinary command failure, 2 for a usage error, 126/127 for a command
# it could not run, 128+n for a signal, and — this is the one that matters —
# 1 for an unbound-variable or other EXPANSION failure, which happens before
# the command runs and is therefore invisible to a `||` guard and to an ERR
# trap alike. It never produces 3. So "not an answer code" is decided by a
# property of the shell rather than by an enumeration of the ways a step can
# go wrong, and a step added later without a guard, or with a guard that
# cannot fire, lands on NOT MEASURED by construction.
#
# That is the whole reason INACTIVE is not 1. It was, and three review rounds
# on this function each found a narrower way for a shell-generated 1 to be read
# as a measurement — an unguarded command, then a typo'd variable name on a
# line that HAS its guard. Each was closed by narrowing the failure surface,
# which is a game with no last move. Moving the answer off the codes bash can
# generate ends it instead: there is no longer a mutation that turns an error
# into an answer, only mutations that turn an error into a different error.
#
# The guards below stay, and so does the trap, but their job is now reporting
# rather than correctness: they make a failed step land on 2 with its logs
# printed instead of on some incidental status, which is nicer to debug and
# lands in the same place either way.
#
# One piece of that reporting layer is load-bearing and not obvious. A command
# on the left of `||` — or in an `if` condition — runs with errexit suppressed,
# and that suppression propagates into a subshell and is NOT undone by a
# `set -e` inside it (measured on bash 5.3: an unguarded `false` there falls
# through to `exit 0`). Calling with errexit disarmed at the site is the only
# form that lets the subshell re-arm it; hence the `set +e` bracket. The ERR
# trap is then required on top, because a bare `set -e` abort exits with the
# FAILING COMMAND's status, and `false` gives 1.
#
# WHY NOT-MEASURED SKIPS RATHER THAN FAILS. The scenario it gates is the only thing in
# this suite that depends on freshness mode; everything else still runs and
# still catches real regressions. Failing instead would turn a statement about
# one machine's toolchain into a red gate reading "the hardlink scheme is
# broken" across the three repos consuming this action — the same category
# error the three-state split exists to prevent, one level up. What would
# change the answer is not-measured becoming the everyday CI outcome; it is not
# (gitdan-ci reports a measured INACTIVE, by measurement).
CHECKSUM_MODE="off"
CHECKSUM_REASON="no nightly on PATH accepting -Z checksum-freshness"
CARGO_BIN=(cargo)
checksum_freshness_probe() {
local d="$scratch/freshness-probe" t="$scratch/freshness-probe-target"
mkcrate "$d" || return 2 # 2 is simply "not 0 and not 3"; see the header
(
set -e
trap 'exit 2' ERR
cd "$d" || exit 2
printf '%s\n' "$CONTENT_A" > src/lib.rs || exit 2
CARGO_TARGET_DIR="$t" cargo +nightly build -q > "$scratch/freshness-probe-warm.log" 2>&1 || exit 2
printf '%s\n' "$CONTENT_B" > src/lib.rs || exit 2
touch -d '@1000000000' src/lib.rs || exit 2
CARGO_TARGET_DIR="$t" cargo +nightly build -v > "$scratch/freshness-probe.log" 2>&1 || exit 2
# 3, not 1: see the header. This is the only statement in the subshell that
# may report a measurement, and it is the only one that may exit 3.
if grep -qE '^\s+Fresh probe' "$scratch/freshness-probe.log"; then exit 3; fi
exit 0
)
}
if cargo +nightly -Z checksum-freshness locate-project > /dev/null 2>&1; then
export CARGO_UNSTABLE_CHECKSUM_FRESHNESS=true
# Errexit off across the call, so the subshell can arm its own — see the
# header. `probe_rc` is read before it is restored.
probe_rc=0
set +e
checksum_freshness_probe
probe_rc=$?
set -e
case "$probe_rc" in
0)
CARGO_BIN=(cargo +nightly)
CHECKSUM_MODE="on"
CHECKSUM_REASON=""
;;
3)
unset CARGO_UNSTABLE_CHECKSUM_FRESHNESS
CHECKSUM_REASON="this nightly accepts -Z checksum-freshness but resolves freshness by mtime"
;;
*)
unset CARGO_UNSTABLE_CHECKSUM_FRESHNESS
CHECKSUM_MODE="unmeasured"
CHECKSUM_REASON="the probe exited ${probe_rc}, which is not one of its answer codes, so this was NOT MEASURED — this toolchain may or may not resolve freshness by content"
# Loud, because the cost is silently lost coverage on a machine that
# might have had it. The suite continues: everything else it asserts is
# independent of freshness mode.
echo "::warning::hardlink-clone-selftest: could not measure whether this toolchain resolves freshness by content — the probe exited ${probe_rc}. This is a failure to measure, not a finding about Cargo."
tail -n 15 "$scratch/freshness-probe-warm.log" "$scratch/freshness-probe.log" 2>/dev/null | sed 's/^/ /' >&2 || true
;;
esac
fi
cd "$crate_dir"
echo "=== checksum-freshness mode: ${CHECKSUM_MODE}${CHECKSUM_REASON:+ — ${CHECKSUM_REASON}} ==="
build_base() { build_base() {
local dir="$1" local dir="$1"
printf '%s\n' "$CONTENT_A" > src/lib.rs printf '%s\n' "$CONTENT_A" > src/lib.rs
@@ -112,11 +244,26 @@ fi
ok "raw cp -al clone mutates the source ($(printf '%s\n' "$ctl_mutated" | wc -l) paths)" ok "raw cp -al clone mutates the source ($(printf '%s\n' "$ctl_mutated" | wc -l) paths)"
printf '%s\n' "$ctl_mutated" | sed 's/^/ /' printf '%s\n' "$ctl_mutated" | sed 's/^/ /'
# Reported, not asserted, and the distinction is the point. The control's job
# is to prove the hazard exists at all, which the non-empty set above already
# does; this line records WHICH families a given Cargo exhibits.
#
# `.fingerprint/*/dep-*` is the worst of them — it carries the per-source
# checksums, so mutating it through a shared inode turns a hardlink clone into
# silent stale-artifact reuse rather than a slow build. It was measured on
# cargo 1.9x nightly (see unshare_mutable_paths in cache-lib.sh) and is NOT
# reproduced on 1.100.0-nightly (2026-08-25), where the control mutates only
# the build/ and *.d families. Failing on its absence would mean this suite
# goes red whenever upstream stops doing something we never wanted it to do —
# and it would go red in the CONTROL, where a failure reads as "the hazard is
# gone" rather than "upstream changed". Nothing is lost by reporting it: the
# fix scenario below asserts the source is byte-identical after a full rebuild
# in the clone, which covers every family this Cargo has, named or not.
if [ "$CHECKSUM_MODE" = "on" ]; then if [ "$CHECKSUM_MODE" = "on" ]; then
if printf '%s' "$ctl_mutated" | grep -q '\.fingerprint/.*/dep-'; then if printf '%s' "$ctl_mutated" | grep -q '\.fingerprint/.*/dep-'; then
ok "control confirms the checksum-freshness dep-info file is among the mutated set" echo " note: this cargo DOES rewrite .fingerprint/*/dep-* in place under checksum freshness"
else else
fail "expected .fingerprint/*/dep-* in the control's mutated set under checksum freshness" echo " note: this cargo does NOT rewrite .fingerprint/*/dep-* in place; only the build/ and *.d families appear above"
fi fi
fi fi
@@ -133,7 +280,7 @@ hardlink_clone_into "$base_fix" "$clone_fix" "selftest" || fail "hardlink_clone_
# rebuild would prove nothing — the rebuild replaces those files anyway. # rebuild would prove nothing — the rebuild replaces those files anyway.
shared=0; unshared=0 shared=0; unshared=0
while IFS= read -r f; do while IFS= read -r f; do
rel="${f#$base_fix/}" rel="${f#"$base_fix"/}"
[ -e "$clone_fix/$rel" ] || continue [ -e "$clone_fix/$rel" ] || continue
if [ "$(stat -c '%i' "$f")" = "$(stat -c '%i' "$clone_fix/$rel")" ]; then if [ "$(stat -c '%i' "$f")" = "$(stat -c '%i' "$clone_fix/$rel")" ]; then
case "$rel" in case "$rel" in
@@ -161,18 +308,32 @@ fi
ok "no file in the source changed after a full rebuild in the clone" ok "no file in the source changed after a full rebuild in the clone"
echo echo
echo "=== the whole point: the source's next build is still correct ===" if [ "$CHECKSUM_MODE" = "on" ]; then
# The source's cache holds artifacts built from CONTENT_A. Advance the source echo "=== the whole point: the source's next build is still correct ==="
# to CONTENT_B (as a merge would) and rebuild in it. If the clone had # The source's cache holds artifacts built from CONTENT_A. Advance the
# corrupted its dep-info, Cargo would report Fresh and keep the stale rlib. # source to CONTENT_B (as a merge would) and rebuild in it. If the clone had
printf '%s\n' "$CONTENT_B" > src/lib.rs # corrupted its dep-info, Cargo would report Fresh and keep the stale rlib.
touch -d '@1000000000' src/lib.rs #
log="$scratch/rebuild.log" # CHECKSUM-FRESHNESS ONLY, and the backdated mtime is why. Under checksum
CARGO_TARGET_DIR="$base_fix" "${CARGO_BIN[@]}" build -v > "$log" 2>&1 || { cat "$log"; fail "rebuild in the source failed"; } # freshness the dep-info file's per-source checksums decide, so a 2001
if grep -qE '^\s+Fresh probe' "$log"; then # timestamp on changed content must still rebuild — the assertion below.
fail "source declared its own crate Fresh against sources it has never built — stale-artifact reuse" # Under Cargo's ordinary MTIME freshness the same timestamp means the source
# is older than the artifact, and reporting Fresh is the correct answer;
# asserting otherwise asserts a bug. This scenario was written against a
# machine with a nightly installed and, run without one, failed on that
# correct answer.
printf '%s\n' "$CONTENT_B" > src/lib.rs
touch -d '@1000000000' src/lib.rs
log="$scratch/rebuild.log"
CARGO_TARGET_DIR="$base_fix" "${CARGO_BIN[@]}" build -v > "$log" 2>&1 || { cat "$log"; fail "rebuild in the source failed"; }
if grep -qE '^\s+Fresh probe' "$log"; then
fail "source declared its own crate Fresh against sources it has never built — stale-artifact reuse"
fi
ok "source correctly rebuilt its crate after advancing to the clone's content"
else
echo "=== skipped: the source's-next-build scenario needs content-based freshness ==="
echo " reason: ${CHECKSUM_REASON}"
fi fi
ok "source correctly rebuilt its crate after advancing to the clone's content"
echo echo
echo "hardlink-clone-selftest: ${pass_count} assertions passed" echo "hardlink-clone-selftest: ${pass_count} assertions passed"
+3 -3
View File
@@ -65,10 +65,10 @@ origin="$scratch/origin.git"; git init -q --bare "$origin"
work="$scratch/work"; git init -q "$work" work="$scratch/work"; git init -q "$work"
( (
cd "$work" cd "$work"
git -c user.email=t@t -c user.name=t commit -q --allow-empty -m init git -c user.email=t@t -c user.name=t -c commit.gpgsign=false commit -q --allow-empty -m init
git branch -M main git branch -M main
git checkout -q -b dev; git -c user.email=t@t -c user.name=t commit -q --allow-empty -m dev git checkout -q -b dev; git -c user.email=t@t -c user.name=t -c commit.gpgsign=false commit -q --allow-empty -m dev
git checkout -q -b feat/live; git -c user.email=t@t -c user.name=t commit -q --allow-empty -m live git checkout -q -b feat/live; git -c user.email=t@t -c user.name=t -c commit.gpgsign=false commit -q --allow-empty -m live
git remote add origin "$origin" git remote add origin "$origin"
git push -q origin main dev feat/live git push -q origin main dev feat/live
) )
+15 -1
View File
@@ -75,6 +75,15 @@
# Asserts BOTH jobs correctly recompile the dependency and succeed. # Asserts BOTH jobs correctly recompile the dependency and succeed.
set -euo pipefail set -euo pipefail
# Every assertion below reads cargo's own words out of a build log
# (`Compiling libdep`, `Fresh probe`). A CI image that forces colour splices an
# ANSI reset between the status word and the crate name, at which point every
# one of those greps silently stops matching and the suite reports the
# opposite of what happened — observed on gitdan-ci's runner image, where
# scenario 2 failed while the log it printed plainly showed `Compiling libdep`.
# Pin the format the assertions are written against.
export CARGO_TERM_COLOR=never
script_dir=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd) script_dir=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)
restore_mtimes="$script_dir/restore-mtimes.sh" restore_mtimes="$script_dir/restore-mtimes.sh"
@@ -122,7 +131,7 @@ assert_log_lacks() {
# every run — simulate that before each restore-mtimes.sh pass, exactly as # every run — simulate that before each restore-mtimes.sh pass, exactly as
# CI would see it, so this test exercises the script the same way CI does. # CI would see it, so this test exercises the script the same way CI does.
stamp_checkout_now() { stamp_checkout_now() {
find . -path ./.git -prune -o -type f -print | xargs touch find . -path ./.git -prune -o -type f -print0 | xargs -0 touch
} }
echo "=== building scratch workspace ===" echo "=== building scratch workspace ==="
@@ -131,6 +140,11 @@ cd "$repo"
git init -q git init -q
git config user.email test@example.com git config user.email test@example.com
git config user.name "restore-mtimes-selftest" git config user.name "restore-mtimes-selftest"
# Local to this mktemp'd throwaway repo. Without it the eight commits below
# inherit the developer's GLOBAL commit.gpgsign, which makes whether this gate
# passes depend on their gpg agent — observed as a red run caused by a full
# disk breaking gpg, in a suite that has nothing to say about either.
git config commit.gpgsign false
cat > Cargo.toml <<'EOF' cat > Cargo.toml <<'EOF'
[workspace] [workspace]
+90
View File
@@ -74,6 +74,17 @@
# onto one set of fingerprints. That is the silent stale-reuse bug the # onto one set of fingerprints. That is the silent stale-reuse bug the
# whole scheme exists to prevent, so the failure has to abort the clone # whole scheme exists to prevent, so the failure has to abort the clone
# rather than be swallowed. # rather than be swallowed.
# 11. THE CONSUMER'S MARKER ACTUALLY STOPS THE PUBLISHER — reader_lock_acquire
# is exercised (not faked, unlike publish-snapshot-selftest.sh's scenario
# 6) by a real hardlink_clone_into racing a real, concurrent
# publish-snapshot.sh republish of the exact snapshot being cloned. Pins
# that the publisher OBSERVABLY WAITS on this consumer's marker — its own
# log reports entering the drain wait — rather than only that the run
# succeeds, which stayed green with the marker call deleted (issue #10).
# Kept as its own scenario, not folded into 8a, because 8a already pins
# exactly one property (the identity check) for exactly one mutant, and
# the suite's one-scenario-one-mutant diagonal across 8a to 8d and 10 is
# deliberate.
set -euo pipefail set -euo pipefail
script_dir=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd) script_dir=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)
. "$script_dir/cache-lib.sh" . "$script_dir/cache-lib.sh"
@@ -635,5 +646,84 @@ leftovers=$(find "$root" -maxdepth 1 \( -name '.stage-*' -o -name '.reading-*' \
ok "no staging or reader-marker scratch left behind" ok "no staging or reader-marker scratch left behind"
rm -f "$scratch/bin/cp" rm -f "$scratch/bin/cp"
echo
echo "=== 11: the consumer's marker actually stops the publisher ==="
# hardlink_clone_into's reader_lock_acquire call is exercised here, not
# faked. publish-snapshot-selftest.sh's scenario 6 stands a hand-written
# marker file in for "a consumer whose clone outlasts the grace period" — a
# deliberate simplification that does not need this consumer's clone code at
# all, so it cannot tell reader_lock_acquire apart from no marker existing.
# This scenario forces the two real scripts to race on the same snapshot: a
# consumer hardlink-cloning it, and a publisher republishing it out from
# under that clone, exactly as 8a does to force the identity check — but what
# is pinned here is not the consumer's response to the rotation (8a's
# property), it is that the PUBLISHER, reading the SAME marker this
# consumer's clone wrote, is observed to have entered its drain wait. Deleting
# reader_lock_acquire (issue #10) leaves the marker never written: the
# publisher's live_reader_count sees zero readers on its first check and
# proceeds straight to reclaiming the rotated generation, silently — the run
# still succeeds, and nothing about its own outcome says so, only the absence
# of a line in the publisher's log.
INTLK=$(cache_key feat/interlock-consumer)
INTLK_BASE=$(cache_key release/5)
TAG_INTLK=jobIntlk
rm -rf "$root/snapshot-$INTLK_BASE" "$root/target-$INTLK_BASE"
make_tree "$root/snapshot-$INTLK_BASE" intlk-gen1
make_tree "$root/target-$INTLK_BASE" intlk-gen2
intlk_gen1_inode=$(stat -c '%i' "$root/snapshot-$INTLK_BASE")
cat > "$scratch/bin/cp" <<EOF
#!/usr/bin/env bash
# Fires once, only on the consumer's own top-level hardlink clone — identified
# by its destination, the seed's private staging path. Every other cp in the
# process tree (the unshare copies, and the publisher's own staging clone)
# falls through to the real one.
if [ "\${@: -1}" = "$root/.stage-$TAG_INTLK" ] && [ ! -e "$scratch/firedIntlk" ]; then
: > "$scratch/firedIntlk"
rc=0; "$real_cp" "\$@" || rc=\$?
# Concurrently: a real, second publish of the exact snapshot this consumer
# is cloning — the republish that, without the marker this consumer's clone
# holds, would reclaim the generation out from under it.
( CACHE_READ_GRACE_SECONDS=10 bash "$script_dir/publish-snapshot.sh" "$INTLK_BASE" "$root" pubIntlk \
> "$scratch/logPubIntlk" 2>&1
echo \$? > "$scratch/rcPubIntlk" ) &
# Hand control back only once the swap is on disk, so the identity read
# immediately after this cp is guaranteed to resolve to the new generation
# — the same technique 8a uses to force the interleaving rather than hope
# for it. reader_lock_release does not run until AFTER this script exits,
# so the marker stays live for the publisher's whole swap-and-scan.
deadline=\$(( \$(date +%s) + 60 ))
while [ "\$(stat -c '%i' "$root/snapshot-$INTLK_BASE" 2>/dev/null)" = "$intlk_gen1_inode" ]; do
[ "\$(date +%s)" -lt "\$deadline" ] || { echo "stub cp: the publisher never swapped the snapshot" >&2; exit 92; }
sleep 0.05
done
exit \$rc
fi
exec "$real_cp" "\$@"
EOF
chmod +x "$scratch/bin/cp"
rcIntlk=0
seed_with_stub "$INTLK" "$INTLK_BASE" "$root" "$TAG_INTLK" > "$scratch/logIntlk" 2>&1 || rcIntlk=$?
[ -e "$scratch/firedIntlk" ] \
|| fail "the stubbed cp never fired: the consumer never raced the publisher, so this scenario proves nothing"
ok "the consumer's clone raced a real, concurrent republish of its own source"
wait_for_file "$scratch/rcPubIntlk" "the publisher never finished"
[ "$(cat "$scratch/rcPubIntlk")" = "0" ] || { tail -40 "$scratch/logPubIntlk"; fail "publish-snapshot.sh exited non-zero"; }
[ "$rcIntlk" = "0" ] || { tail -40 "$scratch/logIntlk"; fail "the seed exited non-zero"; }
ok "both sides of the race completed"
# The property under test: not that the run succeeded, but that the publisher
# itself reports having found a live reader and waited on it. This is silent
# in the consumer's own log and in the run's exit status alike — only the
# publisher's log carries it.
grep -q "readers: waiting for 1 in-flight clone(s) of snapshot-${INTLK_BASE}" "$scratch/logPubIntlk" \
|| { tail -40 "$scratch/logPubIntlk"; fail "the publisher never reported waiting on the consumer's reader marker — the interlock did not observably engage"; }
ok "the publisher observably waited on the consumer's own reader marker before reclaiming the rotated snapshot generation"
leftovers=$(find "$root" -maxdepth 1 \( -name '.stage-*' -o -name '.reading-*' -o -name '.publish-*' \) -print)
[ -z "$leftovers" ] || fail "scratch left behind: ${leftovers}"
ok "no staging, reader-marker or deferred-generation scratch left behind"
rm -f "$scratch/bin/cp"
echo echo
echo "seed-target-dir-selftest: ${pass_count} assertions passed" echo "seed-target-dir-selftest: ${pass_count} assertions passed"
+1 -1
View File
@@ -13,7 +13,7 @@ script_dir=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)
FAST=0 FAST=0
[ "${1:-}" = "--fast" ] && FAST=1 [ "${1:-}" = "--fast" ] && FAST=1
FIXTURE_TESTS=(seed-target-dir-selftest.sh publish-snapshot-selftest.sh prune-cache-selftest.sh) FIXTURE_TESTS=(cache-root-selftest.sh seed-target-dir-selftest.sh publish-snapshot-selftest.sh prune-cache-selftest.sh)
CARGO_TESTS=(hardlink-clone-selftest.sh restore-mtimes-selftest.sh) CARGO_TESTS=(hardlink-clone-selftest.sh restore-mtimes-selftest.sh)
TESTS=("${FIXTURE_TESTS[@]}") TESTS=("${FIXTURE_TESTS[@]}")