The shared action's justification over zemyna's and emowheel's schemes was that hardlink-cloning from a published snapshot closes gitdan #911 "by construction, not by the single job slot". Review disproved that. This makes the claim true, and corrects the README where it could only be bounded. Finding 1 (verdict-level) — silent partial clone ------------------------------------------------ `hardlink_clone_into` ran `cp -al` with no exit-status check, and both call sites invoked it as a condition, which suppresses `set -e` for the whole call. A publisher's `rm -rf` of the generation it rotated away therefore unlinked entries beneath an in-flight consumer walk, and the truncated tree was renamed into place and reported as success. Both layers are fixed: * The consumer verifies its own clone. Every attempt checks `cp -al`'s status explicitly, the source directory's inode before and after (a wholesale replacement mid-walk splices two generations), and the entry count — the only signal for a subtree unlinked before its parent was listed, since `cp -al` reports no error for one it never saw. Any failure discards the staging tree and retries; exhausting the attempts returns a distinct status 2 and fails the job rather than seeding a partial cache. `unshare_subtree` / `_unshare_files` now propagate failure too — a swallowed unshare leaves the clone aliasing its source, the exact corruption that step exists to prevent. * The publisher does not unlink under a reader. A consumer publishes a `.reading-<snapshot>-<tag>` marker before it resolves the snapshot path; the publisher scans for markers after its first rename. A consumer holding the old generation therefore published its marker before that scan and cannot be missed; one arriving after the scan necessarily resolves to the new generation. The publisher waits for readers to drain and, on timeout, DEFERS reclamation rather than forcing it — the old generation is left as `.publish-old-<key>-<tag>` and swept by a later publish. So correctness is closed by construction; disk reclamation is bounded, not immediate. The residual is capped at one deferred generation per publisher ref, and the README now says exactly that instead of the disproved claim. Finding 2 — restore-mtimes.sh ran with no errexit ------------------------------------------------- `set -euo pipefail` was glued to the end of a comment (`# soundness.set -euo pipefail`), so it was entirely commented out: a partial failure of the `git log | awk` pipeline would have produced wrong mtimes across the whole restore instead of failing loudly. Moved to its own line. Audited every other script for the same defect — this was the only instance. Independent confirmation: shellcheck's two SC2164 warnings on this file's `cd "$repo_root"` disappear now that errexit is actually in effect. Finding 3 — lock-acquire window ------------------------------- A just-seeded directory was unlocked until a later action step, so a concurrent job's prune pass could evict it. `seed-target-dir.sh` now takes an optional lock-id and writes the lock marker on every path out of the script, including into the staging tree before its rename, so the directory carries a lock the instant it appears under its final name. The action's acquire step stays (it is idempotent and stamps the LRU marker). Also hardened `prune-cache.sh` to treat a directory with live reader markers as locked. Today no reachable configuration prunes a snapshot — only protected refs publish them and protected refs are excluded from every pass — so this is redundant by policy; it is here so that stops being the reason it is safe. Verification ------------ New selftest scenario 8 races a real seed against a real publish rotation, gating the rotation on the seed's *observed* clone progress so the window is hit deterministically rather than on a fast machine's coin flip. Red-proven against the unguarded scripts, three consecutive runs: ASSERTION FAILED: the seeded tree is truncated: 15443 entries against the snapshot's 493 (was 48805 before the rotation) (15443 / 16986 / 16498) Green after the fix, six consecutive runs, catching the clone mid-walk at ~10.5k of 48805 entries each time. Scenario 9 covers deferred reclamation and its later sweep; scenario 10 covers an unreadable source failing loudly. `bash scripts/selftest.sh`: 5 suites, exit 0, 75 assertions (was 63). shellcheck over `scripts/`: no new findings, two SC2164 warnings resolved. Docs: README's republish-safety paragraph replaced with what the code now guarantees, including the bounded disk residual stated explicitly; new `read-grace-seconds` / `reader-stale-seconds` inputs documented in the `cargo-cache-publish` table; the selftest table names the new race. Refs: daniel/gitdan#11, zemyna#911 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Sqh2vscfzisk83VuPVQX9L
15 KiB
gitdan-actions
Shared Gitea Actions composite actions for the gitdan forge.
Currently one thing, done properly: cargo-cache — a persistent,
per-branch Cargo build cache for self-hosted Gitea runners, where a pull
request's cache is a near-free hardlink clone of an immutable snapshot its
base branch published.
Final home: this repository will live at
daniel/gitdan-actions. Pin that path inuses:once the transfer completes.
Why this exists
Two of this forge's Rust projects independently built the same idea and each got one half right.
| seeding mechanism | seed source | |
|---|---|---|
| project A | cp -al hardlink clone — near-free, cost scales with inode count, not bytes |
the base branch's live target dir — races a build that is still writing |
| project B | cp -a full copy — sound, but ~35 GB duplicated per branch |
a published immutable snapshot — nothing ever writes it while it is read |
This action is the diagonal: hardlink-clone from a published snapshot. Cheap like A, sound like B. It also closes a latent race in A by construction (the seed is staged and swapped in with one atomic rename) rather than relying on the runner having a single execution slot.
One thing neither project had, and the reason the clone is not a plain
cp -al: a build inside a hardlink clone does mutate the directory it was
cloned from. Cargo writes real artifacts by replacing them, but writes its
metadata — and build scripts write their OUT_DIR — with a plain truncating
write, straight through the shared inode. Under
CARGO_UNSTABLE_CHECKSUM_FRESHNESS the file that gets corrupted is
.fingerprint/<unit>/dep-<target>, which holds the per-source checksums that
decide freshness, and the failure is silent stale-artifact reuse rather than a
slow build. scripts/hardlink-clone-selftest.sh reproduces it as an explicit
control and asserts the fix. The fix is to hardlink the artifacts (the GB) and
real-copy the metadata (the MB) — about 3.7% of a Bevy-sized target directory,
against 100% for a full copy.
Quick start
name: CI
on:
push:
branches: [main, dev]
pull_request:
branches: [main, dev]
jobs:
ci:
runs-on: ubuntu-latest
# REQUIRED, and it cannot come from the action: `container.volumes` is a
# job-level property, so the persistent cache volume must be declared
# here. Use a volume name unique to this repository.
container:
volumes:
- myrepo-ci-target:/cache
steps:
- uses: actions/checkout@v4
with:
# REQUIRED. The mtime restore walks every commit that ever touched a
# tracked file; a depth-1 checkout makes every file resolve to the
# tip commit and the cache stops working. The action fails loudly
# rather than silently degrading if this is missing.
fetch-depth: 0
- uses: https://gitdan.com/daniel/gitdan-actions/cargo-cache@v1
with:
protected-branches: 'dev main'
# ... toolchain, system deps, and the build itself. CARGO_TARGET_DIR is
# already exported to the job environment by the step above.
- run: cargo clippy --workspace --all-targets -- -D warnings
- run: cargo test --workspace
# After the build succeeds: record the watermark, and publish a snapshot
# if this run is a push to a protected branch.
- uses: https://gitdan.com/daniel/gitdan-actions/cargo-cache-publish@v1
# Release this job's cache lock even when the build failed, so the
# eviction pass does not have to wait out the staleness grace period.
- if: always()
uses: https://gitdan.com/daniel/gitdan-actions/cargo-cache-publish@v1
with:
mode: release-lock
Recommended alongside it, in the workflow's env: block:
env:
CARGO_INCREMENTAL: 0 # per-run bloat on a persistent volume
CARGO_PROFILE_DEV_DEBUG: line-tables-only
CARGO_PROFILE_TEST_DEBUG: line-tables-only
# Nightly only. Content-addressed freshness instead of mtime-based — a
# strictly stronger guarantee, complementary to the mtime restore (which
# still covers directory-form `rerun-if-changed` build-script watches).
CARGO_UNSTABLE_CHECKSUM_FRESHNESS: "true"
How it works
/cache/
target-<key> one per ref. Where a build actually runs.
snapshot-<key> one per publisher ref. Immutable between publishes;
the only thing a consumer ever clones from.
<key> is the ref sanitised to a safe path component, capped at 48
characters, plus an 8-hex SHA-1 prefix of the raw ref. The hash is not
decoration: feat/foo and feat-foo sanitise identically and would otherwise
share one directory.
A pull request run resolves its own key from github.head_ref (not
ref_name, which on a pull_request event is a synthetic merge ref that
changes on every push) and its base key from github.base_ref. If it has no
directory yet, it hardlink-clones snapshot-<base> into a staging path, strips
Cargo's lock files, real-copies everything Cargo writes in place, and renames
the staging path into target-<own>.
A push to a protected branch has no base to layer over. It builds in its
own directory and, if the build goes green, republishes it as
snapshot-<own>: stage a clone, rename the old snapshot aside, rename the new
one in, then reclaim the old one once nothing is still reading it. Consumers
only ever observe a complete snapshot or none at all.
Concurrency, on the destination. Two jobs sharing one cache key each stage
under their own tag and race on one atomic rename; the loser discards its
staging copy. There is no window in which a partially-populated directory is
visible under the final name. Two jobs then building in the same directory is
Cargo's own .cargo-lock territory, which is what that lock is for.
Concurrency, on the source. The atomic rename is necessary and not
sufficient, because renaming a truncated tree publishes a truncated tree
atomically. A clone reads its source over many seconds, and a publisher
rotating that source unlinks the generation being read — at which point
cp -al can silently omit a subtree it never saw, and report success. Two
mechanisms, both required:
- The publisher does not unlink under a reader. A consumer publishes a
.reading-<snapshot>-<tag>marker before it resolves the snapshot path; the publisher scans for markers after its first rename. A consumer holding the old generation therefore published its marker before that scan and cannot be missed, and one that arrives after the scan necessarily resolves to the new generation. The publisher waits for readers to drain (read-grace-seconds, default 300) and, if they do not, defers the reclamation rather than forcing it — the old generation stays on disk and is swept by a later publish. - The consumer verifies its own clone. Every attempt checks
cp -al's exit status, the source directory's inode before and after (a wholesale replacement mid-walk would otherwise splice two generations), and the entry count (the only signal for a subtree unlinked before its parent was listed — there is no error to read). A tree that fails any of the three is deleted and the clone retried; one that fails the last attempt fails the job. A partial tree never reaches the final name.
What this does and does not guarantee. Correctness is closed by construction: no combination of publish and seed timing produces a target directory holding part of one generation, and a source that cannot be read consistently fails the job loudly instead of seeding a truncated cache. Disk reclamation is bounded, not immediate: a consumer slower than the grace period leaves one extra snapshot generation of directory entries on the volume until the next publish of that snapshot sweeps it. That residual is capped at one deferred generation per publisher ref, and its real cost is close to the inode count rather than the byte count, since the artifacts are hardlinked to whatever cloned them.
Eviction runs three passes: caches for branches that no longer exist on
origin are removed unconditionally; then, only if free space is under the
threshold, live caches are evicted oldest-first; then, as a last resort, this
run's own cache. Protected refs and any cache held open by a running job are
never candidates. Within the pressure pass, target-* directories are evicted
before snapshot-* ones — the reverse of the obvious order, because a
snapshot is hardlinked to everything cloned from it, so removing one frees
almost no real bytes while costing every future PR its warm start.
File mtimes. actions/checkout stamps every file with "now", which makes
every crate look changed to Cargo's mtime-based freshness check — a persistent
target directory buys nothing without fixing that. Each tracked file is
restored to the timestamp of the most recent commit that touched it, plus a
watermark override: for any file that changed since this cache's own last
successful build, "now" is stamped instead. That override is what makes a
merge safe, since a merge can introduce a commit authored before this cache's
last build, where the historically-correct mtime is exactly the wrong answer.
Inputs
cargo-cache
| input | default | meaning |
|---|---|---|
cache-root |
/cache |
mount point of the persistent volume inside the job container |
protected-branches |
dev main |
refs that publish snapshots and are never evicted |
min-free-percent |
10 |
prune when free space drops below this |
restore-mtimes |
true |
restore tracked-file mtimes from git history |
prune |
true |
run the eviction pass |
liveness-prune |
true |
within eviction, remove caches for branches gone from origin |
own-ref |
(auto) | override; defaults to github.head_ref, else github.ref_name |
base-ref |
(auto) | override; defaults to github.base_ref (empty on push) |
seed-fallback-dir |
(empty) | absolute path to seed from when no snapshot exists — for migrating off an existing flat cache |
watermark-file |
.ci-watermark-<job>-sha |
must differ per job when two jobs share one cache key |
lock-id |
<job>-<run_id> |
identifies this job's cache lock |
stale-lock-seconds |
7200 |
age past which another job's lock is treated as abandoned |
Outputs: target-dir, cache-key, seeded-from (own | base-snapshot |
own-snapshot | fallback-dir | concurrent-peer | cold).
Exports to the job environment: CARGO_TARGET_DIR, CARGO_CACHE_ROOT,
CARGO_CACHE_KEY, CARGO_CACHE_LOCK_ID, CARGO_CACHE_SCRIPTS,
CI_WATERMARK_FILE.
cargo-cache-publish
| input | default | meaning |
|---|---|---|
cache-root |
/cache |
must match the consume action |
protected-branches |
dev main |
refs that publish snapshots |
mode |
publish |
publish, or release-lock for the if: always() step |
own-ref |
(auto) | override; defaults to github.head_ref, else github.ref_name |
publish-on-events |
push |
events on which a protected ref actually publishes |
read-grace-seconds |
300 |
how long the swap waits for in-flight clones of the generation it replaces before reclaiming it; on timeout the reclamation is deferred, never forced |
reader-stale-seconds |
7200 |
age past which a consumer's read marker is treated as abandoned by a killed job |
record-watermark |
true |
record HEAD as this cache's watermark (PR runs too) |
publish-on-events defaults to push on purpose: a pull_request run from
dev into main has own-ref dev and would otherwise publish a snapshot
of a merge-preview build, which is not what dev is.
Multiple jobs in one workflow
Jobs sharing a cache key (a ci job and a wasm job on the same branch, say)
each need their own watermark file. A shared one breaks the moment two
jobs run in sequence within one trigger: job A advances the watermark to HEAD,
and job B then reads that just-advanced value, computes an empty diff, and
loses the merge protection entirely. The default (.ci-watermark-<job>-sha)
already gives each job its own; only override watermark-file if you also
override lock-id, and then keep both distinct per job.
Constraints of this runner
- The repository must be public. act_runner fetches actions by anonymous git clone and has no credentialed-fetch option, so a private action repository simply fails to resolve. Nothing secret goes in here.
uses:needs the absolute URL. A bareowner/reporesolves against github.com, because Gitea'sDEFAULT_ACTIONS_URLis unset — and it has to stay unset, oractions/checkout,dtolnay/rust-toolchainandtaiki-e/install-actionstop resolving.- The runner pre-fetches every referenced action before running any step, so a bad action reference fails the job at step 0 rather than where it is used.
container.volumesis job-level and cannot be set from inside a composite action. The consuming workflow declares it; see the quick start.- The cache volume is ext4 — no reflink support, which is precisely why hardlinks are the mechanism that makes cloning cheap.
Versioning
Pin @v1. It is a moving major tag: fixes and backward-compatible inputs move
it forward, and anything that would break an existing consumer gets v2
instead. Pin a commit SHA if you want a frozen version.
Development
bash scripts/selftest.sh # everything (~1 min; needs cargo)
bash scripts/selftest.sh --fast # fixture-only suites, no compiler
| suite | covers |
|---|---|
hardlink-clone-selftest.sh |
that a build in a clone cannot mutate its source — with a control proving a raw cp -al does. Needs a real compiler. |
seed-target-dir-selftest.sh |
seed-source preference, lock-file stripping, two jobs racing on one cache key, and a seed racing a publisher's rotation of the source it is reading — the race that actually truncates a tree |
publish-snapshot-selftest.sh |
the atomic swap, and that a live consumer survives a republish |
prune-cache-selftest.sh |
liveness, protection, locking, eviction order, self-clear — against a real scratch origin |
restore-mtimes-selftest.sh |
the merge hazard and the watermark that closes it, including the two-jobs-one-namespace case. Needs a real compiler. |
Every suite runs the actual script, not a reimplementation of its logic, and every fix scenario is paired with a control that reproduces the bug — a scenario that passes either way proves nothing. The concurrency scenarios race real processes rather than mocking the interleaving, and gate the interfering step on observed progress of the step it interferes with, so the window is hit deterministically instead of on a fast machine's coin flip.
The action YAML holds no logic beyond wiring; everything testable lives in
scripts/. A composite action needs shell: bash on every run: step, and
the actions reach their shared scripts through
${{ github.action_path }}/../scripts, which works because the runner clones
the whole repository when it fetches an action.