Files
gitdan-actions/README.md
T
claudeandClaude Opus 5 719475831b fix(cargo-cache): scope the safety claim to what the code actually prevents
Follow-up to f57e2a6, addressing the reviewer's sharpest question: is the race
"closed by construction", or merely detected and retried? The honest answer is
"both, on different paths", and the README said only the first half.

* README now states three claims separately instead of collapsing them:
  a publisher rotating a snapshot cannot tear a clone of it (by construction —
  the marker ordering prevents the unlink, and the consumer's verification is
  a redundant second check on that path); every OTHER way the source can
  change mid-clone is detected, not prevented (the eviction pass's reader
  check is check-then-delete, and a seed-fallback-dir has no interlock at all
  — there, verification plus a bounded retry and a loud failure is the whole
  guard); and disk reclamation is bounded rather than immediate. Overclaiming
  this property once was the finding; overclaiming it twice would be worse.

* publish-snapshot-selftest.sh now covers the PUBLISHER's half of the race,
  where a reader of the swap belongs. Scenario 3 only ever covered a consumer
  that had already FINISHED cloning — safe for free, since its own hardlinks
  keep the inodes alive. New scenario 6 covers a reader still in flight past
  the grace period (the generation is left on disk, the deferral is warned
  about, and an earlier consumer is still unaffected); scenario 7 covers the
  sweep, so "we defer instead of forcing" cannot quietly become a disk leak.
  Red-proven against 248af306's scripts:

    ASSERTION FAILED: the previous generation was unlinked while a reader
    still held it

  The consumer's half stays in seed-target-dir-selftest.sh scenario 8, which
  still red-proves at 16693 of 48805 entries against the same scripts.

* seed-target-dir-selftest.sh now asserts what happens when the retries are
  EXHAUSTED, not just what hardlink_clone_into returns: an unreadable source
  makes the seed script exit non-zero, name the reason, leave no target dir,
  and — the one that matters — not fall through to its cold-start branch. A
  corrupt-cache bug degrading into an invisible 4x-slower CI job is the
  failure mode worth pinning down. Skipped when running as root, where mode
  bits deny nothing.

* usage_kb: a directory we cannot read measured as the empty string, which was
  then spliced into usage_gb's awk program and made it a syntax error at the
  exact moment something was already going wrong. Now measures 0.

Verification: `bash scripts/selftest.sh` — 5 suites, exit 0, 82 assertions
(was 75 after f57e2a6, 63 before). shellcheck over scripts/: no new findings.

Measured the cost the reviewer asked about, on ext4, warm cache, 78,554
entries: `cp -al` 3126 ms against 44 ms for one `find | wc -l`. Two counts per
attempt is ~2.8% on top of the clone. Not measured on the CI runner's volume.

Refs: daniel/gitdan#11, zemyna#911

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Sqh2vscfzisk83VuPVQX9L
2026-08-23 16:49:44 -05:00

16 KiB

gitdan-actions

Shared Gitea Actions composite actions for the gitdan forge.

Currently one thing, done properly: cargo-cache — a persistent, per-branch Cargo build cache for self-hosted Gitea runners, where a pull request's cache is a near-free hardlink clone of an immutable snapshot its base branch published.

Final home: this repository will live at daniel/gitdan-actions. Pin that path in uses: once the transfer completes.


Why this exists

Two of this forge's Rust projects independently built the same idea and each got one half right.

seeding mechanism seed source
project A cp -al hardlink clone — near-free, cost scales with inode count, not bytes the base branch's live target dir — races a build that is still writing
project B cp -a full copy — sound, but ~35 GB duplicated per branch a published immutable snapshot — nothing ever writes it while it is read

This action is the diagonal: hardlink-clone from a published snapshot. Cheap like A, sound like B. It also closes a latent race in A by construction (the seed is staged and swapped in with one atomic rename) rather than relying on the runner having a single execution slot.

One thing neither project had, and the reason the clone is not a plain cp -al: a build inside a hardlink clone does mutate the directory it was cloned from. Cargo writes real artifacts by replacing them, but writes its metadata — and build scripts write their OUT_DIR — with a plain truncating write, straight through the shared inode. Under CARGO_UNSTABLE_CHECKSUM_FRESHNESS the file that gets corrupted is .fingerprint/<unit>/dep-<target>, which holds the per-source checksums that decide freshness, and the failure is silent stale-artifact reuse rather than a slow build. scripts/hardlink-clone-selftest.sh reproduces it as an explicit control and asserts the fix. The fix is to hardlink the artifacts (the GB) and real-copy the metadata (the MB) — about 3.7% of a Bevy-sized target directory, against 100% for a full copy.


Quick start

name: CI
on:
  push:
    branches: [main, dev]
  pull_request:
    branches: [main, dev]

jobs:
  ci:
    runs-on: ubuntu-latest
    # REQUIRED, and it cannot come from the action: `container.volumes` is a
    # job-level property, so the persistent cache volume must be declared
    # here. Use a volume name unique to this repository.
    container:
      volumes:
        - myrepo-ci-target:/cache
    steps:
      - uses: actions/checkout@v4
        with:
          # REQUIRED. The mtime restore walks every commit that ever touched a
          # tracked file; a depth-1 checkout makes every file resolve to the
          # tip commit and the cache stops working. The action fails loudly
          # rather than silently degrading if this is missing.
          fetch-depth: 0

      - uses: https://gitdan.com/daniel/gitdan-actions/cargo-cache@v1
        with:
          protected-branches: 'dev main'

      # ... toolchain, system deps, and the build itself. CARGO_TARGET_DIR is
      # already exported to the job environment by the step above.
      - run: cargo clippy --workspace --all-targets -- -D warnings
      - run: cargo test --workspace

      # After the build succeeds: record the watermark, and publish a snapshot
      # if this run is a push to a protected branch.
      - uses: https://gitdan.com/daniel/gitdan-actions/cargo-cache-publish@v1

      # Release this job's cache lock even when the build failed, so the
      # eviction pass does not have to wait out the staleness grace period.
      - if: always()
        uses: https://gitdan.com/daniel/gitdan-actions/cargo-cache-publish@v1
        with:
          mode: release-lock

Recommended alongside it, in the workflow's env: block:

env:
  CARGO_INCREMENTAL: 0                  # per-run bloat on a persistent volume
  CARGO_PROFILE_DEV_DEBUG: line-tables-only
  CARGO_PROFILE_TEST_DEBUG: line-tables-only
  # Nightly only. Content-addressed freshness instead of mtime-based — a
  # strictly stronger guarantee, complementary to the mtime restore (which
  # still covers directory-form `rerun-if-changed` build-script watches).
  CARGO_UNSTABLE_CHECKSUM_FRESHNESS: "true"

How it works

/cache/
  target-<key>       one per ref. Where a build actually runs.
  snapshot-<key>     one per publisher ref. Immutable between publishes;
                     the only thing a consumer ever clones from.

<key> is the ref sanitised to a safe path component, capped at 48 characters, plus an 8-hex SHA-1 prefix of the raw ref. The hash is not decoration: feat/foo and feat-foo sanitise identically and would otherwise share one directory.

A pull request run resolves its own key from github.head_ref (not ref_name, which on a pull_request event is a synthetic merge ref that changes on every push) and its base key from github.base_ref. If it has no directory yet, it hardlink-clones snapshot-<base> into a staging path, strips Cargo's lock files, real-copies everything Cargo writes in place, and renames the staging path into target-<own>.

A push to a protected branch has no base to layer over. It builds in its own directory and, if the build goes green, republishes it as snapshot-<own>: stage a clone, rename the old snapshot aside, rename the new one in, then reclaim the old one once nothing is still reading it. Consumers only ever observe a complete snapshot or none at all.

Concurrency, on the destination. Two jobs sharing one cache key each stage under their own tag and race on one atomic rename; the loser discards its staging copy. There is no window in which a partially-populated directory is visible under the final name. Two jobs then building in the same directory is Cargo's own .cargo-lock territory, which is what that lock is for.

Concurrency, on the source. The atomic rename is necessary and not sufficient, because renaming a truncated tree publishes a truncated tree atomically. A clone reads its source over many seconds, and a publisher rotating that source unlinks the generation being read — at which point cp -al can silently omit a subtree it never saw, and report success. Two mechanisms, both required:

  • The publisher does not unlink under a reader. A consumer publishes a .reading-<snapshot>-<tag> marker before it resolves the snapshot path; the publisher scans for markers after its first rename. A consumer holding the old generation therefore published its marker before that scan and cannot be missed, and one that arrives after the scan necessarily resolves to the new generation. The publisher waits for readers to drain (read-grace-seconds, default 300) and, if they do not, defers the reclamation rather than forcing it — the old generation stays on disk and is swept by a later publish.
  • The consumer verifies its own clone. Every attempt checks cp -al's exit status, the source directory's inode before and after (a wholesale replacement mid-walk would otherwise splice two generations), and the entry count (the only signal for a subtree unlinked before its parent was listed — there is no error to read). A tree that fails any of the three is deleted and the clone retried; one that fails the last attempt fails the job. A partial tree never reaches the final name.

What this does and does not guarantee. Three separate claims, deliberately not collapsed into one:

  • A publisher rotating a snapshot cannot tear a clone of it — by construction. This is the case zemyna #911 is about, and the marker ordering above is what closes it: the publisher's scan cannot miss a consumer that resolved the old generation, and on timeout it defers the unlink rather than forcing it. On this path the consumer's own verification is a redundant second check, not the thing holding the guarantee up.
  • Every other way the source can change mid-clone is detected, not prevented. The eviction pass's reader check is check-then-delete, so a consumer publishing its marker inside that gap is narrowed but not excluded — unreachable today only because snapshots belong to protected refs and protected refs are never eviction candidates, which is policy rather than structure. A seed-fallback-dir pointing at a directory something else writes has no interlock at all. There, the per-attempt verification is what stands between a torn read and a corrupt cache: the clone is retried (CACHE_CLONE_ATTEMPTS, default 4) and then fails the job loudly — never seeded partially, and never degraded to a silent cold build.
  • Disk reclamation is bounded, not immediate. A consumer slower than the grace period leaves one extra snapshot generation of directory entries on the volume until a later publish sweeps it; a consumer whose job was killed outright holds it until its marker passes reader-stale-seconds. The residual is capped at one deferred generation per publisher ref, and its real cost is close to inode count rather than byte count, since the artifacts are hardlinked to whatever cloned them.

Eviction runs three passes: caches for branches that no longer exist on origin are removed unconditionally; then, only if free space is under the threshold, live caches are evicted oldest-first; then, as a last resort, this run's own cache. Protected refs and any cache held open by a running job are never candidates. Within the pressure pass, target-* directories are evicted before snapshot-* ones — the reverse of the obvious order, because a snapshot is hardlinked to everything cloned from it, so removing one frees almost no real bytes while costing every future PR its warm start.

File mtimes. actions/checkout stamps every file with "now", which makes every crate look changed to Cargo's mtime-based freshness check — a persistent target directory buys nothing without fixing that. Each tracked file is restored to the timestamp of the most recent commit that touched it, plus a watermark override: for any file that changed since this cache's own last successful build, "now" is stamped instead. That override is what makes a merge safe, since a merge can introduce a commit authored before this cache's last build, where the historically-correct mtime is exactly the wrong answer.


Inputs

cargo-cache

input default meaning
cache-root /cache mount point of the persistent volume inside the job container
protected-branches dev main refs that publish snapshots and are never evicted
min-free-percent 10 prune when free space drops below this
restore-mtimes true restore tracked-file mtimes from git history
prune true run the eviction pass
liveness-prune true within eviction, remove caches for branches gone from origin
own-ref (auto) override; defaults to github.head_ref, else github.ref_name
base-ref (auto) override; defaults to github.base_ref (empty on push)
seed-fallback-dir (empty) absolute path to seed from when no snapshot exists — for migrating off an existing flat cache
watermark-file .ci-watermark-<job>-sha must differ per job when two jobs share one cache key
lock-id <job>-<run_id> identifies this job's cache lock
stale-lock-seconds 7200 age past which another job's lock is treated as abandoned

Outputs: target-dir, cache-key, seeded-from (own | base-snapshot | own-snapshot | fallback-dir | concurrent-peer | cold).

Exports to the job environment: CARGO_TARGET_DIR, CARGO_CACHE_ROOT, CARGO_CACHE_KEY, CARGO_CACHE_LOCK_ID, CARGO_CACHE_SCRIPTS, CI_WATERMARK_FILE.

cargo-cache-publish

input default meaning
cache-root /cache must match the consume action
protected-branches dev main refs that publish snapshots
mode publish publish, or release-lock for the if: always() step
own-ref (auto) override; defaults to github.head_ref, else github.ref_name
publish-on-events push events on which a protected ref actually publishes
read-grace-seconds 300 how long the swap waits for in-flight clones of the generation it replaces before reclaiming it; on timeout the reclamation is deferred, never forced
reader-stale-seconds 7200 age past which a consumer's read marker is treated as abandoned by a killed job
record-watermark true record HEAD as this cache's watermark (PR runs too)

publish-on-events defaults to push on purpose: a pull_request run from dev into main has own-ref dev and would otherwise publish a snapshot of a merge-preview build, which is not what dev is.


Multiple jobs in one workflow

Jobs sharing a cache key (a ci job and a wasm job on the same branch, say) each need their own watermark file. A shared one breaks the moment two jobs run in sequence within one trigger: job A advances the watermark to HEAD, and job B then reads that just-advanced value, computes an empty diff, and loses the merge protection entirely. The default (.ci-watermark-<job>-sha) already gives each job its own; only override watermark-file if you also override lock-id, and then keep both distinct per job.


Constraints of this runner

  • The repository must be public. act_runner fetches actions by anonymous git clone and has no credentialed-fetch option, so a private action repository simply fails to resolve. Nothing secret goes in here.
  • uses: needs the absolute URL. A bare owner/repo resolves against github.com, because Gitea's DEFAULT_ACTIONS_URL is unset — and it has to stay unset, or actions/checkout, dtolnay/rust-toolchain and taiki-e/install-action stop resolving.
  • The runner pre-fetches every referenced action before running any step, so a bad action reference fails the job at step 0 rather than where it is used.
  • container.volumes is job-level and cannot be set from inside a composite action. The consuming workflow declares it; see the quick start.
  • The cache volume is ext4 — no reflink support, which is precisely why hardlinks are the mechanism that makes cloning cheap.

Versioning

Pin @v1. It is a moving major tag: fixes and backward-compatible inputs move it forward, and anything that would break an existing consumer gets v2 instead. Pin a commit SHA if you want a frozen version.


Development

bash scripts/selftest.sh          # everything (~1 min; needs cargo)
bash scripts/selftest.sh --fast   # fixture-only suites, no compiler
suite covers
hardlink-clone-selftest.sh that a build in a clone cannot mutate its source — with a control proving a raw cp -al does. Needs a real compiler.
seed-target-dir-selftest.sh seed-source preference, lock-file stripping, two jobs racing on one cache key, and a seed racing a publisher's rotation of the source it is reading — the race that actually truncates a tree
publish-snapshot-selftest.sh the atomic swap, that a live consumer survives a republish, and the publisher's side of the rotation race: deferred reclamation under a live reader, and its sweep once the reader is gone
prune-cache-selftest.sh liveness, protection, locking, eviction order, self-clear — against a real scratch origin
restore-mtimes-selftest.sh the merge hazard and the watermark that closes it, including the two-jobs-one-namespace case. Needs a real compiler.

Every suite runs the actual script, not a reimplementation of its logic, and every fix scenario is paired with a control that reproduces the bug — a scenario that passes either way proves nothing. The concurrency scenarios race real processes rather than mocking the interleaving, and gate the interfering step on observed progress of the step it interferes with, so the window is hit deterministically instead of on a fast machine's coin flip.

The action YAML holds no logic beyond wiring; everything testable lives in scripts/. A composite action needs shell: bash on every run: step, and the actions reach their shared scripts through ${{ github.action_path }}/../scripts, which works because the runner clones the whole repository when it fetches an action.