The shared action's justification over zemyna's and emowheel's schemes was that hardlink-cloning from a published snapshot closes gitdan #911 "by construction, not by the single job slot". Review disproved that. This makes the claim true, and corrects the README where it could only be bounded. Finding 1 (verdict-level) — silent partial clone ------------------------------------------------ `hardlink_clone_into` ran `cp -al` with no exit-status check, and both call sites invoked it as a condition, which suppresses `set -e` for the whole call. A publisher's `rm -rf` of the generation it rotated away therefore unlinked entries beneath an in-flight consumer walk, and the truncated tree was renamed into place and reported as success. Both layers are fixed: * The consumer verifies its own clone. Every attempt checks `cp -al`'s status explicitly, the source directory's inode before and after (a wholesale replacement mid-walk splices two generations), and the entry count — the only signal for a subtree unlinked before its parent was listed, since `cp -al` reports no error for one it never saw. Any failure discards the staging tree and retries; exhausting the attempts returns a distinct status 2 and fails the job rather than seeding a partial cache. `unshare_subtree` / `_unshare_files` now propagate failure too — a swallowed unshare leaves the clone aliasing its source, the exact corruption that step exists to prevent. * The publisher does not unlink under a reader. A consumer publishes a `.reading-<snapshot>-<tag>` marker before it resolves the snapshot path; the publisher scans for markers after its first rename. A consumer holding the old generation therefore published its marker before that scan and cannot be missed; one arriving after the scan necessarily resolves to the new generation. The publisher waits for readers to drain and, on timeout, DEFERS reclamation rather than forcing it — the old generation is left as `.publish-old-<key>-<tag>` and swept by a later publish. So correctness is closed by construction; disk reclamation is bounded, not immediate. The residual is capped at one deferred generation per publisher ref, and the README now says exactly that instead of the disproved claim. Finding 2 — restore-mtimes.sh ran with no errexit ------------------------------------------------- `set -euo pipefail` was glued to the end of a comment (`# soundness.set -euo pipefail`), so it was entirely commented out: a partial failure of the `git log | awk` pipeline would have produced wrong mtimes across the whole restore instead of failing loudly. Moved to its own line. Audited every other script for the same defect — this was the only instance. Independent confirmation: shellcheck's two SC2164 warnings on this file's `cd "$repo_root"` disappear now that errexit is actually in effect. Finding 3 — lock-acquire window ------------------------------- A just-seeded directory was unlocked until a later action step, so a concurrent job's prune pass could evict it. `seed-target-dir.sh` now takes an optional lock-id and writes the lock marker on every path out of the script, including into the staging tree before its rename, so the directory carries a lock the instant it appears under its final name. The action's acquire step stays (it is idempotent and stamps the LRU marker). Also hardened `prune-cache.sh` to treat a directory with live reader markers as locked. Today no reachable configuration prunes a snapshot — only protected refs publish them and protected refs are excluded from every pass — so this is redundant by policy; it is here so that stops being the reason it is safe. Verification ------------ New selftest scenario 8 races a real seed against a real publish rotation, gating the rotation on the seed's *observed* clone progress so the window is hit deterministically rather than on a fast machine's coin flip. Red-proven against the unguarded scripts, three consecutive runs: ASSERTION FAILED: the seeded tree is truncated: 15443 entries against the snapshot's 493 (was 48805 before the rotation) (15443 / 16986 / 16498) Green after the fix, six consecutive runs, catching the clone mid-walk at ~10.5k of 48805 entries each time. Scenario 9 covers deferred reclamation and its later sweep; scenario 10 covers an unreadable source failing loudly. `bash scripts/selftest.sh`: 5 suites, exit 0, 75 assertions (was 63). shellcheck over `scripts/`: no new findings, two SC2164 warnings resolved. Docs: README's republish-safety paragraph replaced with what the code now guarantees, including the bounded disk residual stated explicitly; new `read-grace-seconds` / `reader-stale-seconds` inputs documented in the `cargo-cache-publish` table; the selftest table names the new race. Refs: daniel/gitdan#11, zemyna#911 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Sqh2vscfzisk83VuPVQX9L
308 lines
15 KiB
Markdown
308 lines
15 KiB
Markdown
# gitdan-actions
|
|
|
|
Shared Gitea Actions composite actions for the gitdan forge.
|
|
|
|
Currently one thing, done properly: **`cargo-cache`** — a persistent,
|
|
per-branch Cargo build cache for self-hosted Gitea runners, where a pull
|
|
request's cache is a near-free hardlink clone of an immutable snapshot its
|
|
base branch published.
|
|
|
|
> **Final home:** this repository will live at `daniel/gitdan-actions`. Pin
|
|
> that path in `uses:` once the transfer completes.
|
|
|
|
---
|
|
|
|
## Why this exists
|
|
|
|
Two of this forge's Rust projects independently built the same idea and each
|
|
got one half right.
|
|
|
|
| | seeding mechanism | seed source |
|
|
|---|---|---|
|
|
| project A | `cp -al` hardlink clone — near-free, cost scales with inode count, not bytes | the base branch's **live** target dir — races a build that is still writing |
|
|
| project B | `cp -a` full copy — sound, but ~35 GB duplicated per branch | a **published immutable snapshot** — nothing ever writes it while it is read |
|
|
|
|
This action is the diagonal: **hardlink-clone from a published snapshot.**
|
|
Cheap like A, sound like B. It also closes a latent race in A by construction
|
|
(the seed is staged and swapped in with one atomic rename) rather than relying
|
|
on the runner having a single execution slot.
|
|
|
|
One thing neither project had, and the reason the clone is not a plain
|
|
`cp -al`: **a build inside a hardlink clone does mutate the directory it was
|
|
cloned from.** Cargo writes real artifacts by replacing them, but writes its
|
|
metadata — and build scripts write their `OUT_DIR` — with a plain truncating
|
|
write, straight through the shared inode. Under
|
|
`CARGO_UNSTABLE_CHECKSUM_FRESHNESS` the file that gets corrupted is
|
|
`.fingerprint/<unit>/dep-<target>`, which holds the per-source checksums that
|
|
decide freshness, and the failure is silent stale-artifact reuse rather than a
|
|
slow build. `scripts/hardlink-clone-selftest.sh` reproduces it as an explicit
|
|
control and asserts the fix. The fix is to hardlink the artifacts (the GB) and
|
|
real-copy the metadata (the MB) — about 3.7% of a Bevy-sized target directory,
|
|
against 100% for a full copy.
|
|
|
|
---
|
|
|
|
## Quick start
|
|
|
|
```yaml
|
|
name: CI
|
|
on:
|
|
push:
|
|
branches: [main, dev]
|
|
pull_request:
|
|
branches: [main, dev]
|
|
|
|
jobs:
|
|
ci:
|
|
runs-on: ubuntu-latest
|
|
# REQUIRED, and it cannot come from the action: `container.volumes` is a
|
|
# job-level property, so the persistent cache volume must be declared
|
|
# here. Use a volume name unique to this repository.
|
|
container:
|
|
volumes:
|
|
- myrepo-ci-target:/cache
|
|
steps:
|
|
- uses: actions/checkout@v4
|
|
with:
|
|
# REQUIRED. The mtime restore walks every commit that ever touched a
|
|
# tracked file; a depth-1 checkout makes every file resolve to the
|
|
# tip commit and the cache stops working. The action fails loudly
|
|
# rather than silently degrading if this is missing.
|
|
fetch-depth: 0
|
|
|
|
- uses: https://gitdan.com/daniel/gitdan-actions/cargo-cache@v1
|
|
with:
|
|
protected-branches: 'dev main'
|
|
|
|
# ... toolchain, system deps, and the build itself. CARGO_TARGET_DIR is
|
|
# already exported to the job environment by the step above.
|
|
- run: cargo clippy --workspace --all-targets -- -D warnings
|
|
- run: cargo test --workspace
|
|
|
|
# After the build succeeds: record the watermark, and publish a snapshot
|
|
# if this run is a push to a protected branch.
|
|
- uses: https://gitdan.com/daniel/gitdan-actions/cargo-cache-publish@v1
|
|
|
|
# Release this job's cache lock even when the build failed, so the
|
|
# eviction pass does not have to wait out the staleness grace period.
|
|
- if: always()
|
|
uses: https://gitdan.com/daniel/gitdan-actions/cargo-cache-publish@v1
|
|
with:
|
|
mode: release-lock
|
|
```
|
|
|
|
Recommended alongside it, in the workflow's `env:` block:
|
|
|
|
```yaml
|
|
env:
|
|
CARGO_INCREMENTAL: 0 # per-run bloat on a persistent volume
|
|
CARGO_PROFILE_DEV_DEBUG: line-tables-only
|
|
CARGO_PROFILE_TEST_DEBUG: line-tables-only
|
|
# Nightly only. Content-addressed freshness instead of mtime-based — a
|
|
# strictly stronger guarantee, complementary to the mtime restore (which
|
|
# still covers directory-form `rerun-if-changed` build-script watches).
|
|
CARGO_UNSTABLE_CHECKSUM_FRESHNESS: "true"
|
|
```
|
|
|
|
---
|
|
|
|
## How it works
|
|
|
|
```
|
|
/cache/
|
|
target-<key> one per ref. Where a build actually runs.
|
|
snapshot-<key> one per publisher ref. Immutable between publishes;
|
|
the only thing a consumer ever clones from.
|
|
```
|
|
|
|
`<key>` is the ref sanitised to a safe path component, capped at 48
|
|
characters, plus an 8-hex SHA-1 prefix of the *raw* ref. The hash is not
|
|
decoration: `feat/foo` and `feat-foo` sanitise identically and would otherwise
|
|
share one directory.
|
|
|
|
**A pull request run** resolves its own key from `github.head_ref` (not
|
|
`ref_name`, which on a `pull_request` event is a synthetic merge ref that
|
|
changes on every push) and its base key from `github.base_ref`. If it has no
|
|
directory yet, it hardlink-clones `snapshot-<base>` into a staging path, strips
|
|
Cargo's lock files, real-copies everything Cargo writes in place, and renames
|
|
the staging path into `target-<own>`.
|
|
|
|
**A push to a protected branch** has no base to layer over. It builds in its
|
|
own directory and, if the build goes green, republishes it as
|
|
`snapshot-<own>`: stage a clone, rename the old snapshot aside, rename the new
|
|
one in, then reclaim the old one once nothing is still reading it. Consumers
|
|
only ever observe a complete snapshot or none at all.
|
|
|
|
**Concurrency, on the destination.** Two jobs sharing one cache key each stage
|
|
under their own tag and race on one atomic rename; the loser discards its
|
|
staging copy. There is no window in which a partially-populated directory is
|
|
visible under the final name. Two jobs then building in the same directory is
|
|
Cargo's own `.cargo-lock` territory, which is what that lock is for.
|
|
|
|
**Concurrency, on the source.** The atomic rename is necessary and not
|
|
sufficient, because renaming a *truncated* tree publishes a truncated tree
|
|
atomically. A clone reads its source over many seconds, and a publisher
|
|
rotating that source unlinks the generation being read — at which point
|
|
`cp -al` can silently omit a subtree it never saw, and report success. Two
|
|
mechanisms, both required:
|
|
|
|
- **The publisher does not unlink under a reader.** A consumer publishes a
|
|
`.reading-<snapshot>-<tag>` marker *before* it resolves the snapshot path;
|
|
the publisher scans for markers *after* its first rename. A consumer holding
|
|
the old generation therefore published its marker before that scan and
|
|
cannot be missed, and one that arrives after the scan necessarily resolves
|
|
to the new generation. The publisher waits for readers to drain
|
|
(`read-grace-seconds`, default 300) and, if they do not, **defers** the
|
|
reclamation rather than forcing it — the old generation stays on disk and is
|
|
swept by a later publish.
|
|
- **The consumer verifies its own clone.** Every attempt checks `cp -al`'s
|
|
exit status, the source directory's inode before and after (a wholesale
|
|
replacement mid-walk would otherwise splice two generations), and the entry
|
|
count (the only signal for a subtree unlinked before its parent was listed —
|
|
there is no error to read). A tree that fails any of the three is deleted
|
|
and the clone retried; one that fails the last attempt fails the job. A
|
|
partial tree never reaches the final name.
|
|
|
|
**What this does and does not guarantee.** *Correctness* is closed by
|
|
construction: no combination of publish and seed timing produces a target
|
|
directory holding part of one generation, and a source that cannot be read
|
|
consistently fails the job loudly instead of seeding a truncated cache.
|
|
*Disk reclamation* is bounded, not immediate: a consumer slower than the grace
|
|
period leaves one extra snapshot generation of directory entries on the volume
|
|
until the next publish of that snapshot sweeps it. That residual is capped at
|
|
one deferred generation per publisher ref, and its real cost is close to the
|
|
inode count rather than the byte count, since the artifacts are hardlinked to
|
|
whatever cloned them.
|
|
|
|
**Eviction** runs three passes: caches for branches that no longer exist on
|
|
origin are removed unconditionally; then, only if free space is under the
|
|
threshold, live caches are evicted oldest-first; then, as a last resort, this
|
|
run's own cache. Protected refs and any cache held open by a running job are
|
|
never candidates. Within the pressure pass, `target-*` directories are evicted
|
|
before `snapshot-*` ones — the reverse of the obvious order, because a
|
|
snapshot is hardlinked to everything cloned from it, so removing one frees
|
|
almost no real bytes while costing every future PR its warm start.
|
|
|
|
**File mtimes.** `actions/checkout` stamps every file with "now", which makes
|
|
every crate look changed to Cargo's mtime-based freshness check — a persistent
|
|
target directory buys nothing without fixing that. Each tracked file is
|
|
restored to the timestamp of the most recent commit that touched it, plus a
|
|
watermark override: for any file that changed since *this cache's own last
|
|
successful build*, "now" is stamped instead. That override is what makes a
|
|
merge safe, since a merge can introduce a commit authored before this cache's
|
|
last build, where the historically-correct mtime is exactly the wrong answer.
|
|
|
|
---
|
|
|
|
## Inputs
|
|
|
|
### `cargo-cache`
|
|
|
|
| input | default | meaning |
|
|
|---|---|---|
|
|
| `cache-root` | `/cache` | mount point of the persistent volume inside the job container |
|
|
| `protected-branches` | `dev main` | refs that publish snapshots and are never evicted |
|
|
| `min-free-percent` | `10` | prune when free space drops below this |
|
|
| `restore-mtimes` | `true` | restore tracked-file mtimes from git history |
|
|
| `prune` | `true` | run the eviction pass |
|
|
| `liveness-prune` | `true` | within eviction, remove caches for branches gone from origin |
|
|
| `own-ref` | *(auto)* | override; defaults to `github.head_ref`, else `github.ref_name` |
|
|
| `base-ref` | *(auto)* | override; defaults to `github.base_ref` (empty on push) |
|
|
| `seed-fallback-dir` | *(empty)* | absolute path to seed from when no snapshot exists — for migrating off an existing flat cache |
|
|
| `watermark-file` | `.ci-watermark-<job>-sha` | must differ per job when two jobs share one cache key |
|
|
| `lock-id` | `<job>-<run_id>` | identifies this job's cache lock |
|
|
| `stale-lock-seconds` | `7200` | age past which another job's lock is treated as abandoned |
|
|
|
|
Outputs: `target-dir`, `cache-key`, `seeded-from` (`own` \| `base-snapshot` \|
|
|
`own-snapshot` \| `fallback-dir` \| `concurrent-peer` \| `cold`).
|
|
|
|
Exports to the job environment: `CARGO_TARGET_DIR`, `CARGO_CACHE_ROOT`,
|
|
`CARGO_CACHE_KEY`, `CARGO_CACHE_LOCK_ID`, `CARGO_CACHE_SCRIPTS`,
|
|
`CI_WATERMARK_FILE`.
|
|
|
|
### `cargo-cache-publish`
|
|
|
|
| input | default | meaning |
|
|
|---|---|---|
|
|
| `cache-root` | `/cache` | must match the consume action |
|
|
| `protected-branches` | `dev main` | refs that publish snapshots |
|
|
| `mode` | `publish` | `publish`, or `release-lock` for the `if: always()` step |
|
|
| `own-ref` | *(auto)* | override; defaults to `github.head_ref`, else `github.ref_name` |
|
|
| `publish-on-events` | `push` | events on which a protected ref actually publishes |
|
|
| `read-grace-seconds` | `300` | how long the swap waits for in-flight clones of the generation it replaces before reclaiming it; on timeout the reclamation is deferred, never forced |
|
|
| `reader-stale-seconds` | `7200` | age past which a consumer's read marker is treated as abandoned by a killed job |
|
|
| `record-watermark` | `true` | record HEAD as this cache's watermark (PR runs too) |
|
|
|
|
`publish-on-events` defaults to `push` on purpose: a `pull_request` run from
|
|
`dev` into `main` has `own-ref` `dev` and would otherwise publish a snapshot
|
|
of a merge-preview build, which is not what `dev` is.
|
|
|
|
---
|
|
|
|
## Multiple jobs in one workflow
|
|
|
|
Jobs sharing a cache key (a `ci` job and a `wasm` job on the same branch, say)
|
|
each need their **own** watermark file. A shared one breaks the moment two
|
|
jobs run in sequence within one trigger: job A advances the watermark to HEAD,
|
|
and job B then reads that just-advanced value, computes an empty diff, and
|
|
loses the merge protection entirely. The default (`.ci-watermark-<job>-sha`)
|
|
already gives each job its own; only override `watermark-file` if you also
|
|
override `lock-id`, and then keep both distinct per job.
|
|
|
|
---
|
|
|
|
## Constraints of this runner
|
|
|
|
- **The repository must be public.** act_runner fetches actions by anonymous
|
|
git clone and has no credentialed-fetch option, so a private action
|
|
repository simply fails to resolve. Nothing secret goes in here.
|
|
- **`uses:` needs the absolute URL.** A bare `owner/repo` resolves against
|
|
github.com, because Gitea's `DEFAULT_ACTIONS_URL` is unset — and it has to
|
|
stay unset, or `actions/checkout`, `dtolnay/rust-toolchain` and
|
|
`taiki-e/install-action` stop resolving.
|
|
- **The runner pre-fetches every referenced action before running any step**,
|
|
so a bad action reference fails the job at step 0 rather than where it is
|
|
used.
|
|
- **`container.volumes` is job-level** and cannot be set from inside a
|
|
composite action. The consuming workflow declares it; see the quick start.
|
|
- **The cache volume is ext4** — no reflink support, which is precisely why
|
|
hardlinks are the mechanism that makes cloning cheap.
|
|
|
|
---
|
|
|
|
## Versioning
|
|
|
|
Pin `@v1`. It is a moving major tag: fixes and backward-compatible inputs move
|
|
it forward, and anything that would break an existing consumer gets `v2`
|
|
instead. Pin a commit SHA if you want a frozen version.
|
|
|
|
---
|
|
|
|
## Development
|
|
|
|
```bash
|
|
bash scripts/selftest.sh # everything (~1 min; needs cargo)
|
|
bash scripts/selftest.sh --fast # fixture-only suites, no compiler
|
|
```
|
|
|
|
| suite | covers |
|
|
|---|---|
|
|
| `hardlink-clone-selftest.sh` | that a build in a clone cannot mutate its source — with a control proving a raw `cp -al` does. Needs a real compiler. |
|
|
| `seed-target-dir-selftest.sh` | seed-source preference, lock-file stripping, two jobs racing on one cache key, **and a seed racing a publisher's rotation of the source it is reading** — the race that actually truncates a tree |
|
|
| `publish-snapshot-selftest.sh` | the atomic swap, and that a live consumer survives a republish |
|
|
| `prune-cache-selftest.sh` | liveness, protection, locking, eviction order, self-clear — against a real scratch `origin` |
|
|
| `restore-mtimes-selftest.sh` | the merge hazard and the watermark that closes it, including the two-jobs-one-namespace case. Needs a real compiler. |
|
|
|
|
Every suite runs the actual script, not a reimplementation of its logic, and
|
|
every fix scenario is paired with a control that reproduces the bug — a
|
|
scenario that passes either way proves nothing. The concurrency scenarios race
|
|
real processes rather than mocking the interleaving, and gate the interfering
|
|
step on *observed* progress of the step it interferes with, so the window is
|
|
hit deterministically instead of on a fast machine's coin flip.
|
|
|
|
The action YAML holds no logic beyond wiring; everything testable lives in
|
|
`scripts/`. A composite action needs `shell: bash` on every `run:` step, and
|
|
the actions reach their shared scripts through
|
|
`${{ github.action_path }}/../scripts`, which works because the runner clones
|
|
the whole repository when it fetches an action.
|