The shared action's justification over zemyna's and emowheel's schemes was that hardlink-cloning from a published snapshot closes gitdan #911 "by construction, not by the single job slot". Review disproved that. This makes the claim true, and corrects the README where it could only be bounded. Finding 1 (verdict-level) — silent partial clone ------------------------------------------------ `hardlink_clone_into` ran `cp -al` with no exit-status check, and both call sites invoked it as a condition, which suppresses `set -e` for the whole call. A publisher's `rm -rf` of the generation it rotated away therefore unlinked entries beneath an in-flight consumer walk, and the truncated tree was renamed into place and reported as success. Both layers are fixed: * The consumer verifies its own clone. Every attempt checks `cp -al`'s status explicitly, the source directory's inode before and after (a wholesale replacement mid-walk splices two generations), and the entry count — the only signal for a subtree unlinked before its parent was listed, since `cp -al` reports no error for one it never saw. Any failure discards the staging tree and retries; exhausting the attempts returns a distinct status 2 and fails the job rather than seeding a partial cache. `unshare_subtree` / `_unshare_files` now propagate failure too — a swallowed unshare leaves the clone aliasing its source, the exact corruption that step exists to prevent. * The publisher does not unlink under a reader. A consumer publishes a `.reading-<snapshot>-<tag>` marker before it resolves the snapshot path; the publisher scans for markers after its first rename. A consumer holding the old generation therefore published its marker before that scan and cannot be missed; one arriving after the scan necessarily resolves to the new generation. The publisher waits for readers to drain and, on timeout, DEFERS reclamation rather than forcing it — the old generation is left as `.publish-old-<key>-<tag>` and swept by a later publish. So correctness is closed by construction; disk reclamation is bounded, not immediate. The residual is capped at one deferred generation per publisher ref, and the README now says exactly that instead of the disproved claim. Finding 2 — restore-mtimes.sh ran with no errexit ------------------------------------------------- `set -euo pipefail` was glued to the end of a comment (`# soundness.set -euo pipefail`), so it was entirely commented out: a partial failure of the `git log | awk` pipeline would have produced wrong mtimes across the whole restore instead of failing loudly. Moved to its own line. Audited every other script for the same defect — this was the only instance. Independent confirmation: shellcheck's two SC2164 warnings on this file's `cd "$repo_root"` disappear now that errexit is actually in effect. Finding 3 — lock-acquire window ------------------------------- A just-seeded directory was unlocked until a later action step, so a concurrent job's prune pass could evict it. `seed-target-dir.sh` now takes an optional lock-id and writes the lock marker on every path out of the script, including into the staging tree before its rename, so the directory carries a lock the instant it appears under its final name. The action's acquire step stays (it is idempotent and stamps the LRU marker). Also hardened `prune-cache.sh` to treat a directory with live reader markers as locked. Today no reachable configuration prunes a snapshot — only protected refs publish them and protected refs are excluded from every pass — so this is redundant by policy; it is here so that stops being the reason it is safe. Verification ------------ New selftest scenario 8 races a real seed against a real publish rotation, gating the rotation on the seed's *observed* clone progress so the window is hit deterministically rather than on a fast machine's coin flip. Red-proven against the unguarded scripts, three consecutive runs: ASSERTION FAILED: the seeded tree is truncated: 15443 entries against the snapshot's 493 (was 48805 before the rotation) (15443 / 16986 / 16498) Green after the fix, six consecutive runs, catching the clone mid-walk at ~10.5k of 48805 entries each time. Scenario 9 covers deferred reclamation and its later sweep; scenario 10 covers an unreadable source failing loudly. `bash scripts/selftest.sh`: 5 suites, exit 0, 75 assertions (was 63). shellcheck over `scripts/`: no new findings, two SC2164 warnings resolved. Docs: README's republish-safety paragraph replaced with what the code now guarantees, including the bounded disk residual stated explicitly; new `read-grace-seconds` / `reader-stale-seconds` inputs documented in the `cargo-cache-publish` table; the selftest table names the new race. Refs: daniel/gitdan#11, zemyna#911 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Sqh2vscfzisk83VuPVQX9L
205 lines
8.7 KiB
YAML
205 lines
8.7 KiB
YAML
name: 'Cargo cache (consume)'
|
|
description: >-
|
|
Per-ref Cargo target-dir cache for Gitea Actions runners: hardlink-clones
|
|
this ref's target directory from its base branch's published, immutable
|
|
snapshot, restores git-history file mtimes, and prunes the volume.
|
|
author: 'gitdan'
|
|
|
|
inputs:
|
|
cache-root:
|
|
description: 'Mount point of the persistent cache volume inside the job container.'
|
|
required: false
|
|
default: '/cache'
|
|
protected-branches:
|
|
description: >-
|
|
Space-separated refs that publish snapshots and are never evicted.
|
|
These are the branches PR caches layer over.
|
|
required: false
|
|
default: 'dev main'
|
|
min-free-percent:
|
|
description: 'Prune when free space on the cache volume drops below this percentage.'
|
|
required: false
|
|
default: '10'
|
|
restore-mtimes:
|
|
description: >-
|
|
Restore every tracked file's mtime from git history. Requires a
|
|
full-history checkout (fetch-depth: 0). Set to false only if the build
|
|
does not use Cargo's mtime-based freshness at all.
|
|
required: false
|
|
default: 'true'
|
|
prune:
|
|
description: 'Run the eviction pass (dead-branch liveness + disk pressure).'
|
|
required: false
|
|
default: 'true'
|
|
liveness-prune:
|
|
description: >-
|
|
Within the prune pass, remove caches for branches that no longer exist
|
|
on origin. Set to false on a runner that cannot reach origin.
|
|
required: false
|
|
default: 'true'
|
|
own-ref:
|
|
description: 'Override this run''s ref. Defaults to github.head_ref, else github.ref_name.'
|
|
required: false
|
|
default: ''
|
|
base-ref:
|
|
description: 'Override the ref to layer over. Defaults to github.base_ref (empty on push).'
|
|
required: false
|
|
default: ''
|
|
seed-fallback-dir:
|
|
description: >-
|
|
Absolute path to seed from when no snapshot exists yet — a pre-existing
|
|
flat cache directory during a migration. Optional.
|
|
required: false
|
|
default: ''
|
|
watermark-file:
|
|
description: >-
|
|
Name of this job's build-watermark file inside the target dir. MUST be
|
|
distinct per job when two jobs share one cache key. Defaults to
|
|
.ci-watermark-<job>-sha.
|
|
required: false
|
|
default: ''
|
|
lock-id:
|
|
description: 'Identifier for this job''s cache lock. Defaults to <job>-<run_id>.'
|
|
required: false
|
|
default: ''
|
|
stale-lock-seconds:
|
|
description: 'Age past which another job''s cache lock is treated as abandoned.'
|
|
required: false
|
|
default: '7200'
|
|
|
|
outputs:
|
|
target-dir:
|
|
description: 'Resolved CARGO_TARGET_DIR. Also exported to the job environment.'
|
|
value: ${{ steps.resolve.outputs.target-dir }}
|
|
cache-key:
|
|
description: 'Sanitized cache key for this run''s own ref.'
|
|
value: ${{ steps.resolve.outputs.cache-key }}
|
|
seeded-from:
|
|
description: 'Where the target dir came from: own | base-snapshot | own-snapshot | fallback-dir | concurrent-peer | cold.'
|
|
value: ${{ steps.seed.outputs.seeded-from }}
|
|
|
|
runs:
|
|
using: 'composite'
|
|
steps:
|
|
# Resolves both cache keys and exports the environment every later step
|
|
# (and the consuming job's own build steps) reads. Must run before
|
|
# anything that touches CARGO_TARGET_DIR, which is why it is first.
|
|
#
|
|
# `head_ref || ref_name` rather than `ref_name` alone: on a pull_request
|
|
# event `ref_name` is a synthetic merge-ref name that changes on every
|
|
# push to the PR, so keying on it would give the same PR a different cache
|
|
# directory every time — defeating the reuse this action exists to
|
|
# provide. On a push event `head_ref` is empty and `ref_name` is the real
|
|
# branch, which is what we want there.
|
|
#
|
|
# `base_ref` is populated only for pull_request events. A push run has
|
|
# nothing to layer over: its own ref IS the reference branch. It
|
|
# publishes, it does not consume.
|
|
- id: resolve
|
|
shell: bash
|
|
run: |
|
|
set -euo pipefail
|
|
SCRIPTS=$(cd "${{ github.action_path }}/.." && pwd)/scripts
|
|
[ -d "$SCRIPTS" ] || { echo "::error::cargo-cache: scripts/ not found at $SCRIPTS"; exit 1; }
|
|
echo "CARGO_CACHE_SCRIPTS=${SCRIPTS}" >> "$GITHUB_ENV"
|
|
OWN_REF="${{ inputs.own-ref }}"
|
|
[ -n "$OWN_REF" ] || OWN_REF="${{ github.head_ref || github.ref_name }}"
|
|
BASE_REF="${{ inputs.base-ref }}"
|
|
[ -n "$BASE_REF" ] || BASE_REF="${{ github.base_ref }}"
|
|
|
|
OWN_KEY=$(bash "${SCRIPTS}/branch-cache-key.sh" "$OWN_REF")
|
|
BASE_KEY=""
|
|
if [ -n "$BASE_REF" ]; then
|
|
BASE_KEY=$(bash "${SCRIPTS}/branch-cache-key.sh" "$BASE_REF")
|
|
echo "cache: own ref '${OWN_REF}' -> ${OWN_KEY}; layering over base ref '${BASE_REF}' -> ${BASE_KEY}"
|
|
else
|
|
echo "cache: own ref '${OWN_REF}' -> ${OWN_KEY} (no base ref — this ref publishes, it does not consume)"
|
|
fi
|
|
|
|
TARGET_DIR="${{ inputs.cache-root }}/target-${OWN_KEY}"
|
|
WATERMARK="${{ inputs.watermark-file }}"
|
|
[ -n "$WATERMARK" ] || WATERMARK=".ci-watermark-${{ github.job }}-sha"
|
|
LOCK_ID="${{ inputs.lock-id }}"
|
|
[ -n "$LOCK_ID" ] || LOCK_ID="${{ github.job }}-${{ github.run_id }}"
|
|
|
|
{
|
|
echo "target-dir=${TARGET_DIR}"
|
|
echo "cache-key=${OWN_KEY}"
|
|
echo "base-key=${BASE_KEY}"
|
|
echo "lock-id=${LOCK_ID}"
|
|
echo "watermark-file=${WATERMARK}"
|
|
} >> "$GITHUB_OUTPUT"
|
|
{
|
|
echo "CARGO_TARGET_DIR=${TARGET_DIR}"
|
|
echo "CARGO_CACHE_ROOT=${{ inputs.cache-root }}"
|
|
echo "CARGO_CACHE_KEY=${OWN_KEY}"
|
|
echo "CARGO_CACHE_LOCK_ID=${LOCK_ID}"
|
|
echo "CI_WATERMARK_FILE=${WATERMARK}"
|
|
} >> "$GITHUB_ENV"
|
|
|
|
# Seeds this ref's target dir from the base's published snapshot. See
|
|
# scripts/seed-target-dir.sh — the staging-then-atomic-rename is what
|
|
# makes concurrent jobs sharing one cache key safe by construction rather
|
|
# than by the runner happening to have a single execution slot, and the
|
|
# clone's own consistency check plus the publish side's reader interlock
|
|
# are what make it safe against the base republishing MID-CLONE.
|
|
#
|
|
# The lock id is passed here as well as acquired in the next step: the
|
|
# seed writes it into the staging tree, so the directory carries a lock
|
|
# the instant it appears under its final name rather than a step later.
|
|
- id: seed
|
|
shell: bash
|
|
run: |
|
|
set -euo pipefail
|
|
SCRIPTS=$(cd "${{ github.action_path }}/.." && pwd)/scripts
|
|
bash "${SCRIPTS}/seed-target-dir.sh" \
|
|
"${{ steps.resolve.outputs.cache-key }}" \
|
|
"${{ steps.resolve.outputs.base-key }}" \
|
|
"${{ inputs.cache-root }}" \
|
|
"${{ github.job }}-${{ github.run_id }}-$$" \
|
|
"${{ inputs.seed-fallback-dir }}" \
|
|
"${{ steps.resolve.outputs.lock-id }}"
|
|
|
|
# Re-stamps the lock the seed step already wrote (acquiring is idempotent
|
|
# — it rewrites the timestamp) and stamps the LRU marker. The marker is
|
|
# touched unconditionally every run: a run that hits the cache for every
|
|
# crate may write nothing at all inside the tree, which would make a
|
|
# just-used directory look stale to the eviction pass.
|
|
- shell: bash
|
|
run: |
|
|
set -euo pipefail
|
|
SCRIPTS=$(cd "${{ github.action_path }}/.." && pwd)/scripts
|
|
bash "${SCRIPTS}/cache-lock.sh" acquire \
|
|
"${{ steps.resolve.outputs.target-dir }}" "${{ steps.resolve.outputs.lock-id }}"
|
|
touch "${{ steps.resolve.outputs.target-dir }}/.cache-last-used"
|
|
|
|
# Runs AFTER seeding, deliberately: restore-mtimes.sh reads the build
|
|
# watermark out of the target dir, so that directory has to be in its
|
|
# final form for this run (reused, seeded, or freshly created) before the
|
|
# watermark it may carry can be read.
|
|
# CARGO_TARGET_DIR and CI_WATERMARK_FILE are passed explicitly rather
|
|
# than read from the job environment the resolve step exported: the export
|
|
# is what the consuming workflow's own build steps rely on, but a step
|
|
# inside this action should not depend on cross-step propagation working
|
|
# when it can just be handed the value.
|
|
- if: ${{ inputs.restore-mtimes == 'true' }}
|
|
shell: bash
|
|
env:
|
|
CARGO_TARGET_DIR: ${{ steps.resolve.outputs.target-dir }}
|
|
CI_WATERMARK_FILE: ${{ steps.resolve.outputs.watermark-file }}
|
|
run: bash "$(cd "${{ github.action_path }}/.." && pwd)/scripts/restore-mtimes.sh"
|
|
|
|
- if: ${{ inputs.prune == 'true' }}
|
|
shell: bash
|
|
env:
|
|
STALE_LOCK_SECONDS: ${{ inputs.stale-lock-seconds }}
|
|
CACHE_LIVENESS: ${{ inputs.liveness-prune }}
|
|
run: |
|
|
set -euo pipefail
|
|
SCRIPTS=$(cd "${{ github.action_path }}/.." && pwd)/scripts
|
|
bash "${SCRIPTS}/prune-cache.sh" \
|
|
"${{ inputs.cache-root }}" \
|
|
"${{ steps.resolve.outputs.target-dir }}" \
|
|
"${{ inputs.protected-branches }}" \
|
|
"${{ inputs.min-free-percent }}"
|