Files
gitdan-actions/cargo-cache/action.yml
T
claudeandClaude Sonnet 5 dc473f0d0c
CI / shellcheck + selftests (pull_request) Successful in 1m48s
fix(prune): stop protecting publisher target dirs under disk pressure
A publisher branch's target-<ref> was protected identically to its
snapshot-<ref>, so it was never a pressure-pass candidate however old
and however tight the disk. With one branch's target dir permanently
resident alongside its snapshot, a second branch had no room to seed,
and every PR that night needed a hand eviction between runs
(daniel/gitdan-actions#24).

A publisher's target dir is a convenience cache its own next run
reseeds from the snapshot, so losing it under pressure is cheap;
nothing downstream depends on it surviving. Only the snapshot stays
protected in the pressure and self-clear passes.

Unprotecting the target dir outright surfaced a second bug the fix
would otherwise have shipped: a branch's tip is trivially an ancestor
of itself, so once a protected ref's target dir was no longer skipped
before reaching the merged-branch check, pass 1 read it as "merged
into itself" and deleted it unconditionally on every run, independent
of disk pressure. is_protected_from_liveness keeps a protected ref's
target dir out of pass 1 alone, so it stays an ordinary pressure-pass
candidate without ever reaching that check. Scenario 3b in the
selftest red-proves this against the unprotect-only version of the
fix.

Also corrects the header's inode-sharing claim, measured false on the
live volume by daniel/zemyna#1073: publish-snapshot.sh unshares every
executable after its cp -al, and executables are most of the tree by
bytes, so a snapshot eviction is a real, large disk cost rather than
the near-free one the old text described — the target-before-snapshot
ordering still holds, now for the warm-start reason alone plus that
cost.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 10:42:08 -05:00

282 lines
13 KiB
YAML

name: 'Cargo cache (consume)'
description: >-
Per-ref Cargo target-dir cache for Gitea Actions runners: hardlink-clones
this ref's target directory from its base branch's published, immutable
snapshot, restores git-history file mtimes, and prunes the volume.
author: 'gitdan'
inputs:
cache-root:
description: 'Mount point of the persistent cache volume inside the job container.'
required: false
default: '/cache'
cache-lineage:
description: >-
Distinguishes two jobs that build the SAME ref for different targets or
profiles (a host build and a wasm32 build, say) and would otherwise
resolve to one CARGO_TARGET_DIR and serialise on Cargo's exclusive
build-directory lock. Names one directory level under cache-root:
<cache-root>/<lineage>/target-<key>. Must be a single path component;
several names are refused outright because a reader elsewhere would stop
seeing the caches under them (see validate_cache_lineage in
scripts/cache-lib.sh). Pass the same value to cargo-cache-publish.
required: false
default: ''
protected-branches:
description: >-
Space-separated refs that publish snapshots. A protected ref's
snapshot is never evicted; its own target dir is an ordinary
pressure-pass candidate, reseeded from the snapshot on its next run.
These are the branches PR caches layer over.
required: false
default: 'dev main'
min-free-percent:
description: >-
An ADDITIONAL free-space floor, as a percentage of the cache volume.
The prune's own requirement is derived per run from what the seed is
about to clone — the mutable set it has to real-copy out of the source
snapshot, which is a property of that snapshot and not of the volume —
and this floor only ever raises it. 0, the default, leaves the derived
requirement as the only gate. Set it to keep headroom for something
other than the clone (the build's own output, another job on the same
volume); it will not make the clone fit, because it does not know how
big the clone is.
required: false
default: '0'
restore-mtimes:
description: >-
Restore every tracked file's mtime from git history. Requires a
full-history checkout (fetch-depth: 0). Set to false only if the build
does not use Cargo's mtime-based freshness at all.
required: false
default: 'true'
prune:
description: >-
Run the eviction pass (dead-branch liveness + disk pressure). It runs
BEFORE the seed step, so what it frees is available to the clone that
step makes.
required: false
default: 'true'
liveness-prune:
description: >-
Within the prune pass, remove caches for branches that are dead: gone
from origin, or still on origin with a tip already merged into a
protected branch. The second signal is what reclaims anything at all on
a forge that keeps branches after merge, and it needs the protected
branches' commits in the checkout — a shallow one withholds it and says
so. Set to false on a runner that cannot reach origin.
required: false
default: 'true'
own-ref:
description: 'Override this run''s ref. Defaults to github.head_ref, else github.ref_name.'
required: false
default: ''
base-ref:
description: 'Override the ref to layer over. Defaults to github.base_ref (empty on push).'
required: false
default: ''
seed-fallback-dir:
description: >-
Absolute path to seed from when no snapshot exists yet — a pre-existing
flat cache directory during a migration. Optional.
required: false
default: ''
watermark-file:
description: >-
Name of this job's build-watermark file inside the target dir. MUST be
distinct per job when two jobs share one target directory — which two
jobs no longer need to do; cache-lineage gives them separate ones.
Defaults to .ci-watermark-<job>-sha, already distinct per job.
required: false
default: ''
lock-id:
description: 'Identifier for this job''s cache lock. Defaults to <job>-<run_id>.'
required: false
default: ''
stale-lock-seconds:
description: 'Age past which another job''s cache lock is treated as abandoned.'
required: false
default: '7200'
outputs:
target-dir:
description: 'Resolved CARGO_TARGET_DIR. Also exported to the job environment.'
value: ${{ steps.resolve.outputs.target-dir }}
cache-key:
description: 'Sanitized cache key for this run''s own ref.'
value: ${{ steps.resolve.outputs.cache-key }}
cache-root:
description: >-
Resolved cache root — cache-root, plus the lineage directory when one is
set. Every directory this action reads or writes is under it. Also
exported as CARGO_CACHE_ROOT.
value: ${{ steps.resolve.outputs.cache-root }}
seeded-from:
description: 'Where the target dir came from: own | base-snapshot | own-snapshot | fallback-dir | concurrent-peer | cold.'
value: ${{ steps.seed.outputs.seeded-from }}
runs:
using: 'composite'
steps:
# Resolves both cache keys and exports the environment every later step
# (and the consuming job's own build steps) reads. Must run before
# anything that touches CARGO_TARGET_DIR, which is why it is first.
#
# `head_ref || ref_name` rather than `ref_name` alone: on a pull_request
# event `ref_name` is a synthetic merge-ref name that changes on every
# push to the PR, so keying on it would give the same PR a different cache
# directory every time — defeating the reuse this action exists to
# provide. On a push event `head_ref` is empty and `ref_name` is the real
# branch, which is what we want there.
#
# `base_ref` is populated only for pull_request events. A push run has
# nothing to layer over: its own ref IS the reference branch. It
# publishes, it does not consume.
#
# The cache root is resolved first because every path below hangs off it.
# A lineage nests one directory level (`<root>/<lineage>/target-<key>`),
# which is what lets two jobs on ONE ref hold two build directories and so
# not serialise on Cargo's exclusive lock. It is resolved through
# cache-root.sh rather than interpolated here so the name is validated —
# several otherwise-reasonable lineage names put their whole subtree out of
# reach of a pass that has to see it. Everything downstream reads the
# resolved value, and it is exported as CARGO_CACHE_ROOT so the publish
# action can check it agrees with its own inputs.
- id: resolve
shell: bash
run: |
set -euo pipefail
SCRIPTS=$(cd "${{ github.action_path }}/.." && pwd)/scripts
[ -d "$SCRIPTS" ] || { echo "::error::cargo-cache: scripts/ not found at $SCRIPTS"; exit 1; }
echo "CARGO_CACHE_SCRIPTS=${SCRIPTS}" >> "$GITHUB_ENV"
CACHE_ROOT=$(bash "${SCRIPTS}/cache-root.sh" resolve \
"${{ inputs.cache-root }}" "${{ inputs.cache-lineage }}")
OWN_REF="${{ inputs.own-ref }}"
[ -n "$OWN_REF" ] || OWN_REF="${{ github.head_ref || github.ref_name }}"
BASE_REF="${{ inputs.base-ref }}"
[ -n "$BASE_REF" ] || BASE_REF="${{ github.base_ref }}"
OWN_KEY=$(bash "${SCRIPTS}/branch-cache-key.sh" "$OWN_REF")
BASE_KEY=""
if [ -n "$BASE_REF" ]; then
BASE_KEY=$(bash "${SCRIPTS}/branch-cache-key.sh" "$BASE_REF")
echo "cache: own ref '${OWN_REF}' -> ${OWN_KEY}; layering over base ref '${BASE_REF}' -> ${BASE_KEY}"
else
echo "cache: own ref '${OWN_REF}' -> ${OWN_KEY} (no base ref — this ref publishes, it does not consume)"
fi
TARGET_DIR="${CACHE_ROOT}/target-${OWN_KEY}"
WATERMARK="${{ inputs.watermark-file }}"
[ -n "$WATERMARK" ] || WATERMARK=".ci-watermark-${{ github.job }}-sha"
LOCK_ID="${{ inputs.lock-id }}"
[ -n "$LOCK_ID" ] || LOCK_ID="${{ github.job }}-${{ github.run_id }}"
{
echo "target-dir=${TARGET_DIR}"
echo "cache-key=${OWN_KEY}"
echo "base-key=${BASE_KEY}"
echo "lock-id=${LOCK_ID}"
echo "watermark-file=${WATERMARK}"
echo "cache-root=${CACHE_ROOT}"
} >> "$GITHUB_OUTPUT"
{
echo "CARGO_TARGET_DIR=${TARGET_DIR}"
echo "CARGO_CACHE_ROOT=${CACHE_ROOT}"
echo "CARGO_CACHE_KEY=${OWN_KEY}"
echo "CARGO_CACHE_LOCK_ID=${LOCK_ID}"
echo "CI_WATERMARK_FILE=${WATERMARK}"
} >> "$GITHUB_ENV"
# Runs BEFORE the seed, which is the only order in which its work can
# help: the eviction it performs is what makes room for the clone the seed
# step is about to make, and the requirement it evicts against is measured
# off the snapshot that clone will read. Running afterwards — where this
# step used to be — meant every run freed space for the NEXT one and the
# seed met whatever the last run happened to leave.
#
# Two things this ordering has to be safe against, and is:
#
# The source it is about to read is excluded from every pass by name
# (see protected_reason in prune-cache.sh), so pass 1 cannot take the
# snapshot out from under the seed that follows it.
# A concurrent job's target dir carries its lock from the instant it
# appears under its final name — the seed writes it into the staging
# tree before the rename — so there is no window in which running this
# earlier sees an unlocked directory somebody is using.
- if: ${{ inputs.prune == 'true' }}
shell: bash
env:
STALE_LOCK_SECONDS: ${{ inputs.stale-lock-seconds }}
CACHE_LIVENESS: ${{ inputs.liveness-prune }}
run: |
set -euo pipefail
SCRIPTS=$(cd "${{ github.action_path }}/.." && pwd)/scripts
bash "${SCRIPTS}/prune-cache.sh" \
"${{ steps.resolve.outputs.cache-root }}" \
"${{ steps.resolve.outputs.target-dir }}" \
"${{ inputs.protected-branches }}" \
"${{ inputs.min-free-percent }}" \
"${{ steps.resolve.outputs.cache-key }}" \
"${{ steps.resolve.outputs.base-key }}" \
"${{ inputs.seed-fallback-dir }}"
# Seeds this ref's target dir from the base's published snapshot. See
# scripts/seed-target-dir.sh — the staging-then-atomic-rename is what
# makes concurrent jobs sharing one cache key safe by construction rather
# than by the runner happening to have a single execution slot, and the
# clone's own consistency check plus the publish side's reader interlock
# are what make it safe against the base republishing MID-CLONE.
#
# The lock id is passed here as well as acquired in the next step: the
# seed writes it into the staging tree, so the directory carries a lock
# the instant it appears under its final name rather than a step later.
- id: seed
shell: bash
run: |
set -euo pipefail
SCRIPTS=$(cd "${{ github.action_path }}/.." && pwd)/scripts
bash "${SCRIPTS}/seed-target-dir.sh" \
"${{ steps.resolve.outputs.cache-key }}" \
"${{ steps.resolve.outputs.base-key }}" \
"${{ steps.resolve.outputs.cache-root }}" \
"${{ github.job }}-${{ github.run_id }}-$$" \
"${{ inputs.seed-fallback-dir }}" \
"${{ steps.resolve.outputs.lock-id }}"
# Re-stamps the lock the seed step already wrote (acquiring is idempotent
# — it rewrites the timestamp) and stamps the LRU marker. The marker is
# touched unconditionally every run: a run that hits the cache for every
# crate may write nothing at all inside the tree, which would make a
# just-used directory look stale to the eviction pass.
#
# `.cache-last-used` IS A CROSS-REPO CONTRACT NAME with daniel/gitdan's
# `scripts/ci-cache-reclaim.sh`, which reads this exact marker for its own
# LRU ordering and to recognise a directory as a cache dir at all. Rename
# it here without a matching change there and that script silently falls
# back to directory mtime for both — see README.md's "Scratch names in a
# cache root" section and the `DEPENDED-UPON NAMES CONTRACT` block in
# gitdan's script.
- shell: bash
run: |
set -euo pipefail
SCRIPTS=$(cd "${{ github.action_path }}/.." && pwd)/scripts
bash "${SCRIPTS}/cache-lock.sh" acquire \
"${{ steps.resolve.outputs.target-dir }}" "${{ steps.resolve.outputs.lock-id }}"
touch "${{ steps.resolve.outputs.target-dir }}/.cache-last-used"
# Runs AFTER seeding, deliberately: restore-mtimes.sh reads the build
# watermark out of the target dir, so that directory has to be in its
# final form for this run (reused, seeded, or freshly created) before the
# watermark it may carry can be read.
# CARGO_TARGET_DIR and CI_WATERMARK_FILE are passed explicitly rather
# than read from the job environment the resolve step exported: the export
# is what the consuming workflow's own build steps rely on, but a step
# inside this action should not depend on cross-step propagation working
# when it can just be handed the value.
- if: ${{ inputs.restore-mtimes == 'true' }}
shell: bash
env:
CARGO_TARGET_DIR: ${{ steps.resolve.outputs.target-dir }}
CI_WATERMARK_FILE: ${{ steps.resolve.outputs.watermark-file }}
run: bash "$(cd "${{ github.action_path }}/.." && pwd)/scripts/restore-mtimes.sh"