Files
claudeandClaude Fable 5.1 07ba53ca79 fix(prune): reclaim merged branches, and size the volume for the clone
Two assumptions in the eviction pass did not hold on this forge, and
between them a volume filled up three times in three days with nothing
reclaimed automatically. Both are replaced here; the pass also moves
ahead of the seed, which is the only order in which its work can help
the run performing it.

LIVENESS. Pass 1 evicted a cache only when its branch was gone from
origin. Gitea keeps a PR's branch after the merge unless the repo opts
into delete-on-merge, and zemyna does not — so ls-remote reports fifty
merged branches and the signal fires for none of them. A second signal
is added beside it: a branch still on origin whose tip is an ancestor of
a protected branch's tip holds no commit that branch does not, so its
cache will never be read again and goes in the same unconditional pass.

Ancestry is answered from the commits in the job's own checkout, so the
answer "cannot tell" exists and stays distinct from "not merged" at both
granularities. A shallow checkout withholds the signal entirely, since a
missing object is its normal case rather than evidence. A single branch
whose tip is not in the checkout is kept, with a warning naming it. A
squash or rebase merge leaves no ancestry and reads as live until the
branch is deleted. All three are missed reclamations, which cost disk;
the other direction costs a branch its cache mid-build.

HEADROOM. Passes 2 and 3 gated on a percentage of the volume, which
cannot express the failure they have to prevent: a clone runs out of
disk while unsharing its mutable paths, and how much that needs is a
property of the snapshot rather than of the disk. Staging failed at 34 G
free and passed at 74 G, so a 10% floor — 19 G here — never fired first.
The requirement is now measured per run off the very source the seed
will read, the pass evicts oldest-first until it is met and stops there,
and falling short of it fails with the shortfall and every directory it
kept, rather than letting the seed fail seconds later against a staging
path that names none of that. min-free-percent survives as an additional
floor, defaulting to 0, and falling short of that one is still a warning
and a self-clear.

ORDERING. The prune step ran after the seed, so each run freed space for
the next one. It now runs between resolve and seed. Two things that
makes newly reachable are closed: the source about to be cloned is
excluded from every pass by name, and a concurrent job's target dir
already carries its lock from the instant it appears under its final
name, so nothing is seen unlocked that is in use.

Red-proven: sixteen assertions across the four new scenarios fail
against the pre-fix scripts, including the zemyna layout evicting
nothing where it should evict exactly one directory, and the headroom
scenario exiting 0 where it should exit 1.

Refs #20.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JXMQCJ5Eg5f9G9cfYzyh4Z
2026-09-07 00:21:12 -05:00

280 lines
13 KiB
YAML

name: 'Cargo cache (consume)'
description: >-
Per-ref Cargo target-dir cache for Gitea Actions runners: hardlink-clones
this ref's target directory from its base branch's published, immutable
snapshot, restores git-history file mtimes, and prunes the volume.
author: 'gitdan'
inputs:
cache-root:
description: 'Mount point of the persistent cache volume inside the job container.'
required: false
default: '/cache'
cache-lineage:
description: >-
Distinguishes two jobs that build the SAME ref for different targets or
profiles (a host build and a wasm32 build, say) and would otherwise
resolve to one CARGO_TARGET_DIR and serialise on Cargo's exclusive
build-directory lock. Names one directory level under cache-root:
<cache-root>/<lineage>/target-<key>. Must be a single path component;
several names are refused outright because a reader elsewhere would stop
seeing the caches under them (see validate_cache_lineage in
scripts/cache-lib.sh). Pass the same value to cargo-cache-publish.
required: false
default: ''
protected-branches:
description: >-
Space-separated refs that publish snapshots and are never evicted.
These are the branches PR caches layer over.
required: false
default: 'dev main'
min-free-percent:
description: >-
An ADDITIONAL free-space floor, as a percentage of the cache volume.
The prune's own requirement is derived per run from what the seed is
about to clone — the mutable set it has to real-copy out of the source
snapshot, which is a property of that snapshot and not of the volume —
and this floor only ever raises it. 0, the default, leaves the derived
requirement as the only gate. Set it to keep headroom for something
other than the clone (the build's own output, another job on the same
volume); it will not make the clone fit, because it does not know how
big the clone is.
required: false
default: '0'
restore-mtimes:
description: >-
Restore every tracked file's mtime from git history. Requires a
full-history checkout (fetch-depth: 0). Set to false only if the build
does not use Cargo's mtime-based freshness at all.
required: false
default: 'true'
prune:
description: >-
Run the eviction pass (dead-branch liveness + disk pressure). It runs
BEFORE the seed step, so what it frees is available to the clone that
step makes.
required: false
default: 'true'
liveness-prune:
description: >-
Within the prune pass, remove caches for branches that are dead: gone
from origin, or still on origin with a tip already merged into a
protected branch. The second signal is what reclaims anything at all on
a forge that keeps branches after merge, and it needs the protected
branches' commits in the checkout — a shallow one withholds it and says
so. Set to false on a runner that cannot reach origin.
required: false
default: 'true'
own-ref:
description: 'Override this run''s ref. Defaults to github.head_ref, else github.ref_name.'
required: false
default: ''
base-ref:
description: 'Override the ref to layer over. Defaults to github.base_ref (empty on push).'
required: false
default: ''
seed-fallback-dir:
description: >-
Absolute path to seed from when no snapshot exists yet — a pre-existing
flat cache directory during a migration. Optional.
required: false
default: ''
watermark-file:
description: >-
Name of this job's build-watermark file inside the target dir. MUST be
distinct per job when two jobs share one target directory — which two
jobs no longer need to do; cache-lineage gives them separate ones.
Defaults to .ci-watermark-<job>-sha, already distinct per job.
required: false
default: ''
lock-id:
description: 'Identifier for this job''s cache lock. Defaults to <job>-<run_id>.'
required: false
default: ''
stale-lock-seconds:
description: 'Age past which another job''s cache lock is treated as abandoned.'
required: false
default: '7200'
outputs:
target-dir:
description: 'Resolved CARGO_TARGET_DIR. Also exported to the job environment.'
value: ${{ steps.resolve.outputs.target-dir }}
cache-key:
description: 'Sanitized cache key for this run''s own ref.'
value: ${{ steps.resolve.outputs.cache-key }}
cache-root:
description: >-
Resolved cache root — cache-root, plus the lineage directory when one is
set. Every directory this action reads or writes is under it. Also
exported as CARGO_CACHE_ROOT.
value: ${{ steps.resolve.outputs.cache-root }}
seeded-from:
description: 'Where the target dir came from: own | base-snapshot | own-snapshot | fallback-dir | concurrent-peer | cold.'
value: ${{ steps.seed.outputs.seeded-from }}
runs:
using: 'composite'
steps:
# Resolves both cache keys and exports the environment every later step
# (and the consuming job's own build steps) reads. Must run before
# anything that touches CARGO_TARGET_DIR, which is why it is first.
#
# `head_ref || ref_name` rather than `ref_name` alone: on a pull_request
# event `ref_name` is a synthetic merge-ref name that changes on every
# push to the PR, so keying on it would give the same PR a different cache
# directory every time — defeating the reuse this action exists to
# provide. On a push event `head_ref` is empty and `ref_name` is the real
# branch, which is what we want there.
#
# `base_ref` is populated only for pull_request events. A push run has
# nothing to layer over: its own ref IS the reference branch. It
# publishes, it does not consume.
#
# The cache root is resolved first because every path below hangs off it.
# A lineage nests one directory level (`<root>/<lineage>/target-<key>`),
# which is what lets two jobs on ONE ref hold two build directories and so
# not serialise on Cargo's exclusive lock. It is resolved through
# cache-root.sh rather than interpolated here so the name is validated —
# several otherwise-reasonable lineage names put their whole subtree out of
# reach of a pass that has to see it. Everything downstream reads the
# resolved value, and it is exported as CARGO_CACHE_ROOT so the publish
# action can check it agrees with its own inputs.
- id: resolve
shell: bash
run: |
set -euo pipefail
SCRIPTS=$(cd "${{ github.action_path }}/.." && pwd)/scripts
[ -d "$SCRIPTS" ] || { echo "::error::cargo-cache: scripts/ not found at $SCRIPTS"; exit 1; }
echo "CARGO_CACHE_SCRIPTS=${SCRIPTS}" >> "$GITHUB_ENV"
CACHE_ROOT=$(bash "${SCRIPTS}/cache-root.sh" resolve \
"${{ inputs.cache-root }}" "${{ inputs.cache-lineage }}")
OWN_REF="${{ inputs.own-ref }}"
[ -n "$OWN_REF" ] || OWN_REF="${{ github.head_ref || github.ref_name }}"
BASE_REF="${{ inputs.base-ref }}"
[ -n "$BASE_REF" ] || BASE_REF="${{ github.base_ref }}"
OWN_KEY=$(bash "${SCRIPTS}/branch-cache-key.sh" "$OWN_REF")
BASE_KEY=""
if [ -n "$BASE_REF" ]; then
BASE_KEY=$(bash "${SCRIPTS}/branch-cache-key.sh" "$BASE_REF")
echo "cache: own ref '${OWN_REF}' -> ${OWN_KEY}; layering over base ref '${BASE_REF}' -> ${BASE_KEY}"
else
echo "cache: own ref '${OWN_REF}' -> ${OWN_KEY} (no base ref — this ref publishes, it does not consume)"
fi
TARGET_DIR="${CACHE_ROOT}/target-${OWN_KEY}"
WATERMARK="${{ inputs.watermark-file }}"
[ -n "$WATERMARK" ] || WATERMARK=".ci-watermark-${{ github.job }}-sha"
LOCK_ID="${{ inputs.lock-id }}"
[ -n "$LOCK_ID" ] || LOCK_ID="${{ github.job }}-${{ github.run_id }}"
{
echo "target-dir=${TARGET_DIR}"
echo "cache-key=${OWN_KEY}"
echo "base-key=${BASE_KEY}"
echo "lock-id=${LOCK_ID}"
echo "watermark-file=${WATERMARK}"
echo "cache-root=${CACHE_ROOT}"
} >> "$GITHUB_OUTPUT"
{
echo "CARGO_TARGET_DIR=${TARGET_DIR}"
echo "CARGO_CACHE_ROOT=${CACHE_ROOT}"
echo "CARGO_CACHE_KEY=${OWN_KEY}"
echo "CARGO_CACHE_LOCK_ID=${LOCK_ID}"
echo "CI_WATERMARK_FILE=${WATERMARK}"
} >> "$GITHUB_ENV"
# Runs BEFORE the seed, which is the only order in which its work can
# help: the eviction it performs is what makes room for the clone the seed
# step is about to make, and the requirement it evicts against is measured
# off the snapshot that clone will read. Running afterwards — where this
# step used to be — meant every run freed space for the NEXT one and the
# seed met whatever the last run happened to leave.
#
# Two things this ordering has to be safe against, and is:
#
# The source it is about to read is excluded from every pass by name
# (see protected_reason in prune-cache.sh), so pass 1 cannot take the
# snapshot out from under the seed that follows it.
# A concurrent job's target dir carries its lock from the instant it
# appears under its final name — the seed writes it into the staging
# tree before the rename — so there is no window in which running this
# earlier sees an unlocked directory somebody is using.
- if: ${{ inputs.prune == 'true' }}
shell: bash
env:
STALE_LOCK_SECONDS: ${{ inputs.stale-lock-seconds }}
CACHE_LIVENESS: ${{ inputs.liveness-prune }}
run: |
set -euo pipefail
SCRIPTS=$(cd "${{ github.action_path }}/.." && pwd)/scripts
bash "${SCRIPTS}/prune-cache.sh" \
"${{ steps.resolve.outputs.cache-root }}" \
"${{ steps.resolve.outputs.target-dir }}" \
"${{ inputs.protected-branches }}" \
"${{ inputs.min-free-percent }}" \
"${{ steps.resolve.outputs.cache-key }}" \
"${{ steps.resolve.outputs.base-key }}" \
"${{ inputs.seed-fallback-dir }}"
# Seeds this ref's target dir from the base's published snapshot. See
# scripts/seed-target-dir.sh — the staging-then-atomic-rename is what
# makes concurrent jobs sharing one cache key safe by construction rather
# than by the runner happening to have a single execution slot, and the
# clone's own consistency check plus the publish side's reader interlock
# are what make it safe against the base republishing MID-CLONE.
#
# The lock id is passed here as well as acquired in the next step: the
# seed writes it into the staging tree, so the directory carries a lock
# the instant it appears under its final name rather than a step later.
- id: seed
shell: bash
run: |
set -euo pipefail
SCRIPTS=$(cd "${{ github.action_path }}/.." && pwd)/scripts
bash "${SCRIPTS}/seed-target-dir.sh" \
"${{ steps.resolve.outputs.cache-key }}" \
"${{ steps.resolve.outputs.base-key }}" \
"${{ steps.resolve.outputs.cache-root }}" \
"${{ github.job }}-${{ github.run_id }}-$$" \
"${{ inputs.seed-fallback-dir }}" \
"${{ steps.resolve.outputs.lock-id }}"
# Re-stamps the lock the seed step already wrote (acquiring is idempotent
# — it rewrites the timestamp) and stamps the LRU marker. The marker is
# touched unconditionally every run: a run that hits the cache for every
# crate may write nothing at all inside the tree, which would make a
# just-used directory look stale to the eviction pass.
#
# `.cache-last-used` IS A CROSS-REPO CONTRACT NAME with daniel/gitdan's
# `scripts/ci-cache-reclaim.sh`, which reads this exact marker for its own
# LRU ordering and to recognise a directory as a cache dir at all. Rename
# it here without a matching change there and that script silently falls
# back to directory mtime for both — see README.md's "Scratch names in a
# cache root" section and the `DEPENDED-UPON NAMES CONTRACT` block in
# gitdan's script.
- shell: bash
run: |
set -euo pipefail
SCRIPTS=$(cd "${{ github.action_path }}/.." && pwd)/scripts
bash "${SCRIPTS}/cache-lock.sh" acquire \
"${{ steps.resolve.outputs.target-dir }}" "${{ steps.resolve.outputs.lock-id }}"
touch "${{ steps.resolve.outputs.target-dir }}/.cache-last-used"
# Runs AFTER seeding, deliberately: restore-mtimes.sh reads the build
# watermark out of the target dir, so that directory has to be in its
# final form for this run (reused, seeded, or freshly created) before the
# watermark it may carry can be read.
# CARGO_TARGET_DIR and CI_WATERMARK_FILE are passed explicitly rather
# than read from the job environment the resolve step exported: the export
# is what the consuming workflow's own build steps rely on, but a step
# inside this action should not depend on cross-step propagation working
# when it can just be handed the value.
- if: ${{ inputs.restore-mtimes == 'true' }}
shell: bash
env:
CARGO_TARGET_DIR: ${{ steps.resolve.outputs.target-dir }}
CI_WATERMARK_FILE: ${{ steps.resolve.outputs.watermark-file }}
run: bash "$(cd "${{ github.action_path }}/.." && pwd)/scripts/restore-mtimes.sh"