feat(cargo-cache): hardlink-clone a per-ref Cargo cache from a published snapshot
Replaces the phase-0 resolution probe with the real actions, merging the two independent per-branch Cargo cache implementations on this forge into the design neither of them had. ## The merge - zemyna seeds a PR branch by `cp -al` hardlink clone (near-free: cost scales with inode count, not bytes) from the base branch's LIVE target dir — a torn read waiting for a second job slot (its own #911). - emowheel seeds from a PUBLISHED IMMUTABLE SNAPSHOT (no race by construction) but with `cp -a`, duplicating ~35 GB per branch. This ships hardlink-clone FROM a published snapshot: zemyna's cost profile, emowheel's soundness, and #911 closed structurally rather than by the runner happening to have one execution slot. ## The bug both implementations have A build inside a `cp -al` clone DOES mutate the directory it was cloned from. Cargo replaces real artifacts, but writes its metadata — and build scripts write their OUT_DIR — with a plain truncating write, straight through the shared inode. Measured set: `.fingerprint/<unit>/dep-<target>` (under CARGO_UNSTABLE_CHECKSUM_FRESHNESS), `build/<pkg>/{output,root-output,out/**}`, `deps/*.d` and `<profile>/*.d`. The checksum-freshness case is a wrong answer, not a slow build: a PR clone rewrites the base's dep-info to describe the PR's sources while the base's cache still holds the artifact built from the base's; once the PR merges, the base's next run finds the checksums match, reports `Fresh`, and links a binary built from the pre-merge code. Reproduced end to end. Fix: hardlink the artifacts (the GB), real-copy the metadata (the MB) — about 3.7% of a 6.9 GB Bevy target dir, against 100% for a full copy. ## Contents - `cargo-cache/action.yml` — consume: resolve keys, seed from the base's snapshot via staging + one atomic rename, strip Cargo lock files, unshare the mutable paths, restore mtimes from git history, lock, prune. - `cargo-cache-publish/action.yml` — publish: record the build watermark, atomically republish the snapshot on a protected branch, release the lock (`mode: release-lock` for the `if: always()` step). - `scripts/` — all logic, so it is testable standalone; the YAML is wiring. - `scripts/*selftest.sh` + `selftest.sh` — five suites, 63 assertions, every fix paired with a control that reproduces the bug. All green locally. Eviction merges emowheel's liveness pass (dead branches pruned unconditionally, not gated on disk pressure) with LRU-under-pressure, but inverts the order within the pressure pass: `target-*` before `snapshot-*`, because a snapshot is hardlinked to everything cloned from it, so evicting one frees almost no real bytes while costing every future PR its warm start. restore-mtimes.sh is ported from emowheel (the watermark variant, which closes the merge hazard zemyna's copy still has) with its provenance de-projectised. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Sqh2vscfzisk83VuPVQX9L
This commit is contained in:
Executable
+211
@@ -0,0 +1,211 @@
|
||||
#!/usr/bin/env bash
|
||||
# Eviction for the per-ref cache directories on the persistent volume.
|
||||
#
|
||||
# Usage: prune-cache.sh <cache-root> <own-target-dir> <protected-branches> <min-free-percent>
|
||||
# protected-branches space-separated raw refs (e.g. "dev main")
|
||||
#
|
||||
# Optional environment:
|
||||
# STALE_LOCK_SECONDS age past which a .ci-lock-* marker is treated as
|
||||
# abandoned (default 7200)
|
||||
# CACHE_LIVENESS "false"/"0" to skip the liveness pass entirely
|
||||
# CACHE_DF_OVERRIDE "<total_kb> <free_kb>", for the selftest
|
||||
#
|
||||
# Three passes, in order:
|
||||
#
|
||||
# 1. LIVENESS — every target-*/snapshot-* directory whose branch no longer
|
||||
# exists on origin is removed UNCONDITIONALLY, not gated on free space.
|
||||
# A directory for a branch deleted days ago is pure loss: nothing will
|
||||
# ever read it again, since a merged PR's branch cannot be reopened.
|
||||
# Waiting for disk pressure to notice means paying for it until then.
|
||||
# Skipped entirely, loudly, if the liveness signal itself is
|
||||
# unavailable — "couldn't determine" is never folded into "dead".
|
||||
# 2. PRESSURE — if free space is still under the threshold, evict remaining
|
||||
# (now necessarily live) directories oldest-first until it clears.
|
||||
# 3. SELF-CLEAR — if pass 2 still isn't enough, wipe this run's own target
|
||||
# dir and pay a cold rebuild, reported to the job summary as well as the
|
||||
# log, because a warning on a green run is what lets a silently 4x-slower
|
||||
# job go unnoticed.
|
||||
#
|
||||
# Reactive-only, with no hard cap on cache size: a workspace's natural working
|
||||
# set is what it is, and bounding the footprint preemptively means wiping
|
||||
# useful content before it is actually causing host pressure. The bound is
|
||||
# physics, not an arbitrary GB number.
|
||||
#
|
||||
# EVICTION ORDER, and why it is the reverse of the obvious one: within the
|
||||
# pressure pass, `target-*` directories are evicted BEFORE `snapshot-*` ones.
|
||||
# A snapshot is a hardlink clone of a live target dir and of every consumer
|
||||
# cloned from it, so removing it frees almost no real bytes — its inodes stay
|
||||
# alive through those other links — while costing every future PR its warm
|
||||
# start. Evicting snapshots first would be nearly pure loss. Target
|
||||
# directories are where a branch's own divergent artifacts actually live, so
|
||||
# they are what freeing space means.
|
||||
#
|
||||
# Two exclusions every pass respects:
|
||||
#
|
||||
# PROTECTED — the publisher branches' target and snapshot directories, and
|
||||
# this run's own target dir, are never candidates in any pass. Evicting a
|
||||
# publisher's snapshot doesn't free real disk (every open PR's clone keeps
|
||||
# the data alive) but does force every subsequent PR to start cold, which is
|
||||
# the entire benefit this scheme exists to deliver.
|
||||
#
|
||||
# LOCKED — a directory carrying a .ci-lock-* marker younger than
|
||||
# STALE_LOCK_SECONDS is held open by a running job and is skipped by every
|
||||
# pass, however dead and however tight the disk. This is what makes eviction
|
||||
# safe on a runner with more than one execution slot. An older marker is
|
||||
# treated as abandoned and logged as such, so an actually-still-running job
|
||||
# that somehow exceeds the threshold is visible in the log rather than
|
||||
# silently losing its cache mid-build.
|
||||
#
|
||||
# Liveness is resolved by `git ls-remote --heads origin`, wrapped in a
|
||||
# timeout. A directory name cannot be inverted back to a branch name (the
|
||||
# sanitiser is lossy and the disambiguating suffix is a one-way hash), so this
|
||||
# goes the other direction: it recomputes the expected directory names for
|
||||
# every branch origin reports, using cache-lib.sh's OWN cache_key function —
|
||||
# the same one the seed step used to create them. Reusing that function rather
|
||||
# than reimplementing a lookalike is what makes the classification sound; any
|
||||
# drift between two spellings would silently misclassify every directory.
|
||||
#
|
||||
# Only ever globs inside <cache-root>. Another project's volume is a different
|
||||
# Docker named volume and is not mounted in this container at all, so "stays
|
||||
# scoped to this repo's cache" holds structurally, not by convention.
|
||||
set -euo pipefail
|
||||
. "$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/cache-lib.sh"
|
||||
|
||||
ROOT="${1:?usage: prune-cache.sh <cache-root> <own-target-dir> <protected-branches> <min-free-percent>}"
|
||||
OWN_DIR="${2:?}"
|
||||
PROTECTED_REFS="${3:-}"
|
||||
MIN_FREE_PCT="${4:-10}"
|
||||
STALE_LOCK_SECONDS="${STALE_LOCK_SECONDS:-7200}"
|
||||
|
||||
declare -A protected_ns=()
|
||||
for ref in $PROTECTED_REFS; do
|
||||
suffix=$(cache_key "$ref")
|
||||
protected_ns["target-${suffix}"]=1
|
||||
protected_ns["snapshot-${suffix}"]=1
|
||||
done
|
||||
|
||||
is_protected() {
|
||||
local dir="$1" name
|
||||
name=$(basename "$dir")
|
||||
[ "$dir" = "$OWN_DIR" ] && return 0
|
||||
[ -n "${protected_ns[$name]:-}" ] && return 0
|
||||
return 1
|
||||
}
|
||||
|
||||
is_locked() {
|
||||
local dir="$1" now lock_file lock_age locked=1
|
||||
now=$(date +%s)
|
||||
for lock_file in "$dir"/.ci-lock-*; do
|
||||
[ -e "$lock_file" ] || continue
|
||||
lock_age=$(( now - $(stat -c '%Y' "$lock_file") ))
|
||||
if [ "$lock_age" -lt "$STALE_LOCK_SECONDS" ]; then
|
||||
echo " $(basename "$dir"): held open by $(basename "$lock_file") (${lock_age}s old)"
|
||||
locked=0
|
||||
else
|
||||
echo " $(basename "$dir"): ignoring stale lock $(basename "$lock_file") (${lock_age}s old > ${STALE_LOCK_SECONDS}s) — treating as abandoned"
|
||||
fi
|
||||
done
|
||||
return "$locked"
|
||||
}
|
||||
|
||||
# Directories oldest-first, target-* before snapshot-* (see the header).
|
||||
# The sort key is a rank digit followed by a zero-padded mtime, so the two
|
||||
# groups sort as blocks rather than interleaving by age. `.cache-last-used`
|
||||
# is the marker every run touches; a run that hits the cache for every crate
|
||||
# may write nothing at all inside the tree, which would make a
|
||||
# just-used directory look stale without it. Fall back to the directory's
|
||||
# own mtime when the marker is missing (a partially-written directory from an
|
||||
# interrupted run is still a valid, if less precise, "last touched" signal).
|
||||
list_by_lru() {
|
||||
local rank d ts
|
||||
for rank in 0:target 1:snapshot; do
|
||||
for d in "$ROOT"/"${rank#*:}"-*; do
|
||||
[ -d "$d" ] || continue
|
||||
if [ -e "$d/.cache-last-used" ]; then ts=$(stat -c '%Y' "$d/.cache-last-used" 2>/dev/null || echo 0)
|
||||
else ts=$(stat -c '%Y' "$d" 2>/dev/null || echo 0); fi
|
||||
printf '%s%012d\t%s\n' "${rank%%:*}" "$ts" "$d"
|
||||
done
|
||||
done | sort | cut -f2-
|
||||
}
|
||||
|
||||
echo "=== pass 1: liveness (unconditional, not gated on free space) ==="
|
||||
declare -A live_ns=()
|
||||
LIVENESS_AVAILABLE=0
|
||||
LIVENESS_REASON=""
|
||||
if [ "${CACHE_LIVENESS:-true}" = "0" ] || [ "${CACHE_LIVENESS:-true}" = "false" ]; then
|
||||
LIVENESS_REASON="disabled via CACHE_LIVENESS=${CACHE_LIVENESS}"
|
||||
elif remote_heads=$(timeout 20 git ls-remote --heads origin 2>&1); then
|
||||
LIVENESS_AVAILABLE=1
|
||||
branch_count=0
|
||||
while IFS= read -r line; do
|
||||
[ -n "$line" ] || continue
|
||||
case "$line" in *refs/heads/*) ;; *) continue ;; esac
|
||||
branch="${line#*refs/heads/}"
|
||||
[ -n "$branch" ] || continue
|
||||
suffix=$(cache_key "$branch")
|
||||
live_ns["target-${suffix}"]=1
|
||||
live_ns["snapshot-${suffix}"]=1
|
||||
branch_count=$((branch_count + 1))
|
||||
done <<< "$remote_heads"
|
||||
echo "liveness: ${branch_count} live branches on origin"
|
||||
else
|
||||
LIVENESS_REASON="git ls-remote --heads origin failed or timed out"
|
||||
fi
|
||||
|
||||
if [ "$LIVENESS_AVAILABLE" = "1" ]; then
|
||||
pruned_any=0
|
||||
for dir in "$ROOT"/target-* "$ROOT"/snapshot-*; do
|
||||
[ -d "$dir" ] || continue
|
||||
name=$(basename "$dir")
|
||||
is_protected "$dir" && continue
|
||||
[ -n "${live_ns[$name]:-}" ] && continue
|
||||
is_locked "$dir" && continue
|
||||
dir_gb=$(usage_gb "$dir")
|
||||
echo "::warning::pruning dead-branch cache ${name} (${dir_gb} GB) — no matching branch on origin"
|
||||
summary_line "- pruned dead-branch cache \`${name}\` (${dir_gb} GB) — branch no longer exists on origin"
|
||||
rm -rf "$dir"
|
||||
pruned_any=1
|
||||
done
|
||||
[ "$pruned_any" = "1" ] || echo "no dead-branch caches found"
|
||||
else
|
||||
echo "liveness: ${LIVENESS_REASON} — treating as UNAVAILABLE (not as \"no branches\"); pass 1 skipped"
|
||||
fi
|
||||
|
||||
echo
|
||||
echo "=== pass 2/3: disk pressure (threshold: free < ${MIN_FREE_PCT}%) ==="
|
||||
read -r TOTAL_KB FREE_KB <<< "$(read_df "$ROOT")"
|
||||
THRESHOLD_KB=$(( TOTAL_KB * MIN_FREE_PCT / 100 ))
|
||||
|
||||
if [ "$FREE_KB" -ge "$THRESHOLD_KB" ]; then
|
||||
echo "cache: $(basename "$OWN_DIR") $(usage_gb "$OWN_DIR") GB | $(report_df host "$FREE_KB" "$TOTAL_KB")"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
echo "::warning::$(report_df disk "$FREE_KB" "$TOTAL_KB") < ${MIN_FREE_PCT}% threshold"
|
||||
|
||||
# A plain loop over a pre-materialised, pre-sorted list rather than a live
|
||||
# `find | while` pipeline, so `rm -rf` inside the loop cannot make a running
|
||||
# directory walk re-observe its own deletions.
|
||||
mapfile -t LRU < <(list_by_lru)
|
||||
for dir in "${LRU[@]}"; do
|
||||
[ -d "$dir" ] || continue
|
||||
is_protected "$dir" && continue
|
||||
read -r TOTAL_KB FREE_KB <<< "$(read_df "$ROOT")"
|
||||
[ "$FREE_KB" -ge "$THRESHOLD_KB" ] && break
|
||||
is_locked "$dir" && continue
|
||||
dir_gb=$(usage_gb "$dir")
|
||||
echo "::warning::evicting $(basename "$dir") (${dir_gb} GB, LRU under disk pressure)"
|
||||
summary_line "- evicted \`$(basename "$dir")\` (${dir_gb} GB, LRU under disk pressure)"
|
||||
rm -rf "$dir"
|
||||
done
|
||||
|
||||
read -r TOTAL_KB FREE_KB <<< "$(read_df "$ROOT")"
|
||||
if [ "$FREE_KB" -lt "$THRESHOLD_KB" ]; then
|
||||
OWN_GB=$(usage_gb "$OWN_DIR")
|
||||
echo "::warning::still under threshold after evicting every eligible sibling; clearing own $(basename "$OWN_DIR") (was ${OWN_GB} GB) — this run pays a cold rebuild"
|
||||
summary_line "- **self-clear**: \`$(basename "$OWN_DIR")\` (was ${OWN_GB} GB) wiped — this run pays a cold rebuild"
|
||||
rm -rf "$OWN_DIR"
|
||||
mkdir -p "$OWN_DIR"
|
||||
else
|
||||
echo "$(report_df post-eviction "$FREE_KB" "$TOTAL_KB") — sibling eviction recovered enough space"
|
||||
fi
|
||||
Reference in New Issue
Block a user