Replaces the phase-0 resolution probe with the real actions, merging the two independent per-branch Cargo cache implementations on this forge into the design neither of them had. ## The merge - zemyna seeds a PR branch by `cp -al` hardlink clone (near-free: cost scales with inode count, not bytes) from the base branch's LIVE target dir — a torn read waiting for a second job slot (its own #911). - emowheel seeds from a PUBLISHED IMMUTABLE SNAPSHOT (no race by construction) but with `cp -a`, duplicating ~35 GB per branch. This ships hardlink-clone FROM a published snapshot: zemyna's cost profile, emowheel's soundness, and #911 closed structurally rather than by the runner happening to have one execution slot. ## The bug both implementations have A build inside a `cp -al` clone DOES mutate the directory it was cloned from. Cargo replaces real artifacts, but writes its metadata — and build scripts write their OUT_DIR — with a plain truncating write, straight through the shared inode. Measured set: `.fingerprint/<unit>/dep-<target>` (under CARGO_UNSTABLE_CHECKSUM_FRESHNESS), `build/<pkg>/{output,root-output,out/**}`, `deps/*.d` and `<profile>/*.d`. The checksum-freshness case is a wrong answer, not a slow build: a PR clone rewrites the base's dep-info to describe the PR's sources while the base's cache still holds the artifact built from the base's; once the PR merges, the base's next run finds the checksums match, reports `Fresh`, and links a binary built from the pre-merge code. Reproduced end to end. Fix: hardlink the artifacts (the GB), real-copy the metadata (the MB) — about 3.7% of a 6.9 GB Bevy target dir, against 100% for a full copy. ## Contents - `cargo-cache/action.yml` — consume: resolve keys, seed from the base's snapshot via staging + one atomic rename, strip Cargo lock files, unshare the mutable paths, restore mtimes from git history, lock, prune. - `cargo-cache-publish/action.yml` — publish: record the build watermark, atomically republish the snapshot on a protected branch, release the lock (`mode: release-lock` for the `if: always()` step). - `scripts/` — all logic, so it is testable standalone; the YAML is wiring. - `scripts/*selftest.sh` + `selftest.sh` — five suites, 63 assertions, every fix paired with a control that reproduces the bug. All green locally. Eviction merges emowheel's liveness pass (dead branches pruned unconditionally, not gated on disk pressure) with LRU-under-pressure, but inverts the order within the pressure pass: `target-*` before `snapshot-*`, because a snapshot is hardlinked to everything cloned from it, so evicting one frees almost no real bytes while costing every future PR its warm start. restore-mtimes.sh is ported from emowheel (the watermark variant, which closes the merge hazard zemyna's copy still has) with its provenance de-projectised. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Sqh2vscfzisk83VuPVQX9L
212 lines
9.4 KiB
Bash
Executable File
212 lines
9.4 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# Eviction for the per-ref cache directories on the persistent volume.
|
|
#
|
|
# Usage: prune-cache.sh <cache-root> <own-target-dir> <protected-branches> <min-free-percent>
|
|
# protected-branches space-separated raw refs (e.g. "dev main")
|
|
#
|
|
# Optional environment:
|
|
# STALE_LOCK_SECONDS age past which a .ci-lock-* marker is treated as
|
|
# abandoned (default 7200)
|
|
# CACHE_LIVENESS "false"/"0" to skip the liveness pass entirely
|
|
# CACHE_DF_OVERRIDE "<total_kb> <free_kb>", for the selftest
|
|
#
|
|
# Three passes, in order:
|
|
#
|
|
# 1. LIVENESS — every target-*/snapshot-* directory whose branch no longer
|
|
# exists on origin is removed UNCONDITIONALLY, not gated on free space.
|
|
# A directory for a branch deleted days ago is pure loss: nothing will
|
|
# ever read it again, since a merged PR's branch cannot be reopened.
|
|
# Waiting for disk pressure to notice means paying for it until then.
|
|
# Skipped entirely, loudly, if the liveness signal itself is
|
|
# unavailable — "couldn't determine" is never folded into "dead".
|
|
# 2. PRESSURE — if free space is still under the threshold, evict remaining
|
|
# (now necessarily live) directories oldest-first until it clears.
|
|
# 3. SELF-CLEAR — if pass 2 still isn't enough, wipe this run's own target
|
|
# dir and pay a cold rebuild, reported to the job summary as well as the
|
|
# log, because a warning on a green run is what lets a silently 4x-slower
|
|
# job go unnoticed.
|
|
#
|
|
# Reactive-only, with no hard cap on cache size: a workspace's natural working
|
|
# set is what it is, and bounding the footprint preemptively means wiping
|
|
# useful content before it is actually causing host pressure. The bound is
|
|
# physics, not an arbitrary GB number.
|
|
#
|
|
# EVICTION ORDER, and why it is the reverse of the obvious one: within the
|
|
# pressure pass, `target-*` directories are evicted BEFORE `snapshot-*` ones.
|
|
# A snapshot is a hardlink clone of a live target dir and of every consumer
|
|
# cloned from it, so removing it frees almost no real bytes — its inodes stay
|
|
# alive through those other links — while costing every future PR its warm
|
|
# start. Evicting snapshots first would be nearly pure loss. Target
|
|
# directories are where a branch's own divergent artifacts actually live, so
|
|
# they are what freeing space means.
|
|
#
|
|
# Two exclusions every pass respects:
|
|
#
|
|
# PROTECTED — the publisher branches' target and snapshot directories, and
|
|
# this run's own target dir, are never candidates in any pass. Evicting a
|
|
# publisher's snapshot doesn't free real disk (every open PR's clone keeps
|
|
# the data alive) but does force every subsequent PR to start cold, which is
|
|
# the entire benefit this scheme exists to deliver.
|
|
#
|
|
# LOCKED — a directory carrying a .ci-lock-* marker younger than
|
|
# STALE_LOCK_SECONDS is held open by a running job and is skipped by every
|
|
# pass, however dead and however tight the disk. This is what makes eviction
|
|
# safe on a runner with more than one execution slot. An older marker is
|
|
# treated as abandoned and logged as such, so an actually-still-running job
|
|
# that somehow exceeds the threshold is visible in the log rather than
|
|
# silently losing its cache mid-build.
|
|
#
|
|
# Liveness is resolved by `git ls-remote --heads origin`, wrapped in a
|
|
# timeout. A directory name cannot be inverted back to a branch name (the
|
|
# sanitiser is lossy and the disambiguating suffix is a one-way hash), so this
|
|
# goes the other direction: it recomputes the expected directory names for
|
|
# every branch origin reports, using cache-lib.sh's OWN cache_key function —
|
|
# the same one the seed step used to create them. Reusing that function rather
|
|
# than reimplementing a lookalike is what makes the classification sound; any
|
|
# drift between two spellings would silently misclassify every directory.
|
|
#
|
|
# Only ever globs inside <cache-root>. Another project's volume is a different
|
|
# Docker named volume and is not mounted in this container at all, so "stays
|
|
# scoped to this repo's cache" holds structurally, not by convention.
|
|
set -euo pipefail
|
|
. "$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/cache-lib.sh"
|
|
|
|
ROOT="${1:?usage: prune-cache.sh <cache-root> <own-target-dir> <protected-branches> <min-free-percent>}"
|
|
OWN_DIR="${2:?}"
|
|
PROTECTED_REFS="${3:-}"
|
|
MIN_FREE_PCT="${4:-10}"
|
|
STALE_LOCK_SECONDS="${STALE_LOCK_SECONDS:-7200}"
|
|
|
|
declare -A protected_ns=()
|
|
for ref in $PROTECTED_REFS; do
|
|
suffix=$(cache_key "$ref")
|
|
protected_ns["target-${suffix}"]=1
|
|
protected_ns["snapshot-${suffix}"]=1
|
|
done
|
|
|
|
is_protected() {
|
|
local dir="$1" name
|
|
name=$(basename "$dir")
|
|
[ "$dir" = "$OWN_DIR" ] && return 0
|
|
[ -n "${protected_ns[$name]:-}" ] && return 0
|
|
return 1
|
|
}
|
|
|
|
is_locked() {
|
|
local dir="$1" now lock_file lock_age locked=1
|
|
now=$(date +%s)
|
|
for lock_file in "$dir"/.ci-lock-*; do
|
|
[ -e "$lock_file" ] || continue
|
|
lock_age=$(( now - $(stat -c '%Y' "$lock_file") ))
|
|
if [ "$lock_age" -lt "$STALE_LOCK_SECONDS" ]; then
|
|
echo " $(basename "$dir"): held open by $(basename "$lock_file") (${lock_age}s old)"
|
|
locked=0
|
|
else
|
|
echo " $(basename "$dir"): ignoring stale lock $(basename "$lock_file") (${lock_age}s old > ${STALE_LOCK_SECONDS}s) — treating as abandoned"
|
|
fi
|
|
done
|
|
return "$locked"
|
|
}
|
|
|
|
# Directories oldest-first, target-* before snapshot-* (see the header).
|
|
# The sort key is a rank digit followed by a zero-padded mtime, so the two
|
|
# groups sort as blocks rather than interleaving by age. `.cache-last-used`
|
|
# is the marker every run touches; a run that hits the cache for every crate
|
|
# may write nothing at all inside the tree, which would make a
|
|
# just-used directory look stale without it. Fall back to the directory's
|
|
# own mtime when the marker is missing (a partially-written directory from an
|
|
# interrupted run is still a valid, if less precise, "last touched" signal).
|
|
list_by_lru() {
|
|
local rank d ts
|
|
for rank in 0:target 1:snapshot; do
|
|
for d in "$ROOT"/"${rank#*:}"-*; do
|
|
[ -d "$d" ] || continue
|
|
if [ -e "$d/.cache-last-used" ]; then ts=$(stat -c '%Y' "$d/.cache-last-used" 2>/dev/null || echo 0)
|
|
else ts=$(stat -c '%Y' "$d" 2>/dev/null || echo 0); fi
|
|
printf '%s%012d\t%s\n' "${rank%%:*}" "$ts" "$d"
|
|
done
|
|
done | sort | cut -f2-
|
|
}
|
|
|
|
echo "=== pass 1: liveness (unconditional, not gated on free space) ==="
|
|
declare -A live_ns=()
|
|
LIVENESS_AVAILABLE=0
|
|
LIVENESS_REASON=""
|
|
if [ "${CACHE_LIVENESS:-true}" = "0" ] || [ "${CACHE_LIVENESS:-true}" = "false" ]; then
|
|
LIVENESS_REASON="disabled via CACHE_LIVENESS=${CACHE_LIVENESS}"
|
|
elif remote_heads=$(timeout 20 git ls-remote --heads origin 2>&1); then
|
|
LIVENESS_AVAILABLE=1
|
|
branch_count=0
|
|
while IFS= read -r line; do
|
|
[ -n "$line" ] || continue
|
|
case "$line" in *refs/heads/*) ;; *) continue ;; esac
|
|
branch="${line#*refs/heads/}"
|
|
[ -n "$branch" ] || continue
|
|
suffix=$(cache_key "$branch")
|
|
live_ns["target-${suffix}"]=1
|
|
live_ns["snapshot-${suffix}"]=1
|
|
branch_count=$((branch_count + 1))
|
|
done <<< "$remote_heads"
|
|
echo "liveness: ${branch_count} live branches on origin"
|
|
else
|
|
LIVENESS_REASON="git ls-remote --heads origin failed or timed out"
|
|
fi
|
|
|
|
if [ "$LIVENESS_AVAILABLE" = "1" ]; then
|
|
pruned_any=0
|
|
for dir in "$ROOT"/target-* "$ROOT"/snapshot-*; do
|
|
[ -d "$dir" ] || continue
|
|
name=$(basename "$dir")
|
|
is_protected "$dir" && continue
|
|
[ -n "${live_ns[$name]:-}" ] && continue
|
|
is_locked "$dir" && continue
|
|
dir_gb=$(usage_gb "$dir")
|
|
echo "::warning::pruning dead-branch cache ${name} (${dir_gb} GB) — no matching branch on origin"
|
|
summary_line "- pruned dead-branch cache \`${name}\` (${dir_gb} GB) — branch no longer exists on origin"
|
|
rm -rf "$dir"
|
|
pruned_any=1
|
|
done
|
|
[ "$pruned_any" = "1" ] || echo "no dead-branch caches found"
|
|
else
|
|
echo "liveness: ${LIVENESS_REASON} — treating as UNAVAILABLE (not as \"no branches\"); pass 1 skipped"
|
|
fi
|
|
|
|
echo
|
|
echo "=== pass 2/3: disk pressure (threshold: free < ${MIN_FREE_PCT}%) ==="
|
|
read -r TOTAL_KB FREE_KB <<< "$(read_df "$ROOT")"
|
|
THRESHOLD_KB=$(( TOTAL_KB * MIN_FREE_PCT / 100 ))
|
|
|
|
if [ "$FREE_KB" -ge "$THRESHOLD_KB" ]; then
|
|
echo "cache: $(basename "$OWN_DIR") $(usage_gb "$OWN_DIR") GB | $(report_df host "$FREE_KB" "$TOTAL_KB")"
|
|
exit 0
|
|
fi
|
|
|
|
echo "::warning::$(report_df disk "$FREE_KB" "$TOTAL_KB") < ${MIN_FREE_PCT}% threshold"
|
|
|
|
# A plain loop over a pre-materialised, pre-sorted list rather than a live
|
|
# `find | while` pipeline, so `rm -rf` inside the loop cannot make a running
|
|
# directory walk re-observe its own deletions.
|
|
mapfile -t LRU < <(list_by_lru)
|
|
for dir in "${LRU[@]}"; do
|
|
[ -d "$dir" ] || continue
|
|
is_protected "$dir" && continue
|
|
read -r TOTAL_KB FREE_KB <<< "$(read_df "$ROOT")"
|
|
[ "$FREE_KB" -ge "$THRESHOLD_KB" ] && break
|
|
is_locked "$dir" && continue
|
|
dir_gb=$(usage_gb "$dir")
|
|
echo "::warning::evicting $(basename "$dir") (${dir_gb} GB, LRU under disk pressure)"
|
|
summary_line "- evicted \`$(basename "$dir")\` (${dir_gb} GB, LRU under disk pressure)"
|
|
rm -rf "$dir"
|
|
done
|
|
|
|
read -r TOTAL_KB FREE_KB <<< "$(read_df "$ROOT")"
|
|
if [ "$FREE_KB" -lt "$THRESHOLD_KB" ]; then
|
|
OWN_GB=$(usage_gb "$OWN_DIR")
|
|
echo "::warning::still under threshold after evicting every eligible sibling; clearing own $(basename "$OWN_DIR") (was ${OWN_GB} GB) — this run pays a cold rebuild"
|
|
summary_line "- **self-clear**: \`$(basename "$OWN_DIR")\` (was ${OWN_GB} GB) wiped — this run pays a cold rebuild"
|
|
rm -rf "$OWN_DIR"
|
|
mkdir -p "$OWN_DIR"
|
|
else
|
|
echo "$(report_df post-eviction "$FREE_KB" "$TOTAL_KB") — sibling eviction recovered enough space"
|
|
fi
|