Files
gitdan-actions/scripts/prune-cache.sh
T
claude b0d63c807b style(prune-cache): rewrap the settle-window comment; drop an overclaim
Review nits on #2, both cosmetic.

The sentence added last round left its paragraph at 127 characters in a file
that otherwise wraps comments at 78-79 — the rewrap after the insert simply
did not happen. `awk 'length($0)>80'` over both files now reports one comment
line, the pre-existing 93-character usage string at :4.

Scenario 14's header called its fixture "the only shape that can tell the two
timestamps apart". The property it needs is old mtime with fresh ctime; an
old directory renamed a moment ago is the production instance of that, not
the only construction of it. "Production's shape" was already carrying the
argument.

No behaviour change and no assertion change.
2026-08-24 09:10:43 -05:00

359 lines
17 KiB
Bash
Executable File

#!/usr/bin/env bash
# Eviction for the per-ref cache directories on the persistent volume.
#
# Usage: prune-cache.sh <cache-root> <own-target-dir> <protected-branches> <min-free-percent>
# protected-branches space-separated raw refs (e.g. "dev main")
#
# Optional environment:
# STALE_LOCK_SECONDS age past which a .ci-lock-* marker is treated as
# abandoned (default 7200)
# CACHE_LIVENESS "false"/"0" to skip the liveness pass entirely
# CACHE_DF_OVERRIDE "<total_kb> <free_kb>", for the selftest
# EVICTION_ASIDE_SETTLE_SECONDS
# how long a directory renamed aside for eviction is
# left alone before another pass may reclaim it
# (default 60)
#
# Three passes, in order:
#
# 1. LIVENESS — every target-*/snapshot-* directory whose branch no longer
# exists on origin is removed UNCONDITIONALLY, not gated on free space.
# A directory for a branch deleted days ago is pure loss: nothing will
# ever read it again, since a merged PR's branch cannot be reopened.
# Waiting for disk pressure to notice means paying for it until then.
# Skipped entirely, loudly, if the liveness signal itself is
# unavailable — "couldn't determine" is never folded into "dead".
# 2. PRESSURE — if free space is still under the threshold, evict remaining
# (now necessarily live) directories oldest-first until it clears.
# 3. SELF-CLEAR — if pass 2 still isn't enough, wipe this run's own target
# dir and pay a cold rebuild, reported to the job summary as well as the
# log, because a warning on a green run is what lets a silently 4x-slower
# job go unnoticed.
#
# Reactive-only, with no hard cap on cache size: a workspace's natural working
# set is what it is, and bounding the footprint preemptively means wiping
# useful content before it is actually causing host pressure. The bound is
# physics, not an arbitrary GB number.
#
# EVICTION ORDER, and why it is the reverse of the obvious one: within the
# pressure pass, `target-*` directories are evicted BEFORE `snapshot-*` ones.
# A snapshot is a hardlink clone of a live target dir and of every consumer
# cloned from it, so removing it frees almost no real bytes — its inodes stay
# alive through those other links — while costing every future PR its warm
# start. Evicting snapshots first would be nearly pure loss. Target
# directories are where a branch's own divergent artifacts actually live, so
# they are what freeing space means.
#
# Two exclusions every pass respects:
#
# PROTECTED — the publisher branches' target and snapshot directories, and
# this run's own target dir, are never candidates in any pass. Evicting a
# publisher's snapshot doesn't free real disk (every open PR's clone keeps
# the data alive) but does force every subsequent PR to start cold, which is
# the entire benefit this scheme exists to deliver.
#
# LOCKED — a directory carrying a .ci-lock-* marker younger than
# STALE_LOCK_SECONDS is held open by a running job, or named by a live
# .reading-<dir>-* marker (a job is hardlink-cloning it this instant), is
# skipped by every pass, however dead and however tight the disk. This is
# what makes eviction safe on a runner with more than one execution slot.
# An older marker is treated as abandoned and logged as such, so an
# actually-still-running job that somehow exceeds the threshold is visible
# in the log rather than silently losing its cache mid-build.
#
# That exclusion is decided TWICE per eviction — once as the cheap filter
# that keeps a held directory out of the pass at all, and once after the
# directory has been renamed aside, which is the decision the unlink
# actually rests on. See evict_dir.
#
# Liveness is resolved by `git ls-remote --heads origin`, wrapped in a
# timeout. A directory name cannot be inverted back to a branch name (the
# sanitiser is lossy and the disambiguating suffix is a one-way hash), so this
# goes the other direction: it recomputes the expected directory names for
# every branch origin reports, using cache-lib.sh's OWN cache_key function —
# the same one the seed step used to create them. Reusing that function rather
# than reimplementing a lookalike is what makes the classification sound; any
# drift between two spellings would silently misclassify every directory.
#
# Only ever globs inside <cache-root>. Another project's volume is a different
# Docker named volume and is not mounted in this container at all, so "stays
# scoped to this repo's cache" holds structurally, not by convention.
set -euo pipefail
. "$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/cache-lib.sh"
ROOT="${1:?usage: prune-cache.sh <cache-root> <own-target-dir> <protected-branches> <min-free-percent>}"
OWN_DIR="${2:?}"
PROTECTED_REFS="${3:-}"
MIN_FREE_PCT="${4:-10}"
STALE_LOCK_SECONDS="${STALE_LOCK_SECONDS:-7200}"
# An aside directory is in flight for one rename plus one marker glob —
# milliseconds. Anything older belongs to a pass that died between the two, so
# an age is what separates "another pass is mid-eviction" from "a leftover",
# and it separates them without having to identify the pass that created it.
# The `$$` in an aside's name was that pass's PID inside its own job
# container, so testing it with `kill -0` from a different one is not
# unreliable, it is meaningless — and PIDs recycle besides. Three orders of
# magnitude of headroom over the operation it covers, and short enough that a
# genuine leftover is reclaimed by the next run rather than lingering while
# the volume is under pressure.
EVICTION_ASIDE_SETTLE_SECONDS="${EVICTION_ASIDE_SETTLE_SECONDS:-60}"
declare -A protected_ns=()
for ref in $PROTECTED_REFS; do
suffix=$(cache_key "$ref")
protected_ns["target-${suffix}"]=1
protected_ns["snapshot-${suffix}"]=1
done
is_protected() {
local dir="$1" name
name=$(basename "$dir")
[ "$dir" = "$OWN_DIR" ] && return 0
[ -n "${protected_ns[$name]:-}" ] && return 0
return 1
}
# is_locked <dir> [name]
#
# `name` is the directory's own name for reporting and for the reader-marker
# lookup, which matters when <dir> has been renamed aside for eviction: the
# markers a consumer publishes are keyed on the name it resolved, not on
# whatever the eviction pass has since called the directory.
is_locked() {
local dir="$1" name="${2:-$(basename "$1")}" now lock_file lock_age locked=1 readers
now=$(date +%s)
# A directory being hardlink-cloned right now carries no .ci-lock-* of its
# own — a snapshot has its locks stripped by construction — so the reader
# markers are the only signal that unlinking it would truncate somebody's
# in-flight clone. Today no reachable configuration prunes a snapshot (only
# protected refs publish them, and protected refs are excluded from every
# pass), which makes this guard redundant *by policy* — evict_dir's second
# look is what stops policy being the reason it is safe.
readers=$(live_reader_count "$ROOT" "$name")
if [ "$readers" -gt 0 ]; then
echo " ${name}: ${readers} job(s) currently cloning it — not a candidate"
locked=0
fi
for lock_file in "$dir"/.ci-lock-*; do
[ -e "$lock_file" ] || continue
lock_age=$(( now - $(stat -c '%Y' "$lock_file") ))
if [ "$lock_age" -lt "$STALE_LOCK_SECONDS" ]; then
echo " ${name}: held open by $(basename "$lock_file") (${lock_age}s old)"
locked=0
else
echo " ${name}: ignoring stale lock $(basename "$lock_file") (${lock_age}s old > ${STALE_LOCK_SECONDS}s) — treating as abandoned"
fi
done
return "$locked"
}
# evict_dir <dir>
#
# Unlinks <dir>, or declines to and says why. Returns 0 only if the directory
# is actually gone.
#
# The caller has already established that <dir> is a candidate, which is not
# the same as establishing that it is still one at the instant of the unlink:
# a consumer publishes its reader marker whenever it starts a clone, and the
# `du` that measures the directory between those two points runs for seconds
# on a multi-GB tree. Unlinking under a live clone truncates it silently —
# `cp -al` never reports a subtree that was removed before it read the
# parent's listing (see cache-lib.sh's reader-marker section).
#
# So the directory is renamed aside first and only then re-examined, which is
# what makes the second look conclusive rather than merely closer to the
# unlink. It is publish-snapshot.sh's rotation, and it rests on the same
# ordering proof: a consumer publishes its marker BEFORE it resolves the
# source path, so one that resolved this directory did so before the rename
# and therefore published its marker before the scan below, which happens
# strictly after that rename. A consumer arriving after the rename cannot
# resolve the path at all and falls through to its own cold-start path — the
# same safe degrade the publisher's swap window already produces.
#
# The rename disturbs nothing already in flight: it unlinks no entry and
# leaves the source inode unchanged, which is precisely why the publish side
# can rotate a snapshot out from under a live reader. A declined eviction
# therefore costs a deferred eviction and nothing else.
evict_dir() {
local dir="$1" name aside
name=$(basename "$dir")
aside="${ROOT}/.evicting-${name}-$$"
mv -T "$dir" "$aside" 2>/dev/null || {
echo " ${name}: could not be set aside for eviction — skipped this pass"
return 1
}
if is_locked "$aside" "$name"; then
if mv -T "$aside" "$dir" 2>/dev/null; then
echo " ${name}: claimed by a job while its eviction was in flight — restored, not evicted"
elif [ -d "$aside" ]; then
echo "::warning::prune: ${name} was claimed mid-eviction and its own name is taken again — leaving $(basename "$aside") for a later pass to reclaim once its readers drain"
else
echo "::warning::prune: ${name} was claimed mid-eviction and is already gone — another pass reclaimed it after its readers drained"
fi
return 1
fi
rm -rf "$aside"
return 0
}
# Directories oldest-first, target-* before snapshot-* (see the header).
# The sort key is a rank digit followed by a zero-padded mtime, so the two
# groups sort as blocks rather than interleaving by age. `.cache-last-used`
# is the marker every run touches; a run that hits the cache for every crate
# may write nothing at all inside the tree, which would make a
# just-used directory look stale without it. Fall back to the directory's
# own mtime when the marker is missing (a partially-written directory from an
# interrupted run is still a valid, if less precise, "last touched" signal).
list_by_lru() {
local rank d ts
for rank in 0:target 1:snapshot; do
for d in "$ROOT"/"${rank#*:}"-*; do
[ -d "$d" ] || continue
if [ -e "$d/.cache-last-used" ]; then ts=$(stat -c '%Y' "$d/.cache-last-used" 2>/dev/null || echo 0)
else ts=$(stat -c '%Y' "$d" 2>/dev/null || echo 0); fi
printf '%s%012d\t%s\n' "${rank%%:*}" "$ts" "$d"
done
done | sort | cut -f2-
}
# Deferred reclamations from an earlier pass: a directory renamed aside for
# eviction that could not be unlinked, because a job claimed it inside the
# window and its own name was taken again before it could be restored — or
# whose run was killed between the rename and the unlink. Nothing below would
# ever see one: every pass globs target-*/snapshot-*, which a dotted name does
# not match. Left unswept it is permanently unreclaimable disk on the one
# volume whose entire problem is disk.
#
# TWO DISTINCT PROPERTIES HOLD HERE, and neither implies the other.
#
# No clone can be truncated by the unlink below, by construction: the aside
# name only comes into existence after the evicting pass's rename, so a
# consumer that resolved the directory published its marker strictly before
# that rename and therefore before the count below — it cannot be missed. A
# consumer that arrives later cannot resolve the path at all. This holds under
# any interleaving and needs no settle window.
#
# No pass mid-eviction is mistaken for a leftover, by bound rather than by
# construction, and this is what the settle window is for. Without it a
# concurrent pass could unlink an aside its owner is still deciding about, and
# `rm -rf` traverses fd-relative: the owner's restore can then republish a
# half-emptied tree under a live cache name. What the window guarantees is that
# the two cannot be confused within it. What it does not guarantee is the
# pathological case beyond it — an evicting pass suspended past the window and
# then resumed finds its aside reclaimed and fails its restore, saying so.
now=$(date +%s)
for aside in "$ROOT"/.evicting-*; do
[ -d "$aside" ] || continue
# Change time, not modification time. `mv` leaves a directory's mtime alone
# (a cache last written days ago keeps a days-old mtime, which is what
# list_by_lru wants and exactly the wrong signal here) but rename(2) does
# update ctime, so %Z is when this directory was set aside and %Y is not.
aside_age=$(( now - $(stat -c '%Z' "$aside" 2>/dev/null || echo "$now") ))
if [ "$aside_age" -lt "$EVICTION_ASIDE_SETTLE_SECONDS" ]; then
echo "prune: $(basename "$aside") was set aside ${aside_age}s ago — another pass may still be evicting it, leaving it alone"
continue
fi
aside_name=$(basename "$aside"); aside_name="${aside_name#.evicting-}"; aside_name="${aside_name%-*}"
if [ "$(live_reader_count "$ROOT" "$aside_name")" -gt 0 ]; then
echo "prune: $(basename "$aside") is still being read — deferring its reclamation again"
continue
fi
echo "prune: reclaiming deferred eviction $(basename "$aside")"
rm -rf "$aside"
done
echo "=== pass 1: liveness (unconditional, not gated on free space) ==="
declare -A live_ns=()
LIVENESS_AVAILABLE=0
LIVENESS_REASON=""
if [ "${CACHE_LIVENESS:-true}" = "0" ] || [ "${CACHE_LIVENESS:-true}" = "false" ]; then
LIVENESS_REASON="disabled via CACHE_LIVENESS=${CACHE_LIVENESS}"
elif remote_heads=$(timeout 20 git ls-remote --heads origin 2>&1); then
LIVENESS_AVAILABLE=1
branch_count=0
while IFS= read -r line; do
[ -n "$line" ] || continue
case "$line" in *refs/heads/*) ;; *) continue ;; esac
branch="${line#*refs/heads/}"
[ -n "$branch" ] || continue
suffix=$(cache_key "$branch")
live_ns["target-${suffix}"]=1
live_ns["snapshot-${suffix}"]=1
branch_count=$((branch_count + 1))
done <<< "$remote_heads"
echo "liveness: ${branch_count} live branches on origin"
else
LIVENESS_REASON="git ls-remote --heads origin failed or timed out"
fi
if [ "$LIVENESS_AVAILABLE" = "1" ]; then
pruned_any=0
# Tracked separately so the line below cannot contradict the decline lines
# above it: "none pruned" and "none found" are different outcomes, and a
# pass that declined every dead cache it found has not found none.
spared_any=0
for dir in "$ROOT"/target-* "$ROOT"/snapshot-*; do
[ -d "$dir" ] || continue
name=$(basename "$dir")
is_protected "$dir" && continue
[ -n "${live_ns[$name]:-}" ] && continue
is_locked "$dir" && { spared_any=1; continue; }
dir_gb=$(usage_gb "$dir")
evict_dir "$dir" || { spared_any=1; continue; }
echo "::warning::pruned dead-branch cache ${name} (${dir_gb} GB) — no matching branch on origin"
summary_line "- pruned dead-branch cache \`${name}\` (${dir_gb} GB) — branch no longer exists on origin"
pruned_any=1
done
if [ "$pruned_any" = "0" ]; then
if [ "$spared_any" = "1" ]; then
echo "every dead-branch cache found is still in use — none pruned this pass"
else
echo "no dead-branch caches found"
fi
fi
else
echo "liveness: ${LIVENESS_REASON} — treating as UNAVAILABLE (not as \"no branches\"); pass 1 skipped"
fi
echo
echo "=== pass 2/3: disk pressure (threshold: free < ${MIN_FREE_PCT}%) ==="
read -r TOTAL_KB FREE_KB <<< "$(read_df "$ROOT")"
THRESHOLD_KB=$(( TOTAL_KB * MIN_FREE_PCT / 100 ))
if [ "$FREE_KB" -ge "$THRESHOLD_KB" ]; then
echo "cache: $(basename "$OWN_DIR") $(usage_gb "$OWN_DIR") GB | $(report_df host "$FREE_KB" "$TOTAL_KB")"
exit 0
fi
echo "::warning::$(report_df disk "$FREE_KB" "$TOTAL_KB") < ${MIN_FREE_PCT}% threshold"
# A plain loop over a pre-materialised, pre-sorted list rather than a live
# `find | while` pipeline, so `rm -rf` inside the loop cannot make a running
# directory walk re-observe its own deletions.
mapfile -t LRU < <(list_by_lru)
for dir in "${LRU[@]}"; do
[ -d "$dir" ] || continue
is_protected "$dir" && continue
read -r TOTAL_KB FREE_KB <<< "$(read_df "$ROOT")"
[ "$FREE_KB" -ge "$THRESHOLD_KB" ] && break
is_locked "$dir" && continue
dir_gb=$(usage_gb "$dir")
evict_dir "$dir" || continue
echo "::warning::evicted $(basename "$dir") (${dir_gb} GB, LRU under disk pressure)"
summary_line "- evicted \`$(basename "$dir")\` (${dir_gb} GB, LRU under disk pressure)"
done
read -r TOTAL_KB FREE_KB <<< "$(read_df "$ROOT")"
if [ "$FREE_KB" -lt "$THRESHOLD_KB" ]; then
OWN_GB=$(usage_gb "$OWN_DIR")
echo "::warning::still under threshold after evicting every eligible sibling; clearing own $(basename "$OWN_DIR") (was ${OWN_GB} GB) — this run pays a cold rebuild"
summary_line "- **self-clear**: \`$(basename "$OWN_DIR")\` (was ${OWN_GB} GB) wiped — this run pays a cold rebuild"
rm -rf "$OWN_DIR"
mkdir -p "$OWN_DIR"
else
echo "$(report_df post-eviction "$FREE_KB" "$TOTAL_KB") — sibling eviction recovered enough space"
fi