5 Commits
Author SHA1 Message Date
claudeandClaude Sonnet 5 ea48c03aeb fix(ci): target origin/main's live tip, not the run's own trigger commit
CI / move v1 to main (pull_request) Skipped
CI / shellcheck + selftests (pull_request) Successful in 1m35s
The wake-one-cancel-the-rest mechanism behind the concurrency group
(CancelPreviousJobsByJobConcurrency, models/actions/run_job.go, at
the v1.27.2 tag this instance runs) picks its survivor from
models/actions/run_job_list.go's query, which carries no `ORDER BY`
-- so under three-way contention on a shared 2-slot runner, the
*newest* commit's job can be the one cancelled while an older sibling
survives and, correctly from its own vantage, advances v1 forward
from a stale view. No push regresses v1 (the ancestor check from
dc1e631 already prevented that), but the newest merge goes silently
unreleased behind a cancelled job that reads as benign, not red --
exactly the failure #27 exists to end.

Every job now resolves `origin/main`'s tip fresh, right before the
push, instead of using `${{ github.sha }}`. Re-fetched explicitly
rather than trusted from the checkout step, which can be minutes
stale by this point behind its own selftest job. Every execution that
reaches the push step now converges on the same target regardless of
which job the concurrency group lets through, so which one wins the
wake no longer matters -- the survivor pushes where any of them
would have.

That doesn't make the push safe on its own: two jobs can still read
main at genuinely different moments if it advances between their two
fetches, so whichever read the tip earlier must not overwrite the
other's already-pushed, newer one. The ancestor check from dc1e631 is
kept for exactly this -- its target changed (origin/main's live tip,
not this job's own trigger commit) but its job didn't.

What each guard now protects against, after this change:
- concurrency group: stops two jobs from pushing at the same time --
  wasted work now that a cancelled job costs nothing, not a
  correctness backstop by itself.
- ancestor check: stops a job whose own fetch of the tip is stale
  relative to another job's already-pushed, fresher one from
  regressing v1.

Rewrote both the job-level comments and README's Versioning section,
which described "force-moves v1 to that commit" and said nothing
about the concurrency group or the skip-as-success case.

Re-derived the truth table against the new target (a live origin/main
tip, not a fixed commit) in a scratch origin+clone: no v1 yet (push),
v1 exactly at the tip (skip), the tip moved past v1 because a newer
merge landed (push, to the new tip -- not stuck on any prior commit),
v1 already ahead of the tip (skip, defensive), unrelated histories
(push, defensive). All five resolved as intended; ran the exact
condition and fetch sequence the workflow step uses, not a
simulation of it.

bash scripts/selftest.sh: all 6 suites green (67 prune-cache
assertions, unchanged). shellcheck -x --source-path=scripts
scripts/*.sh: clean -- as before, this does not cover the inline
`run:` shell in ci.yaml.

What remains unverifiable short of a real merge is unchanged from
dc1e631: the workflow step's actual execution inside a real Actions
run, and whether the built-in token has write access at all. Neither
this commit nor the one before it can exercise those.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 12:49:05 -05:00
claudeandClaude Sonnet 5 dc1e6317c6 fix(ci): make the v1 push monotonic, not just mutually exclusive
The concurrency group added in af1233f only excludes two release-tag
jobs that are simultaneously Running/Waiting/Blocked
(services/actions/clear_tasks.go:64-80,
models/actions/run_job.go:641-660 at the v1.27.2 tag this instance
runs) -- it has no notion of commit order between jobs that never
overlap. On a 2-slot runner shared across 4 repos, with a
multi-minute selftest gating each release-tag job, two merges landing
close together routinely finish in the opposite order from the pushes
that triggered them: if the newer commit's job completes and exits
first, the older commit's job later finds no live holder in the
group, is not blocked, and force-pushes v1 backward to itself. The
concurrency comment's "never regress it to an older one" and af1233f's
commit message both asserted the opposite -- true of the simultaneous
case the guard covers, false of the staggered one it doesn't, so
authored-false rather than drift.

Added a merge-base check before the push: skip if v1 already points
at this commit or a descendant of it (`--is-ancestor` treats a commit
as its own ancestor, so "at" and "ahead" are the same branch). A v1
that doesn't exist yet, or shares no history with this commit, falls
through to the push -- the guard only ever skips, never fails. Needs
`fetch-depth: 0` on the checkout: actions/checkout's own description
for that value is "all history for all branches and tags", and its
source (dist/index.js: fetchDepth <= 0 selects
getRefSpecForAllHistory, which includes the tags refspec) confirms
tags are fetched as part of that, not gated behind the separate
fetch-tags input -- so refs/tags/v1 and the history behind it are both
guaranteed present locally without a second fetch step.

Rewrote both false claims: the concurrency comment now says what the
group actually bounds (simultaneous competing pushes, not completion
order), and states plainly that the guard below is what makes the
outcome order-independent.

Verified the check's five cases (no tag yet, tag ahead of this
commit, tag behind this commit, tag equal to this commit, unrelated
history) against a real scratch git repo, extracting the exact
condition used in the workflow step -- all five resolved as intended
(skip only when v1 is already at or ahead). The workflow step itself
cannot be exercised outside a real Actions run.

bash scripts/selftest.sh: all 6 suites green (67 prune-cache
assertions, unchanged). shellcheck -x --source-path=scripts
scripts/*.sh: clean -- note this does not cover the inline `run:`
shell in ci.yaml, which shellcheck was never wired to check in this
repo (verified against the CI job itself, which shellchecks only
scripts/*.sh).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 12:24:27 -05:00
claudeandClaude Sonnet 5 af1233f14a fix(ci): serialise release-tag across concurrent merges
The workflow-level concurrency group is keyed per-commit (github.sha,
ci.yaml:19-28) so unrelated pushes never block each other -- but that
also means two merges landing close together run two concurrent
release-tag jobs, each force-pushing its own commit to v1. If the
older commit's job finishes last, v1 regresses to a stale-but-green
commit and stays there until the next merge corrects it forward.
Bounded blast radius (never a red commit, self-heals on the next
merge), but a silently wrong v1 is the exact failure #27 exists to
end.

Added a job-level `concurrency:` on release-tag with a fixed group
name and `cancel-in-progress: false`. Confirmed this is additive to
the workflow-level group, not a replacement, by reading gitea's source
at the v1.27.2 tag this instance runs (`tea api version`): run-level
and job-level concurrency are separate model fields
(ActionRunAttempt.ConcurrencyGroup vs ActionRunJob.ConcurrencyGroup),
evaluated by separate functions (EvaluateRunConcurrencyFillModel vs
EvaluateJobConcurrencyFillModel, services/actions/concurrency.go) and
enforced by separate cancellation paths (CancelPreviousJobsByRunConcurrency
in models/actions/run.go:364 vs CancelPreviousJobsByJobConcurrency in
models/actions/run_job.go:641) -- job_emitter.go's checkRunConcurrency
checks both groups independently (services/actions/job_emitter.go:210-236).
A fixed, non-sha group name is what serialises release-tag across
commits without touching the per-sha grouping every other job still
relies on. `cancel-in-progress: false` matters because Gitea's default
queue behaviour (documented same as GitHub's: a new queued job
supersedes an older *queued* one in the same group, but not a
*running* one) already gives the newest commit's push priority once
the running job clears; cancelling the running job too would abandon
whichever merge is currently pushing mid-flight, the same defect
wearing different clothes.

Re-verified: YAML parses, shellcheck clean, `bash scripts/selftest.sh`
all 6 suites green (67 prune-cache assertions, unchanged by this
commit).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 12:07:01 -05:00
claudeandClaude Sonnet 5 21b444121d test(prune): red-prove $OWN_DIR protection under genuine disk pressure
Scenario 9 asserted only that target-$OWN existed, checked right after
a no-pressure run (3b) where nothing was ever a candidate for
eviction, and prune-cache.sh's self-clear step unconditionally
recreates an empty $OWN_DIR whenever the run ends under the percentage
floor regardless of what pass 2 did to it. Either way the existence
check passed whether or not the pass-2 guard (protected_reason,
prune-cache.sh:221) actually protected the directory. Moving that
guard's $OWN_DIR check into the pass-1-only predicate — the exact
mutation gitdan-actions#26 describes — left all 64 assertions green,
confirmed here before the fix.

Rewritten to run under a real, shrinking `df` (the scenario-17
pattern: a fake df that re-measures the fixture with `du` on every
call, so eviction genuinely lowers the reported pressure), with
MIN_FREE_PCT=0 and a clone-headroom floor sized so self-clear's
percentage check can never fire — only pass 2's guard decides the
outcome. The fixture gives $OWN_DIR real content and an older
timestamp than a sibling target dir, sized so the requirement is met
by evicting exactly one of them. With the guard removed, $OWN_DIR is
the one evicted (LRU-oldest, and self-clear is structurally disabled
by MIN_FREE_PCT=0 so there is nothing left to recreate it) — the
scenario now fails loudly on the same mutation that left it green
before. Restored and reverified green (67 assertions, up from 64) with
the guard intact.

bash scripts/selftest.sh: all 6 suites green. shellcheck -x
--source-path=scripts scripts/*.sh: clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 11:52:01 -05:00
claudeandClaude Sonnet 5 7f18cb2436 feat(ci): advance v1 automatically once the gate is green on main
Nothing moved v1 when main advanced, so a merged change was inert
until someone remembered to retag by hand — it happened on 2026-09-22
(PR #25 merged, v1 stayed on the previous release) and was only caught
because a person asked whether the tag had moved.

Option 1 from gitdan-actions#27 (automate it) over option 2 (fail loud
while it lags): a `release-tag` job, gated with `needs: selftest` so a
broken build never reaches it, force-moves v1 to the pushed commit
using the run's built-in GITHUB_TOKEN. If that token turns out not to
have write access, the push fails and the job goes red in the Actions
UI — a loud failure either way, not the silent one this replaces.
Whether the token actually has write access here is unverified short
of a real merge; that merge is the next step for this branch.

README's Versioning section documented the old manual step as a
deliberate decision "never something a merge does by itself" — that
claim is now false, so it's rewritten to describe the automated job
and keeps the manual command as the recovery path for when the job
can't push.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 11:51:51 -05:00
3 changed files with 182 additions and 13 deletions
+98
View File
@@ -101,3 +101,101 @@ jobs:
# file.
- name: Selftests
run: bash scripts/selftest.sh
release-tag:
name: move v1 to main
# `needs:` is what makes this "after the gate is green" rather than
# merely "after a push": a failed selftest skips this job outright, so
# v1 can never advance onto a broken build. The `if:` restricts it to an
# actual push to main -- a pull_request run targeting main shares this
# workflow but has no ref worth tagging.
needs: selftest
if: ${{ github.event_name == 'push' && github.ref == 'refs/heads/main' }}
runs-on: ubuntu-latest
timeout-minutes: 2
# The workflow-level group above is keyed per-commit (github.sha) so
# unrelated commits' CI never blocks each other -- which also means two
# merges landing close together can run two concurrent release-tag jobs.
# A job-level `concurrency:` is a second, independent group scoped to
# this job alone -- it does not replace the workflow-level one, it adds
# to it (confirmed by reading gitea's source at the v1.27.2 tag this
# instance runs: run-level and job-level concurrency are separate model
# fields, evaluated and enforced by separate functions --
# CancelPreviousJobsByRunConcurrency vs CancelPreviousJobsByJobConcurrency
# in models/actions/{run,run_job}.go -- not one overriding the other).
#
# This does NOT decide which of several contending jobs gets to push --
# Gitea wakes exactly one Blocked job in the group and cancels the rest
# outright, with no ordering on which one it picks (no `ORDER BY` in the
# query behind CancelPreviousJobsByJobConcurrency,
# models/actions/run_job_list.go). What it buys is cheaper: every
# execution that does reach the push step targets origin/main's live tip
# (below), never its own trigger commit, so it makes no difference which
# one wins -- the survivor pushes where any of them would have, and a
# cancelled job costs nothing. This group's only job is to stop more than
# one job from pushing AT THE SAME TIME, which is wasted work, not a
# correctness risk on its own.
concurrency:
group: release-tag-v1
cancel-in-progress: false
# Requests write access from the run's built-in token (see README's
# Versioning section for what's actually verified about it). Without
# this the checkout below still succeeds -- it's the push that would be
# rejected, which is a red job, not a silent no-op.
permissions:
contents: write
steps:
# fetch-depth: 0 fetches full history AND all tags (actions/checkout's
# own description: "0 indicates all history for all branches and
# tags") -- REQUIRED so refs/tags/v1 and the ancestry behind it are
# both present locally for the merge-base check below, regardless of
# which commit this run happens to be built from.
- uses: actions/checkout@v4
with:
token: ${{ secrets.GITHUB_TOKEN }}
fetch-depth: 0
# This run's own trigger commit (${{ github.sha }}) is deliberately not
# what gets pushed: whichever job the concurrency group above lets
# through is the one that pushes, and that choice carries no relation
# to commit recency, so every execution has to converge on the SAME
# target regardless of which job it is. `origin/main`'s live tip,
# re-fetched here rather than trusted from the checkout above (which
# can be minutes stale by this point, behind its own selftest job), is
# that common target -- read fresh, every job that reaches this step
# resolves to the same commit whenever main hasn't moved between them,
# and to whatever's newest when it has.
#
# A live target doesn't make the push itself safe on its own: two jobs
# can still read main at genuinely different moments if it advances
# between their two fetches, so the one with the earlier reading must
# not overwrite the other's already-pushed, newer one. That's what the
# ancestor check below still guards -- not "this job's stale trigger
# commit" any more, but "this job's freshly-read tip, which another
# job's fresher read may have already superseded." `--is-ancestor`
# treats a commit as its own ancestor, so "already at" and "already
# ahead" are one case. A v1 that doesn't exist yet, or that shares no
# history with this tip, falls through to the push -- the guard is
# only ever a reason to skip, never a reason to fail.
- name: Determine origin/main's tip and whether v1 needs to move
id: check
run: |
git fetch origin main
TIP=$(git rev-parse origin/main)
echo "tip=$TIP" >> "$GITHUB_OUTPUT"
if git rev-parse -q --verify refs/tags/v1 >/dev/null \
&& git merge-base --is-ancestor "$TIP" refs/tags/v1; then
echo "v1 already at or ahead of origin/main's tip ($TIP) -- nothing to do"
echo "skip=true" >> "$GITHUB_OUTPUT"
else
echo "skip=false" >> "$GITHUB_OUTPUT"
fi
# Lightweight tag, matching what v1 already is (`git cat-file -t v1`
# reports `commit`, not `tag`) -- no identity needed to move it, only
# to push it.
- name: Force v1 to origin/main's tip
if: steps.check.outputs.skip != 'true'
run: |
git tag -f v1 "${{ steps.check.outputs.tip }}"
git push --force origin v1
+26 -10
View File
@@ -634,13 +634,30 @@ entry, another permission — is a `v2`, not a `v1` move. Everything else moves
`v1`: correctness fixes, new optional inputs, and anything internal to
`scripts/`.
**Moving the tag is a release step, and it is the operator's.** Merging to
`main` ships nothing to anybody. `v1` is a lightweight tag and does not follow
a branch, so until it is re-pointed every consumer keeps fetching the commit it
already named, whatever `main` now says. The gap is deliberate: re-pointing
`v1` changes what another repository's CI executes on its next run, so it is a
decision taken once, knowingly, after the merge — never something a merge does
by itself.
**Moving the tag is automatic, gated on the same build that gates a PR.** A
`release-tag` job in `.gitea/workflows/ci.yaml` runs on every push to `main`,
`needs: selftest`, and force-moves `v1` to `origin/main`'s live tip once
selftest succeeds — a broken build never reaches it, so `v1` can't advance
onto one. It pushes with the run's built-in `GITHUB_TOKEN`; if that token
turns out not to have write access, the push step fails and the job goes red
in the Actions UI. That's a loud failure, not the silent one this replaced:
`v1` stays put, and nobody has to notice on their own that it lagged.
Two merges landing close together can start two `release-tag` jobs at once; a
job-level `concurrency` group lets only one push at a time, and every job
targets `origin/main`'s current tip rather than its own trigger commit, so it
makes no difference which one the group lets through. Before pushing, the job
also checks whether `v1` already points at that tip or a descendant of it —
covering the case where two jobs read the tip at genuinely different moments
— and skips as a normal, successful outcome rather than pushing backward. A
run whose log says "nothing to do" did its job correctly; it just found
nothing to move.
This used to be a manual step, treated as a deliberate release decision taken
once, knowingly, after the merge — in practice it was still forgotten
(gitdan-actions#27): PR #25 merged to `main` and `v1` stayed on the previous
release until someone asked whether it had moved. The manual form below is
still the recovery path, for when the automated job can't push:
```bash
git fetch origin
@@ -651,9 +668,8 @@ git ls-remote --tags origin v1 # must equal git rev-parse origin/main
**Downstream** are emowheel, which pins `cargo-cache@v1` and
`cargo-cache-publish@v1` across its CI workflow, and zemyna, migrating to the
same pin. Both pick a move up on their next run with no change on their side,
which is the whole point of the moving pointer and also the reason the move is
not automatic.
same pin. Both pick a move up automatically on their next run with no change
on their side, which is the whole point of the moving pointer.
---
+58 -3
View File
@@ -33,7 +33,16 @@
# so evicting it frees almost nothing while costing every future PR its
# warm start.
# 8. SELF-CLEAR REPORTS LOUDLY to the job summary, not just a log warning.
# 9. OWN CACHE NEVER EVICTED by a sibling pass.
# 9. OWN CACHE NEVER EVICTED by a sibling pass, genuinely under pressure —
# against a real, shrinking `df` (gitdan-actions#26): the static
# CACHE_DF_OVERRIDE every other scenario uses never reflects an
# eviction, so pass 3's self-clear (`rm -rf "$OWN_DIR"; mkdir -p
# "$OWN_DIR"`) fires regardless and recreates an empty OWN_DIR whether
# pass 2 touched it or not — existence survives either way, which is
# why an existence-only assertion here passed even with the pass-2
# guard removed. This one checks CONTENTS, and sizes the requirement
# so it is satisfiable without self-clear at all: only pass 2's guard
# decides the outcome.
# 10. SCOPED TO THE CACHE ROOT — a decoy outside it (standing in for another
# project's volume) is never touched.
# 11. A LIVE READER MARKER PROTECTS A CACHE the same way a lock file does — a
@@ -174,8 +183,54 @@ fi
ok "no protected ref's own target dir reaches the merged-branch check"
echo
echo "=== 9: own cache never evicted by a sibling pass ==="
assert_kept "$root/target-$OWN" "this run's own cache survives"
echo "=== 9: own cache survives a sibling pass genuinely under pressure ==="
# A real, shrinking `df` (the scenario-17 pattern), not CACHE_DF_OVERRIDE:
# eviction has to actually free space for "pressure eases once enough is
# freed" to mean anything. MIN_FREE_PCT=0 and a clone-headroom floor (not
# the percentage floor) drive the requirement, so the requirement is an
# exact, chosen KB rather than a percentage of a volume size this fixture
# would otherwise have to reverse-engineer.
rm -rf "$root"; mkdir -p "$root"
blob_kb=4096
mkdir -p "$root/target-$OWN"
head -c $((blob_kb * 1024)) /dev/zero > "$root/target-$OWN/blob"
touch -d '2020-01-01' "$root/target-$OWN/.cache-last-used"
mkdir -p "$root/target-$LIVE"
head -c $((blob_kb * 1024)) /dev/zero > "$root/target-$LIVE/blob"
touch -d '2021-01-01' "$root/target-$LIVE/.cache-last-used"
# The clone-headroom lookup's base-snapshot candidate — never read for its
# content (CACHE_CLONE_HEADROOM_PERCENT=0 below), only for existing so the
# floor alone becomes the requirement.
mkdir -p "$root/snapshot-$DEV"
real_du=$(command -v du)
used9=$($real_du -sk "$root" | awk '{print $1}')
cap9=$(( used9 + 2048 )) # 2 MB to spare: under the requirement, over nothing else
mkdir -p "$scratch/bin9"
cat > "$scratch/bin9/df" <<DFEOF
#!/usr/bin/env bash
used=\$($real_du -sk "$root" | awk '{print \$1}')
echo "Filesystem 1024-blocks Used Available Capacity Mounted-on"
echo "fake $cap9 \$used \$(( $cap9 - used )) 50% $root"
DFEOF
chmod +x "$scratch/bin9/df"
# Floor sits strictly between "0 evicted" (2048 KB free) and "1 evicted"
# (2048 + blob_kb free) — satisfiable by evicting exactly one candidate.
PATH="$scratch/bin9:$outer_path" \
CACHE_CLONE_HEADROOM_PERCENT=0 CACHE_CLONE_HEADROOM_FLOOR_KB=$(( 2048 + blob_kb / 2 )) \
GITHUB_STEP_SUMMARY="$scratch/summary" \
bash "$prune" "$root" "$root/target-$OWN" "dev main" 0 \
"$(cache_key unused-clone-probe)" "$DEV" "" \
> "$scratch/log" 2>&1 \
|| { cat "$scratch/log"; fail "prune-cache.sh exited non-zero"; }
assert_kept "$root/target-$OWN" "this run's own cache directory survives a genuinely pressured sibling pass"
assert_kept "$root/target-$OWN/blob" "and its contents survive — not a recreated empty directory"
assert_gone "$root/target-$LIVE" "the sibling is evicted instead, to make the same room"
if grep -q 'self-clear' "$scratch/log"; then
fail "own cache was cleared by pass 3, not genuinely spared by pass 2 — this scenario proves nothing"
fi
ok "the requirement was met by pass 2 alone; pass 3 never ran"
echo
echo "=== 7: target dirs evicted before snapshots ==="