From 7f18cb24362a0665c75175bb2f4400c6f7f4bb33 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 22 Sep 2026 11:51:51 -0500 Subject: [PATCH 1/8] feat(ci): advance v1 automatically once the gate is green on main MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Nothing moved v1 when main advanced, so a merged change was inert until someone remembered to retag by hand — it happened on 2026-09-22 (PR #25 merged, v1 stayed on the previous release) and was only caught because a person asked whether the tag had moved. Option 1 from gitdan-actions#27 (automate it) over option 2 (fail loud while it lags): a `release-tag` job, gated with `needs: selftest` so a broken build never reaches it, force-moves v1 to the pushed commit using the run's built-in GITHUB_TOKEN. If that token turns out not to have write access, the push fails and the job goes red in the Actions UI — a loud failure either way, not the silent one this replaces. Whether the token actually has write access here is unverified short of a real merge; that merge is the next step for this branch. README's Versioning section documented the old manual step as a deliberate decision "never something a merge does by itself" — that claim is now false, so it's rewritten to describe the automated job and keeps the manual command as the recovery path for when the job can't push. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb --- .gitea/workflows/ci.yaml | 30 ++++++++++++++++++++++++++++++ README.md | 26 ++++++++++++++++---------- 2 files changed, 46 insertions(+), 10 deletions(-) diff --git a/.gitea/workflows/ci.yaml b/.gitea/workflows/ci.yaml index ac594f0..70053c0 100644 --- a/.gitea/workflows/ci.yaml +++ b/.gitea/workflows/ci.yaml @@ -101,3 +101,33 @@ jobs: # file. - name: Selftests run: bash scripts/selftest.sh + + release-tag: + name: move v1 to main + # `needs:` is what makes this "after the gate is green" rather than + # merely "after a push": a failed selftest skips this job outright, so + # v1 can never advance onto a broken build. The `if:` restricts it to an + # actual push to main -- a pull_request run targeting main shares this + # workflow but has no ref worth tagging. + needs: selftest + if: ${{ github.event_name == 'push' && github.ref == 'refs/heads/main' }} + runs-on: ubuntu-latest + timeout-minutes: 2 + # Requests write access from the run's built-in token (see README's + # Versioning section for what's actually verified about it). Without + # this the checkout below still succeeds -- it's the push that would be + # rejected, which is a red job, not a silent no-op. + permissions: + contents: write + steps: + - uses: actions/checkout@v4 + with: + token: ${{ secrets.GITHUB_TOKEN }} + + # Lightweight tag, matching what v1 already is (`git cat-file -t v1` + # reports `commit`, not `tag`) -- no identity needed to move it, only + # to push it. + - name: Force v1 to this commit + run: | + git tag -f v1 "${{ github.sha }}" + git push --force origin v1 diff --git a/README.md b/README.md index ef66069..9b916af 100644 --- a/README.md +++ b/README.md @@ -634,13 +634,20 @@ entry, another permission — is a `v2`, not a `v1` move. Everything else moves `v1`: correctness fixes, new optional inputs, and anything internal to `scripts/`. -**Moving the tag is a release step, and it is the operator's.** Merging to -`main` ships nothing to anybody. `v1` is a lightweight tag and does not follow -a branch, so until it is re-pointed every consumer keeps fetching the commit it -already named, whatever `main` now says. The gap is deliberate: re-pointing -`v1` changes what another repository's CI executes on its next run, so it is a -decision taken once, knowingly, after the merge — never something a merge does -by itself. +**Moving the tag is automatic, gated on the same build that gates a PR.** A +`release-tag` job in `.gitea/workflows/ci.yaml` runs on every push to `main`, +`needs: selftest`, and force-moves `v1` to that commit once selftest +succeeds — a broken build never reaches it, so `v1` can't advance onto one. +It pushes with the run's built-in `GITHUB_TOKEN`; if that token turns out not +to have write access, the push step fails and the job goes red in the Actions +UI. That's a loud failure, not the silent one this replaced: `v1` stays put, +and nobody has to notice on their own that it lagged. + +This used to be a manual step, treated as a deliberate release decision taken +once, knowingly, after the merge — in practice it was still forgotten +(gitdan-actions#27): PR #25 merged to `main` and `v1` stayed on the previous +release until someone asked whether it had moved. The manual form below is +now the recovery path, for when the automated job can't push: ```bash git fetch origin @@ -651,9 +658,8 @@ git ls-remote --tags origin v1 # must equal git rev-parse origin/main **Downstream** are emowheel, which pins `cargo-cache@v1` and `cargo-cache-publish@v1` across its CI workflow, and zemyna, migrating to the -same pin. Both pick a move up on their next run with no change on their side, -which is the whole point of the moving pointer and also the reason the move is -not automatic. +same pin. Both pick a move up automatically on their next run with no change +on their side, which is the whole point of the moving pointer. --- From 21b444121de16b266cce27a3d58c63cb1862613f Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 22 Sep 2026 11:52:01 -0500 Subject: [PATCH 2/8] test(prune): red-prove $OWN_DIR protection under genuine disk pressure MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Scenario 9 asserted only that target-$OWN existed, checked right after a no-pressure run (3b) where nothing was ever a candidate for eviction, and prune-cache.sh's self-clear step unconditionally recreates an empty $OWN_DIR whenever the run ends under the percentage floor regardless of what pass 2 did to it. Either way the existence check passed whether or not the pass-2 guard (protected_reason, prune-cache.sh:221) actually protected the directory. Moving that guard's $OWN_DIR check into the pass-1-only predicate — the exact mutation gitdan-actions#26 describes — left all 64 assertions green, confirmed here before the fix. Rewritten to run under a real, shrinking `df` (the scenario-17 pattern: a fake df that re-measures the fixture with `du` on every call, so eviction genuinely lowers the reported pressure), with MIN_FREE_PCT=0 and a clone-headroom floor sized so self-clear's percentage check can never fire — only pass 2's guard decides the outcome. The fixture gives $OWN_DIR real content and an older timestamp than a sibling target dir, sized so the requirement is met by evicting exactly one of them. With the guard removed, $OWN_DIR is the one evicted (LRU-oldest, and self-clear is structurally disabled by MIN_FREE_PCT=0 so there is nothing left to recreate it) — the scenario now fails loudly on the same mutation that left it green before. Restored and reverified green (67 assertions, up from 64) with the guard intact. bash scripts/selftest.sh: all 6 suites green. shellcheck -x --source-path=scripts scripts/*.sh: clean. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb --- scripts/prune-cache-selftest.sh | 61 +++++++++++++++++++++++++++++++-- 1 file changed, 58 insertions(+), 3 deletions(-) diff --git a/scripts/prune-cache-selftest.sh b/scripts/prune-cache-selftest.sh index ab0aa17..dc640d0 100755 --- a/scripts/prune-cache-selftest.sh +++ b/scripts/prune-cache-selftest.sh @@ -33,7 +33,16 @@ # so evicting it frees almost nothing while costing every future PR its # warm start. # 8. SELF-CLEAR REPORTS LOUDLY to the job summary, not just a log warning. -# 9. OWN CACHE NEVER EVICTED by a sibling pass. +# 9. OWN CACHE NEVER EVICTED by a sibling pass, genuinely under pressure — +# against a real, shrinking `df` (gitdan-actions#26): the static +# CACHE_DF_OVERRIDE every other scenario uses never reflects an +# eviction, so pass 3's self-clear (`rm -rf "$OWN_DIR"; mkdir -p +# "$OWN_DIR"`) fires regardless and recreates an empty OWN_DIR whether +# pass 2 touched it or not — existence survives either way, which is +# why an existence-only assertion here passed even with the pass-2 +# guard removed. This one checks CONTENTS, and sizes the requirement +# so it is satisfiable without self-clear at all: only pass 2's guard +# decides the outcome. # 10. SCOPED TO THE CACHE ROOT — a decoy outside it (standing in for another # project's volume) is never touched. # 11. A LIVE READER MARKER PROTECTS A CACHE the same way a lock file does — a @@ -174,8 +183,54 @@ fi ok "no protected ref's own target dir reaches the merged-branch check" echo -echo "=== 9: own cache never evicted by a sibling pass ===" -assert_kept "$root/target-$OWN" "this run's own cache survives" +echo "=== 9: own cache survives a sibling pass genuinely under pressure ===" +# A real, shrinking `df` (the scenario-17 pattern), not CACHE_DF_OVERRIDE: +# eviction has to actually free space for "pressure eases once enough is +# freed" to mean anything. MIN_FREE_PCT=0 and a clone-headroom floor (not +# the percentage floor) drive the requirement, so the requirement is an +# exact, chosen KB rather than a percentage of a volume size this fixture +# would otherwise have to reverse-engineer. +rm -rf "$root"; mkdir -p "$root" +blob_kb=4096 +mkdir -p "$root/target-$OWN" +head -c $((blob_kb * 1024)) /dev/zero > "$root/target-$OWN/blob" +touch -d '2020-01-01' "$root/target-$OWN/.cache-last-used" +mkdir -p "$root/target-$LIVE" +head -c $((blob_kb * 1024)) /dev/zero > "$root/target-$LIVE/blob" +touch -d '2021-01-01' "$root/target-$LIVE/.cache-last-used" +# The clone-headroom lookup's base-snapshot candidate — never read for its +# content (CACHE_CLONE_HEADROOM_PERCENT=0 below), only for existing so the +# floor alone becomes the requirement. +mkdir -p "$root/snapshot-$DEV" + +real_du=$(command -v du) +used9=$($real_du -sk "$root" | awk '{print $1}') +cap9=$(( used9 + 2048 )) # 2 MB to spare: under the requirement, over nothing else +mkdir -p "$scratch/bin9" +cat > "$scratch/bin9/df" < "$scratch/log" 2>&1 \ + || { cat "$scratch/log"; fail "prune-cache.sh exited non-zero"; } +assert_kept "$root/target-$OWN" "this run's own cache directory survives a genuinely pressured sibling pass" +assert_kept "$root/target-$OWN/blob" "and its contents survive — not a recreated empty directory" +assert_gone "$root/target-$LIVE" "the sibling is evicted instead, to make the same room" +if grep -q 'self-clear' "$scratch/log"; then + fail "own cache was cleared by pass 3, not genuinely spared by pass 2 — this scenario proves nothing" +fi +ok "the requirement was met by pass 2 alone; pass 3 never ran" echo echo "=== 7: target dirs evicted before snapshots ===" From af1233f14a162db38a0bf81419f165dc17b23344 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 22 Sep 2026 12:07:01 -0500 Subject: [PATCH 3/8] fix(ci): serialise release-tag across concurrent merges The workflow-level concurrency group is keyed per-commit (github.sha, ci.yaml:19-28) so unrelated pushes never block each other -- but that also means two merges landing close together run two concurrent release-tag jobs, each force-pushing its own commit to v1. If the older commit's job finishes last, v1 regresses to a stale-but-green commit and stays there until the next merge corrects it forward. Bounded blast radius (never a red commit, self-heals on the next merge), but a silently wrong v1 is the exact failure #27 exists to end. Added a job-level `concurrency:` on release-tag with a fixed group name and `cancel-in-progress: false`. Confirmed this is additive to the workflow-level group, not a replacement, by reading gitea's source at the v1.27.2 tag this instance runs (`tea api version`): run-level and job-level concurrency are separate model fields (ActionRunAttempt.ConcurrencyGroup vs ActionRunJob.ConcurrencyGroup), evaluated by separate functions (EvaluateRunConcurrencyFillModel vs EvaluateJobConcurrencyFillModel, services/actions/concurrency.go) and enforced by separate cancellation paths (CancelPreviousJobsByRunConcurrency in models/actions/run.go:364 vs CancelPreviousJobsByJobConcurrency in models/actions/run_job.go:641) -- job_emitter.go's checkRunConcurrency checks both groups independently (services/actions/job_emitter.go:210-236). A fixed, non-sha group name is what serialises release-tag across commits without touching the per-sha grouping every other job still relies on. `cancel-in-progress: false` matters because Gitea's default queue behaviour (documented same as GitHub's: a new queued job supersedes an older *queued* one in the same group, but not a *running* one) already gives the newest commit's push priority once the running job clears; cancelling the running job too would abandon whichever merge is currently pushing mid-flight, the same defect wearing different clothes. Re-verified: YAML parses, shellcheck clean, `bash scripts/selftest.sh` all 6 suites green (67 prune-cache assertions, unchanged by this commit). Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb --- .gitea/workflows/ci.yaml | 22 ++++++++++++++++++++++ 1 file changed, 22 insertions(+) diff --git a/.gitea/workflows/ci.yaml b/.gitea/workflows/ci.yaml index 70053c0..7225c59 100644 --- a/.gitea/workflows/ci.yaml +++ b/.gitea/workflows/ci.yaml @@ -113,6 +113,28 @@ jobs: if: ${{ github.event_name == 'push' && github.ref == 'refs/heads/main' }} runs-on: ubuntu-latest timeout-minutes: 2 + # The workflow-level group above is keyed per-commit (github.sha) so + # unrelated commits' CI never blocks each other -- which also means two + # merges landing close together run two concurrent release-tag jobs, each + # force-pushing its OWN commit to v1. If the older one finishes last, v1 + # regresses to a stale-but-green commit until the next merge corrects it. + # + # A job-level `concurrency:` is a second, independent group scoped to + # this job alone -- it does not replace the workflow-level one, it adds + # to it (confirmed by reading gitea's source at the v1.27.2 tag this + # instance runs: run-level and job-level concurrency are separate model + # fields, evaluated and enforced by separate functions -- + # CancelPreviousJobsByRunConcurrency vs CancelPreviousJobsByJobConcurrency + # in models/actions/{run,run_job}.go -- not one overriding the other). + # A fixed, non-sha group name is what serialises across commits: + # `cancel-in-progress: false` keeps a running push from being interrupted + # by a newer one, and Gitea's default "single pending slot" behaviour + # (same as GitHub's) already supersedes any older QUEUED attempt with the + # newest one once the running job clears -- so this can delay v1 landing + # on the newest commit, never regress it to an older one. + concurrency: + group: release-tag-v1 + cancel-in-progress: false # Requests write access from the run's built-in token (see README's # Versioning section for what's actually verified about it). Without # this the checkout below still succeeds -- it's the push that would be From dc1e6317c651429c6e58dd027cb66a6143b67c4d Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 22 Sep 2026 12:24:27 -0500 Subject: [PATCH 4/8] fix(ci): make the v1 push monotonic, not just mutually exclusive The concurrency group added in af1233f only excludes two release-tag jobs that are simultaneously Running/Waiting/Blocked (services/actions/clear_tasks.go:64-80, models/actions/run_job.go:641-660 at the v1.27.2 tag this instance runs) -- it has no notion of commit order between jobs that never overlap. On a 2-slot runner shared across 4 repos, with a multi-minute selftest gating each release-tag job, two merges landing close together routinely finish in the opposite order from the pushes that triggered them: if the newer commit's job completes and exits first, the older commit's job later finds no live holder in the group, is not blocked, and force-pushes v1 backward to itself. The concurrency comment's "never regress it to an older one" and af1233f's commit message both asserted the opposite -- true of the simultaneous case the guard covers, false of the staggered one it doesn't, so authored-false rather than drift. Added a merge-base check before the push: skip if v1 already points at this commit or a descendant of it (`--is-ancestor` treats a commit as its own ancestor, so "at" and "ahead" are the same branch). A v1 that doesn't exist yet, or shares no history with this commit, falls through to the push -- the guard only ever skips, never fails. Needs `fetch-depth: 0` on the checkout: actions/checkout's own description for that value is "all history for all branches and tags", and its source (dist/index.js: fetchDepth <= 0 selects getRefSpecForAllHistory, which includes the tags refspec) confirms tags are fetched as part of that, not gated behind the separate fetch-tags input -- so refs/tags/v1 and the history behind it are both guaranteed present locally without a second fetch step. Rewrote both false claims: the concurrency comment now says what the group actually bounds (simultaneous competing pushes, not completion order), and states plainly that the guard below is what makes the outcome order-independent. Verified the check's five cases (no tag yet, tag ahead of this commit, tag behind this commit, tag equal to this commit, unrelated history) against a real scratch git repo, extracting the exact condition used in the workflow step -- all five resolved as intended (skip only when v1 is already at or ahead). The workflow step itself cannot be exercised outside a real Actions run. bash scripts/selftest.sh: all 6 suites green (67 prune-cache assertions, unchanged). shellcheck -x --source-path=scripts scripts/*.sh: clean -- note this does not cover the inline `run:` shell in ci.yaml, which shellcheck was never wired to check in this repo (verified against the CI job itself, which shellchecks only scripts/*.sh). Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb --- .gitea/workflows/ci.yaml | 48 +++++++++++++++++++++++++++++++--------- 1 file changed, 38 insertions(+), 10 deletions(-) diff --git a/.gitea/workflows/ci.yaml b/.gitea/workflows/ci.yaml index 7225c59..c4c8fbf 100644 --- a/.gitea/workflows/ci.yaml +++ b/.gitea/workflows/ci.yaml @@ -115,10 +115,7 @@ jobs: timeout-minutes: 2 # The workflow-level group above is keyed per-commit (github.sha) so # unrelated commits' CI never blocks each other -- which also means two - # merges landing close together run two concurrent release-tag jobs, each - # force-pushing its OWN commit to v1. If the older one finishes last, v1 - # regresses to a stale-but-green commit until the next merge corrects it. - # + # merges landing close together can run two concurrent release-tag jobs. # A job-level `concurrency:` is a second, independent group scoped to # this job alone -- it does not replace the workflow-level one, it adds # to it (confirmed by reading gitea's source at the v1.27.2 tag this @@ -126,12 +123,13 @@ jobs: # fields, evaluated and enforced by separate functions -- # CancelPreviousJobsByRunConcurrency vs CancelPreviousJobsByJobConcurrency # in models/actions/{run,run_job}.go -- not one overriding the other). - # A fixed, non-sha group name is what serialises across commits: - # `cancel-in-progress: false` keeps a running push from being interrupted - # by a newer one, and Gitea's default "single pending slot" behaviour - # (same as GitHub's) already supersedes any older QUEUED attempt with the - # newest one once the running job clears -- so this can delay v1 landing - # on the newest commit, never regress it to an older one. + # This serialises jobs that are actually running or queued AT THE SAME + # TIME -- it has no notion of commit order between jobs that never + # overlap. Two release-tag runs on a 2-slot, 4-repo runner with a + # multi-minute selftest ahead of each can easily not overlap at all, + # finishing in either order regardless of push order. The step below is + # what makes the OUTCOME order-independent; this group only bounds how + # much work is wasted getting there. concurrency: group: release-tag-v1 cancel-in-progress: false @@ -142,14 +140,44 @@ jobs: permissions: contents: write steps: + # fetch-depth: 0 fetches full history AND all tags (actions/checkout's + # own description: "0 indicates all history for all branches and + # tags") -- REQUIRED so refs/tags/v1 and the ancestry behind it are + # both present locally for the merge-base check below, regardless of + # which commit this run happens to be built from. - uses: actions/checkout@v4 with: token: ${{ secrets.GITHUB_TOKEN }} + fetch-depth: 0 + + # The concurrency group above only excludes a SIMULTANEOUS competing + # push; it does nothing for two release-tag jobs that never overlap + # and finish in the opposite order from the merges that triggered + # them -- which a shared 2-slot runner and a multi-minute selftest + # ahead of this job make routine, not rare. So the push itself has to + # be safe regardless of completion order: skip whenever v1 already + # points at this commit or a descendant of it, rather than trusting + # this run to be the last one to finish. `--is-ancestor` treats a + # commit as its own ancestor, so "already at" and "already ahead" are + # one case. A v1 that doesn't exist yet, or that shares no history + # with this commit, falls through to the push below -- the guard is + # only ever a reason to skip, never a reason to fail. + - name: Skip if v1 already at or ahead of this commit + id: check + run: | + if git rev-parse -q --verify refs/tags/v1 >/dev/null \ + && git merge-base --is-ancestor "${{ github.sha }}" refs/tags/v1; then + echo "v1 already at or ahead of ${{ github.sha }} -- nothing to do" + echo "skip=true" >> "$GITHUB_OUTPUT" + else + echo "skip=false" >> "$GITHUB_OUTPUT" + fi # Lightweight tag, matching what v1 already is (`git cat-file -t v1` # reports `commit`, not `tag`) -- no identity needed to move it, only # to push it. - name: Force v1 to this commit + if: steps.check.outputs.skip != 'true' run: | git tag -f v1 "${{ github.sha }}" git push --force origin v1 From ea48c03aeb11c75473bba472418a2e0944134945 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 22 Sep 2026 12:49:05 -0500 Subject: [PATCH 5/8] fix(ci): target origin/main's live tip, not the run's own trigger commit The wake-one-cancel-the-rest mechanism behind the concurrency group (CancelPreviousJobsByJobConcurrency, models/actions/run_job.go, at the v1.27.2 tag this instance runs) picks its survivor from models/actions/run_job_list.go's query, which carries no `ORDER BY` -- so under three-way contention on a shared 2-slot runner, the *newest* commit's job can be the one cancelled while an older sibling survives and, correctly from its own vantage, advances v1 forward from a stale view. No push regresses v1 (the ancestor check from dc1e631 already prevented that), but the newest merge goes silently unreleased behind a cancelled job that reads as benign, not red -- exactly the failure #27 exists to end. Every job now resolves `origin/main`'s tip fresh, right before the push, instead of using `${{ github.sha }}`. Re-fetched explicitly rather than trusted from the checkout step, which can be minutes stale by this point behind its own selftest job. Every execution that reaches the push step now converges on the same target regardless of which job the concurrency group lets through, so which one wins the wake no longer matters -- the survivor pushes where any of them would have. That doesn't make the push safe on its own: two jobs can still read main at genuinely different moments if it advances between their two fetches, so whichever read the tip earlier must not overwrite the other's already-pushed, newer one. The ancestor check from dc1e631 is kept for exactly this -- its target changed (origin/main's live tip, not this job's own trigger commit) but its job didn't. What each guard now protects against, after this change: - concurrency group: stops two jobs from pushing at the same time -- wasted work now that a cancelled job costs nothing, not a correctness backstop by itself. - ancestor check: stops a job whose own fetch of the tip is stale relative to another job's already-pushed, fresher one from regressing v1. Rewrote both the job-level comments and README's Versioning section, which described "force-moves v1 to that commit" and said nothing about the concurrency group or the skip-as-success case. Re-derived the truth table against the new target (a live origin/main tip, not a fixed commit) in a scratch origin+clone: no v1 yet (push), v1 exactly at the tip (skip), the tip moved past v1 because a newer merge landed (push, to the new tip -- not stuck on any prior commit), v1 already ahead of the tip (skip, defensive), unrelated histories (push, defensive). All five resolved as intended; ran the exact condition and fetch sequence the workflow step uses, not a simulation of it. bash scripts/selftest.sh: all 6 suites green (67 prune-cache assertions, unchanged). shellcheck -x --source-path=scripts scripts/*.sh: clean -- as before, this does not cover the inline `run:` shell in ci.yaml. What remains unverifiable short of a real merge is unchanged from dc1e631: the workflow step's actual execution inside a real Actions run, and whether the built-in token has write access at all. Neither this commit nor the one before it can exercise those. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb --- .gitea/workflows/ci.yaml | 64 +++++++++++++++++++++++++--------------- README.md | 24 ++++++++++----- 2 files changed, 58 insertions(+), 30 deletions(-) diff --git a/.gitea/workflows/ci.yaml b/.gitea/workflows/ci.yaml index c4c8fbf..9f00531 100644 --- a/.gitea/workflows/ci.yaml +++ b/.gitea/workflows/ci.yaml @@ -123,13 +123,18 @@ jobs: # fields, evaluated and enforced by separate functions -- # CancelPreviousJobsByRunConcurrency vs CancelPreviousJobsByJobConcurrency # in models/actions/{run,run_job}.go -- not one overriding the other). - # This serialises jobs that are actually running or queued AT THE SAME - # TIME -- it has no notion of commit order between jobs that never - # overlap. Two release-tag runs on a 2-slot, 4-repo runner with a - # multi-minute selftest ahead of each can easily not overlap at all, - # finishing in either order regardless of push order. The step below is - # what makes the OUTCOME order-independent; this group only bounds how - # much work is wasted getting there. + # + # This does NOT decide which of several contending jobs gets to push -- + # Gitea wakes exactly one Blocked job in the group and cancels the rest + # outright, with no ordering on which one it picks (no `ORDER BY` in the + # query behind CancelPreviousJobsByJobConcurrency, + # models/actions/run_job_list.go). What it buys is cheaper: every + # execution that does reach the push step targets origin/main's live tip + # (below), never its own trigger commit, so it makes no difference which + # one wins -- the survivor pushes where any of them would have, and a + # cancelled job costs nothing. This group's only job is to stop more than + # one job from pushing AT THE SAME TIME, which is wasted work, not a + # correctness risk on its own. concurrency: group: release-tag-v1 cancel-in-progress: false @@ -150,24 +155,37 @@ jobs: token: ${{ secrets.GITHUB_TOKEN }} fetch-depth: 0 - # The concurrency group above only excludes a SIMULTANEOUS competing - # push; it does nothing for two release-tag jobs that never overlap - # and finish in the opposite order from the merges that triggered - # them -- which a shared 2-slot runner and a multi-minute selftest - # ahead of this job make routine, not rare. So the push itself has to - # be safe regardless of completion order: skip whenever v1 already - # points at this commit or a descendant of it, rather than trusting - # this run to be the last one to finish. `--is-ancestor` treats a - # commit as its own ancestor, so "already at" and "already ahead" are - # one case. A v1 that doesn't exist yet, or that shares no history - # with this commit, falls through to the push below -- the guard is + # This run's own trigger commit (${{ github.sha }}) is deliberately not + # what gets pushed: whichever job the concurrency group above lets + # through is the one that pushes, and that choice carries no relation + # to commit recency, so every execution has to converge on the SAME + # target regardless of which job it is. `origin/main`'s live tip, + # re-fetched here rather than trusted from the checkout above (which + # can be minutes stale by this point, behind its own selftest job), is + # that common target -- read fresh, every job that reaches this step + # resolves to the same commit whenever main hasn't moved between them, + # and to whatever's newest when it has. + # + # A live target doesn't make the push itself safe on its own: two jobs + # can still read main at genuinely different moments if it advances + # between their two fetches, so the one with the earlier reading must + # not overwrite the other's already-pushed, newer one. That's what the + # ancestor check below still guards -- not "this job's stale trigger + # commit" any more, but "this job's freshly-read tip, which another + # job's fresher read may have already superseded." `--is-ancestor` + # treats a commit as its own ancestor, so "already at" and "already + # ahead" are one case. A v1 that doesn't exist yet, or that shares no + # history with this tip, falls through to the push -- the guard is # only ever a reason to skip, never a reason to fail. - - name: Skip if v1 already at or ahead of this commit + - name: Determine origin/main's tip and whether v1 needs to move id: check run: | + git fetch origin main + TIP=$(git rev-parse origin/main) + echo "tip=$TIP" >> "$GITHUB_OUTPUT" if git rev-parse -q --verify refs/tags/v1 >/dev/null \ - && git merge-base --is-ancestor "${{ github.sha }}" refs/tags/v1; then - echo "v1 already at or ahead of ${{ github.sha }} -- nothing to do" + && git merge-base --is-ancestor "$TIP" refs/tags/v1; then + echo "v1 already at or ahead of origin/main's tip ($TIP) -- nothing to do" echo "skip=true" >> "$GITHUB_OUTPUT" else echo "skip=false" >> "$GITHUB_OUTPUT" @@ -176,8 +194,8 @@ jobs: # Lightweight tag, matching what v1 already is (`git cat-file -t v1` # reports `commit`, not `tag`) -- no identity needed to move it, only # to push it. - - name: Force v1 to this commit + - name: Force v1 to origin/main's tip if: steps.check.outputs.skip != 'true' run: | - git tag -f v1 "${{ github.sha }}" + git tag -f v1 "${{ steps.check.outputs.tip }}" git push --force origin v1 diff --git a/README.md b/README.md index 9b916af..43ecf61 100644 --- a/README.md +++ b/README.md @@ -636,18 +636,28 @@ entry, another permission — is a `v2`, not a `v1` move. Everything else moves **Moving the tag is automatic, gated on the same build that gates a PR.** A `release-tag` job in `.gitea/workflows/ci.yaml` runs on every push to `main`, -`needs: selftest`, and force-moves `v1` to that commit once selftest -succeeds — a broken build never reaches it, so `v1` can't advance onto one. -It pushes with the run's built-in `GITHUB_TOKEN`; if that token turns out not -to have write access, the push step fails and the job goes red in the Actions -UI. That's a loud failure, not the silent one this replaced: `v1` stays put, -and nobody has to notice on their own that it lagged. +`needs: selftest`, and force-moves `v1` to `origin/main`'s live tip once +selftest succeeds — a broken build never reaches it, so `v1` can't advance +onto one. It pushes with the run's built-in `GITHUB_TOKEN`; if that token +turns out not to have write access, the push step fails and the job goes red +in the Actions UI. That's a loud failure, not the silent one this replaced: +`v1` stays put, and nobody has to notice on their own that it lagged. + +Two merges landing close together can start two `release-tag` jobs at once; a +job-level `concurrency` group lets only one push at a time, and every job +targets `origin/main`'s current tip rather than its own trigger commit, so it +makes no difference which one the group lets through. Before pushing, the job +also checks whether `v1` already points at that tip or a descendant of it — +covering the case where two jobs read the tip at genuinely different moments +— and skips as a normal, successful outcome rather than pushing backward. A +run whose log says "nothing to do" did its job correctly; it just found +nothing to move. This used to be a manual step, treated as a deliberate release decision taken once, knowingly, after the merge — in practice it was still forgotten (gitdan-actions#27): PR #25 merged to `main` and `v1` stayed on the previous release until someone asked whether it had moved. The manual form below is -now the recovery path, for when the automated job can't push: +still the recovery path, for when the automated job can't push: ```bash git fetch origin From 17d87b06471f48021382a7b98ef2e6919ded45ed Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 22 Sep 2026 14:01:28 -0500 Subject: [PATCH 6/8] fix(ci): only ever push a commit this run actually gated; fix vacuous scenario-9 guard ## v1 could advance onto an ungated commit `needs: selftest` gates this run's own commit, but the push targeted origin/main's freshly-fetched tip with nothing comparing the two. Trace: M1 merges green; M2 merges while M1's selftest is still running; M1's release-tag job fetches tip = M2 and pushes v1 = M2, whose own selftest may be queued, running, or red. If M2 is red, its own job is skipped, so v1 sits on a red commit across every consuming project until the next green merge -- with nothing red pointing at the release itself. ci.yaml:107-109 and README.md:640-641 both asserted this couldn't happen; ea48c03's own diff established the precondition (tip "can be minutes stale... behind its own selftest job") without closing it. Fix: skip the push unless origin/main's tip IS this run's own github.sha, checked before the existing v1-monotonicity check (dc1e631) rather than replacing it -- the two compose (tip-mismatch first, since it's the coarser reason to defer; ancestor-check second, for a duplicate run whose own commit is still current). Reverted the push target from the fetched tip back to `${{ github.sha }}`, now that the guard makes them provably equal whenever the push fires. ## Convergence trace: does the newest commit's job still run? Read gitea's source further at the pinned v1.27.2 tag: PrepareToStartJobWithConcurrency (services/actions/clear_tasks.go) calls CancelPreviousJobsByJobConcurrency on every job entering the group, unconditionally cancelling whatever was previously Waiting/Blocked there -- so at most one job sits queued in the group at a time; each new arrival supersedes it. Because job-level concurrency is only evaluated once `needs: selftest` is satisfied (job_emitter.go re-evaluates readiness there), "arrival order" tracks each commit's own selftest-completion time, not raw merge order -- an older commit with a slower selftest can enter the group after a newer one and cancel its queued slot. That cancelled job is gone for good; it will never push. But the commit that's genuinely current at the moment merges stop arriving is always the one still queued when the running job finishes, because every subsequent arrival (from every subsequent merge, not just the "newest" one at any single instant) keeps re-superseding the queue. So v1 always eventually catches up -- "one merge later" in the common case, "at the next merge, whenever that happens" in the adversarial case where a stale survivor runs, finds itself no longer current, defers, and nothing is left queued. It cannot get stuck forever short of the repository never receiving another merge, because every future push re-attempts the same check against whatever's current by then. Stated this plainly in the comment and README rather than repeating the false "never" guarantee in softer words. Re-derived the truth table against the new guard in a scratch origin+clone, six cases: own commit == tip, no v1 (push); v1 already == own commit (skip, duplicate run); tip moved past own gated commit because a newer merge landed (skip, defers); the newer commit's own run once nothing further has landed (push); v1 already ahead of a now-stale gated commit (skip, tip-mismatch catches it first); tip == own commit but v1 independently ahead via a local-only descendant, isolating the second (ancestor) check on its own (skip). All six resolved as intended. ## Scenario 9's pass-3 guard was vacuous scripts/prune-cache-selftest.sh:230 grepped 'self-clear' against $scratch/log. summary_line() (cache-lib.sh) writes only to $GITHUB_STEP_SUMMARY; the self-clear branch (prune-cache.sh:632) writes 'self-clear' there and 'clearing own' to stdout (:631) -- 'self-clear' never appears in $scratch/log at all, so the branch was unreachable and the `ok` unconditional. Scenario 8 already greps the right string against the right file (assert_log "clearing own" ...); scenario 9 now does the same, staying on $scratch/log where it already was -- the string was wrong, not the file. Red-proved by capturing a real prune-cache.sh log where self-clear genuinely fired (scenario 9's own fixture with MIN_FREE_PCT temporarily raised to 100, in a scratch copy, reverted after) and running both patterns against it: `grep -q 'self-clear'` -> no match (the old check's vacuous pass, confirmed); `grep -q 'clearing own'` -> match (the fix's correct fail). Matches the reviewer's own measurement exactly. The real prune-cache-selftest.sh was untouched during this experiment; only the grep string changed in the actual commit. bash scripts/selftest.sh: all 6 suites green (67 prune-cache assertions, unchanged in count -- the fix corrects what scenario 9's existing check compares, not what it asserts). shellcheck -x --source-path=scripts scripts/*.sh: clean, as before not covering the inline ci.yaml shell. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb --- .gitea/workflows/ci.yaml | 74 +++++++++++++++++---------------- README.md | 24 ++++++----- scripts/prune-cache-selftest.sh | 2 +- 3 files changed, 54 insertions(+), 46 deletions(-) diff --git a/.gitea/workflows/ci.yaml b/.gitea/workflows/ci.yaml index 9f00531..5d01627 100644 --- a/.gitea/workflows/ci.yaml +++ b/.gitea/workflows/ci.yaml @@ -128,13 +128,12 @@ jobs: # Gitea wakes exactly one Blocked job in the group and cancels the rest # outright, with no ordering on which one it picks (no `ORDER BY` in the # query behind CancelPreviousJobsByJobConcurrency, - # models/actions/run_job_list.go). What it buys is cheaper: every - # execution that does reach the push step targets origin/main's live tip - # (below), never its own trigger commit, so it makes no difference which - # one wins -- the survivor pushes where any of them would have, and a - # cancelled job costs nothing. This group's only job is to stop more than - # one job from pushing AT THE SAME TIME, which is wasted work, not a - # correctness risk on its own. + # models/actions/run_job_list.go). Every execution still only ever pushes + # its OWN gated commit (below), never another job's, so a job that gets + # cancelled here costs nothing beyond its own wasted run -- it was never + # going to push anyone else's commit either. This group's only job is to + # stop more than one job from pushing AT THE SAME TIME, which is wasted + # work, not a correctness risk on its own. concurrency: group: release-tag-v1 cancel-in-progress: false @@ -155,37 +154,42 @@ jobs: token: ${{ secrets.GITHUB_TOKEN }} fetch-depth: 0 - # This run's own trigger commit (${{ github.sha }}) is deliberately not - # what gets pushed: whichever job the concurrency group above lets - # through is the one that pushes, and that choice carries no relation - # to commit recency, so every execution has to converge on the SAME - # target regardless of which job it is. `origin/main`'s live tip, - # re-fetched here rather than trusted from the checkout above (which - # can be minutes stale by this point, behind its own selftest job), is - # that common target -- read fresh, every job that reaches this step - # resolves to the same commit whenever main hasn't moved between them, - # and to whatever's newest when it has. + # `needs: selftest` only gates THIS run's own commit -- it says nothing + # about whether `origin/main` has since moved on to a merge whose own + # selftest hasn't finished, is still queued, or is red. Two things + # follow, checked in this order: # - # A live target doesn't make the push itself safe on its own: two jobs - # can still read main at genuinely different moments if it advances - # between their two fetches, so the one with the earlier reading must - # not overwrite the other's already-pushed, newer one. That's what the - # ancestor check below still guards -- not "this job's stale trigger - # commit" any more, but "this job's freshly-read tip, which another - # job's fresher read may have already superseded." `--is-ancestor` - # treats a commit as its own ancestor, so "already at" and "already - # ahead" are one case. A v1 that doesn't exist yet, or that shares no - # history with this tip, falls through to the push -- the guard is - # only ever a reason to skip, never a reason to fail. - - name: Determine origin/main's tip and whether v1 needs to move + # 1. If `origin/main`'s live tip (re-fetched here, not trusted from the + # checkout above, which can be minutes stale behind this job's own + # selftest) is no longer THIS run's own `github.sha`, some other + # merge has landed since. Pushing it would release a commit this run + # never gated -- so this run defers instead, unconditionally. The + # commit that IS the live tip has its own run, and that run's own + # guard is what releases it once ITS turn to push comes -- possibly + # only once merges pause for a moment, so v1 can land one merge + # later than the newest one in a busy stretch. It never lands on a + # commit that wasn't gated, and it always catches up once things go + # quiet, because every future merge re-attempts the same check + # against whatever is current by then. + # 2. Only once (1) confirms this IS the live tip does it matter whether + # v1 already covers it -- a duplicate run, or a manual push already + # having done this. `--is-ancestor` treats a commit as its own + # ancestor, so "already at" and "already ahead" are one case. A v1 + # that doesn't exist yet, or shares no history with this commit, + # falls through to the push -- both checks are only ever a reason to + # skip, never a reason to fail. + - name: Determine whether this commit is still current and needs releasing id: check run: | git fetch origin main TIP=$(git rev-parse origin/main) - echo "tip=$TIP" >> "$GITHUB_OUTPUT" - if git rev-parse -q --verify refs/tags/v1 >/dev/null \ - && git merge-base --is-ancestor "$TIP" refs/tags/v1; then - echo "v1 already at or ahead of origin/main's tip ($TIP) -- nothing to do" + SHA="${{ github.sha }}" + if [ "$TIP" != "$SHA" ]; then + echo "origin/main's tip ($TIP) has moved past this run's own gated commit ($SHA) -- deferring to whichever run's own commit is now the live tip" + echo "skip=true" >> "$GITHUB_OUTPUT" + elif git rev-parse -q --verify refs/tags/v1 >/dev/null \ + && git merge-base --is-ancestor "$SHA" refs/tags/v1; then + echo "v1 already at or ahead of $SHA -- nothing to do" echo "skip=true" >> "$GITHUB_OUTPUT" else echo "skip=false" >> "$GITHUB_OUTPUT" @@ -194,8 +198,8 @@ jobs: # Lightweight tag, matching what v1 already is (`git cat-file -t v1` # reports `commit`, not `tag`) -- no identity needed to move it, only # to push it. - - name: Force v1 to origin/main's tip + - name: Force v1 to this commit if: steps.check.outputs.skip != 'true' run: | - git tag -f v1 "${{ steps.check.outputs.tip }}" + git tag -f v1 "${{ github.sha }}" git push --force origin v1 diff --git a/README.md b/README.md index 43ecf61..c9f646f 100644 --- a/README.md +++ b/README.md @@ -636,22 +636,26 @@ entry, another permission — is a `v2`, not a `v1` move. Everything else moves **Moving the tag is automatic, gated on the same build that gates a PR.** A `release-tag` job in `.gitea/workflows/ci.yaml` runs on every push to `main`, -`needs: selftest`, and force-moves `v1` to `origin/main`'s live tip once +`needs: selftest`, and force-moves `v1` to *that run's own commit* once selftest succeeds — a broken build never reaches it, so `v1` can't advance onto one. It pushes with the run's built-in `GITHUB_TOKEN`; if that token turns out not to have write access, the push step fails and the job goes red in the Actions UI. That's a loud failure, not the silent one this replaced: `v1` stays put, and nobody has to notice on their own that it lagged. -Two merges landing close together can start two `release-tag` jobs at once; a -job-level `concurrency` group lets only one push at a time, and every job -targets `origin/main`'s current tip rather than its own trigger commit, so it -makes no difference which one the group lets through. Before pushing, the job -also checks whether `v1` already points at that tip or a descendant of it — -covering the case where two jobs read the tip at genuinely different moments -— and skips as a normal, successful outcome rather than pushing backward. A -run whose log says "nothing to do" did its job correctly; it just found -nothing to move. +Before pushing, the job checks two things and pushes only if both hold: that +`origin/main`'s live tip is still this run's own commit (not some later merge +that landed while this job was queued behind its own selftest), and that `v1` +doesn't already point at that commit or a descendant of it. Either check can +skip the push, as a normal, successful outcome — a run whose log says +"nothing to do" did its job correctly. The first check is what a job-level +`concurrency` group alone can't guarantee: Gitea's wake-one/cancel-rest +handling applies no ordering by commit recency, so an older merge's job can +be the one that survives to run — that job now defers instead of releasing a +commit it never gated. The newer merge's own job releases it once its own +turn comes, which can land `v1` one merge behind the newest during a busy +stretch; it always catches up once merges pause, because every later run +re-checks against whatever is current by then. This used to be a manual step, treated as a deliberate release decision taken once, knowingly, after the merge — in practice it was still forgotten diff --git a/scripts/prune-cache-selftest.sh b/scripts/prune-cache-selftest.sh index dc640d0..bf90c0e 100755 --- a/scripts/prune-cache-selftest.sh +++ b/scripts/prune-cache-selftest.sh @@ -227,7 +227,7 @@ PATH="$scratch/bin9:$outer_path" \ assert_kept "$root/target-$OWN" "this run's own cache directory survives a genuinely pressured sibling pass" assert_kept "$root/target-$OWN/blob" "and its contents survive — not a recreated empty directory" assert_gone "$root/target-$LIVE" "the sibling is evicted instead, to make the same room" -if grep -q 'self-clear' "$scratch/log"; then +if grep -q 'clearing own' "$scratch/log"; then fail "own cache was cleared by pass 3, not genuinely spared by pass 2 — this scenario proves nothing" fi ok "the requirement was met by pass 2 alone; pass 3 never ran" From ca0ee132d94d3b0d814821054f2a353b64950bee Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 22 Sep 2026 15:01:05 -0500 Subject: [PATCH 7/8] fix(ci): add a scheduled v1 sweep and lease every v1 push The release guard in 17d87b0 was safe but not live. Gitea 1.27.2 calls CancelPreviousJobsByJobConcurrency whenever a job's `needs` resolve (services/actions/clear_tasks.go:91, models/actions/run_job.go:641), so a job's place in the `release-tag-v1` group followed when its own selftest finished, not merge order. A newer merge C2 finishing selftest first queued behind the older C1, C1 cancelled it, saw tip = C2, and deferred: nobody pushed, and if merges then stopped v1 stayed stale indefinitely behind a Skipped and a Cancelled job. The "always catches up once merges pause" claim in ci.yaml and README was false. What now holds: - release-sweep.yaml runs on `schedule` every 15 minutes, in its own workflow and concurrency group, so nothing in ci.yaml can cancel it. When v1 already covers main's tip it stops after a checkout and one merge-base. Otherwise it checks out the tip, runs the same shellcheck and selftest.sh as ci.yaml's selftest job, and tags the tip only if they pass; a failing main therefore turns the sweep red on every tick while v1 lags, which is #27's AC1 loud-failure half. It reads the tip itself because a scheduled run's github.sha is the CommitSHA recorded when the schedule was registered on the last push to main (services/actions/notifier_helper.go:569-580, services/actions/schedule_tasks.go:126-141), and ref is the default branch: schedules are registered only from it (notifier_helper.go:120, :531, :603-604). event_name is "schedule" (context.go:71 reads TriggerEvent, set at schedule_tasks.go:136). Cron is 5-field robfig in UTC (models/actions/schedule_spec.go:38-41). - 15 minutes, not 10: the sweep is the fallback, not the release path, and every tick is a run on gitdan-ci's shared slots and a row in the Actions list. 96 no-op runs a day of a few seconds each is the cost; the lag bound it buys is one interval plus one selftest run. - Both writers go through scripts/release-v1.sh and push with --force-with-lease=refs/tags/v1:, so v1 cannot move backwards when the sweep and a merge job race. A lost lease re-reads v1: at or ahead of this run's gated commit is a clean skip (the other writer released something at least as new); still behind it is a retry leased on the new value, up to three attempts, since the other writer may have tagged an older commit and giving up there would leave v1 short of a commit this run did gate; anything else goes red. A rejection with v1 unmoved is diagnosed as a non-lease failure and goes red at once. - release-tag loses its job-level concurrency group. The lease already gives the ordering the group was there for, and the group was what cancelled the one job that could have released the newest merge. Without it each merge's job runs, and the one whose commit is still the tip when it checks releases it. The shell moves out of ci.yaml into scripts/release-v1.sh so shellcheck and selftest.sh cover it. release-v1-selftest.sh runs it against a scratch bare origin: sweep no-op at and ahead of the tip, tag on a green gate, no tag and a failing sweep on a red one, the stranded trace above followed by a catching-up sweep, and each lost-lease outcome, with a control showing an unleased push does step v1 back. Red-proved by seven mutations of release-v1.sh, each failing a named assertion: plain --force, accepting any lost lease, a merge job that never defers, ancestry reduced to equality, no non-lease diagnosis, a sweep that never needs to run, and a retry that does not re-lease. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb --- .gitea/workflows/ci.yaml | 98 +++-------------- .gitea/workflows/release-sweep.yaml | 69 ++++++++++++ README.md | 43 ++++---- scripts/release-v1-selftest.sh | 165 ++++++++++++++++++++++++++++ scripts/release-v1.sh | 100 +++++++++++++++++ scripts/selftest.sh | 2 +- 6 files changed, 371 insertions(+), 106 deletions(-) create mode 100644 .gitea/workflows/release-sweep.yaml create mode 100644 scripts/release-v1-selftest.sh create mode 100644 scripts/release-v1.sh diff --git a/.gitea/workflows/ci.yaml b/.gitea/workflows/ci.yaml index 5d01627..886e515 100644 --- a/.gitea/workflows/ci.yaml +++ b/.gitea/workflows/ci.yaml @@ -104,39 +104,19 @@ jobs: release-tag: name: move v1 to main - # `needs:` is what makes this "after the gate is green" rather than - # merely "after a push": a failed selftest skips this job outright, so - # v1 can never advance onto a broken build. The `if:` restricts it to an - # actual push to main -- a pull_request run targeting main shares this - # workflow but has no ref worth tagging. + # `needs: selftest` is what makes this "after the gate is green": a failed + # selftest skips this job, so v1 never advances onto a broken build. The + # `if:` restricts it to an actual push to main. + # + # No job-level `concurrency:`. The lease in release-v1.sh already keeps v1 + # from moving backwards, and a group here only cancelled queued jobs in + # whatever order their selftests finished -- which could leave no job to + # release the newest merge. Anything this job defers or misses, + # release-sweep.yaml picks up. needs: selftest if: ${{ github.event_name == 'push' && github.ref == 'refs/heads/main' }} runs-on: ubuntu-latest timeout-minutes: 2 - # The workflow-level group above is keyed per-commit (github.sha) so - # unrelated commits' CI never blocks each other -- which also means two - # merges landing close together can run two concurrent release-tag jobs. - # A job-level `concurrency:` is a second, independent group scoped to - # this job alone -- it does not replace the workflow-level one, it adds - # to it (confirmed by reading gitea's source at the v1.27.2 tag this - # instance runs: run-level and job-level concurrency are separate model - # fields, evaluated and enforced by separate functions -- - # CancelPreviousJobsByRunConcurrency vs CancelPreviousJobsByJobConcurrency - # in models/actions/{run,run_job}.go -- not one overriding the other). - # - # This does NOT decide which of several contending jobs gets to push -- - # Gitea wakes exactly one Blocked job in the group and cancels the rest - # outright, with no ordering on which one it picks (no `ORDER BY` in the - # query behind CancelPreviousJobsByJobConcurrency, - # models/actions/run_job_list.go). Every execution still only ever pushes - # its OWN gated commit (below), never another job's, so a job that gets - # cancelled here costs nothing beyond its own wasted run -- it was never - # going to push anyone else's commit either. This group's only job is to - # stop more than one job from pushing AT THE SAME TIME, which is wasted - # work, not a correctness risk on its own. - concurrency: - group: release-tag-v1 - cancel-in-progress: false # Requests write access from the run's built-in token (see README's # Versioning section for what's actually verified about it). Without # this the checkout below still succeeds -- it's the push that would be @@ -144,62 +124,14 @@ jobs: permissions: contents: write steps: - # fetch-depth: 0 fetches full history AND all tags (actions/checkout's - # own description: "0 indicates all history for all branches and - # tags") -- REQUIRED so refs/tags/v1 and the ancestry behind it are - # both present locally for the merge-base check below, regardless of - # which commit this run happens to be built from. + # Full history, so the ancestry checks in release-v1.sh can see how + # this commit relates to v1. - uses: actions/checkout@v4 with: token: ${{ secrets.GITHUB_TOKEN }} fetch-depth: 0 - # `needs: selftest` only gates THIS run's own commit -- it says nothing - # about whether `origin/main` has since moved on to a merge whose own - # selftest hasn't finished, is still queued, or is red. Two things - # follow, checked in this order: - # - # 1. If `origin/main`'s live tip (re-fetched here, not trusted from the - # checkout above, which can be minutes stale behind this job's own - # selftest) is no longer THIS run's own `github.sha`, some other - # merge has landed since. Pushing it would release a commit this run - # never gated -- so this run defers instead, unconditionally. The - # commit that IS the live tip has its own run, and that run's own - # guard is what releases it once ITS turn to push comes -- possibly - # only once merges pause for a moment, so v1 can land one merge - # later than the newest one in a busy stretch. It never lands on a - # commit that wasn't gated, and it always catches up once things go - # quiet, because every future merge re-attempts the same check - # against whatever is current by then. - # 2. Only once (1) confirms this IS the live tip does it matter whether - # v1 already covers it -- a duplicate run, or a manual push already - # having done this. `--is-ancestor` treats a commit as its own - # ancestor, so "already at" and "already ahead" are one case. A v1 - # that doesn't exist yet, or shares no history with this commit, - # falls through to the push -- both checks are only ever a reason to - # skip, never a reason to fail. - - name: Determine whether this commit is still current and needs releasing - id: check - run: | - git fetch origin main - TIP=$(git rev-parse origin/main) - SHA="${{ github.sha }}" - if [ "$TIP" != "$SHA" ]; then - echo "origin/main's tip ($TIP) has moved past this run's own gated commit ($SHA) -- deferring to whichever run's own commit is now the live tip" - echo "skip=true" >> "$GITHUB_OUTPUT" - elif git rev-parse -q --verify refs/tags/v1 >/dev/null \ - && git merge-base --is-ancestor "$SHA" refs/tags/v1; then - echo "v1 already at or ahead of $SHA -- nothing to do" - echo "skip=true" >> "$GITHUB_OUTPUT" - else - echo "skip=false" >> "$GITHUB_OUTPUT" - fi - - # Lightweight tag, matching what v1 already is (`git cat-file -t v1` - # reports `commit`, not `tag`) -- no identity needed to move it, only - # to push it. - - name: Force v1 to this commit - if: steps.check.outputs.skip != 'true' - run: | - git tag -f v1 "${{ github.sha }}" - git push --force origin v1 + # Releases this run's own commit only while it is still main's tip, and + # only forward -- see release-v1.sh. + - name: Move v1 to this commit if it is still main's tip + run: bash scripts/release-v1.sh merge "${{ github.sha }}" diff --git a/.gitea/workflows/release-sweep.yaml b/.gitea/workflows/release-sweep.yaml new file mode 100644 index 0000000..589222e --- /dev/null +++ b/.gitea/workflows/release-sweep.yaml @@ -0,0 +1,69 @@ +name: Release sweep + +# Keeps v1 from lagging main when ci.yaml's release-tag job defers or never +# runs. Each tick either finds v1 already covering main's tip and exits, or +# gates the tip exactly as ci.yaml's selftest job does and moves v1 to it. A +# red run here means v1 is behind a main that fails its gate. +# +# Gitea registers schedules from the default branch only, so this fires once +# it is on main. +on: + schedule: + - cron: '*/15 * * * *' + +# A tick that arrives while another is still gating waits behind it rather +# than gating the same tip twice. +concurrency: + group: release-sweep + cancel-in-progress: false + +jobs: + sweep: + name: move v1 to main if it lags + runs-on: ubuntu-latest + timeout-minutes: 25 + permissions: + contents: write + steps: + - uses: actions/checkout@v4 + with: + token: ${{ secrets.GITHUB_TOKEN }} + fetch-depth: 0 + + # A scheduled run's github.sha is main as of the last push, not + # necessarily its tip, so the tip is read here instead. + - name: Check whether v1 lags main + id: check + run: bash scripts/release-v1.sh sweep-check + + # Everything below runs only when v1 lags, and gates the tip itself, + # not the commit this run was created from. + - name: Check out main's tip + if: steps.check.outputs.needed == 'true' + run: git checkout -q --detach "${{ steps.check.outputs.tip }}" + + - name: Install shellcheck + if: steps.check.outputs.needed == 'true' + uses: taiki-e/install-action@v2 + with: + tool: shellcheck + + # Same toolchains, same order, as ci.yaml's selftest job. + - name: Install Rust nightly + if: steps.check.outputs.needed == 'true' + uses: dtolnay/rust-toolchain@nightly + - name: Install Rust toolchain + if: steps.check.outputs.needed == 'true' + uses: dtolnay/rust-toolchain@stable + + - name: shellcheck + if: steps.check.outputs.needed == 'true' + run: shellcheck -x --source-path=scripts scripts/*.sh + + - name: Selftests + if: steps.check.outputs.needed == 'true' + run: bash scripts/selftest.sh + + - name: Move v1 to the gated tip + if: steps.check.outputs.needed == 'true' + run: bash scripts/release-v1.sh push "${{ steps.check.outputs.tip }}" "${{ steps.check.outputs.v1 }}" diff --git a/README.md b/README.md index c9f646f..b21aa46 100644 --- a/README.md +++ b/README.md @@ -634,34 +634,32 @@ entry, another permission — is a `v2`, not a `v1` move. Everything else moves `v1`: correctness fixes, new optional inputs, and anything internal to `scripts/`. -**Moving the tag is automatic, gated on the same build that gates a PR.** A -`release-tag` job in `.gitea/workflows/ci.yaml` runs on every push to `main`, -`needs: selftest`, and force-moves `v1` to *that run's own commit* once -selftest succeeds — a broken build never reaches it, so `v1` can't advance -onto one. It pushes with the run's built-in `GITHUB_TOKEN`; if that token -turns out not to have write access, the push step fails and the job goes red -in the Actions UI. That's a loud failure, not the silent one this replaced: -`v1` stays put, and nobody has to notice on their own that it lagged. +**Moving the tag is automatic, gated on the same build that gates a PR.** Two +jobs move it, both through `scripts/release-v1.sh`, both with the run's +built-in `GITHUB_TOKEN`: -Before pushing, the job checks two things and pushes only if both hold: that -`origin/main`'s live tip is still this run's own commit (not some later merge -that landed while this job was queued behind its own selftest), and that `v1` -doesn't already point at that commit or a descendant of it. Either check can -skip the push, as a normal, successful outcome — a run whose log says -"nothing to do" did its job correctly. The first check is what a job-level -`concurrency` group alone can't guarantee: Gitea's wake-one/cancel-rest -handling applies no ordering by commit recency, so an older merge's job can -be the one that survives to run — that job now defers instead of releasing a -commit it never gated. The newer merge's own job releases it once its own -turn comes, which can land `v1` one merge behind the newest during a busy -stretch; it always catches up once merges pause, because every later run -re-checks against whatever is current by then. +- **`release-tag`** in `.gitea/workflows/ci.yaml` runs on every push to + `main`, `needs: selftest`, and moves `v1` to that run's own commit — but only + while that commit is still `main`'s tip. A run whose merge has already been + overtaken defers, as a successful no-op, rather than release a commit it + never gated. +- **`release-sweep.yaml`** runs every 15 minutes. When `v1` already points at + `main`'s tip or a descendant of it, it exits after a checkout and one + comparison. Otherwise it runs the same shellcheck and selftests against the + tip and moves `v1` there only if they pass. + +So `v1` trails a green `main` by at most about one sweep interval plus one +selftest run, and **a `main` that fails its gate shows up as a red sweep on every +tick until it is fixed** — as does a push the token is not allowed to +make. Both jobs push with `--force-with-lease` on the `v1` they read, so +neither can move `v1` backwards over the other; a job that loses the lease to +a newer `v1` finishes green. This used to be a manual step, treated as a deliberate release decision taken once, knowingly, after the merge — in practice it was still forgotten (gitdan-actions#27): PR #25 merged to `main` and `v1` stayed on the previous release until someone asked whether it had moved. The manual form below is -still the recovery path, for when the automated job can't push: +still the recovery path, for when neither job can push: ```bash git fetch origin @@ -745,6 +743,7 @@ change here reaches all of them at once. That is what the gate is for. | `seed-target-dir-selftest.sh` | seed-source preference, lock-file stripping, two jobs racing on one cache key, **and one scenario per check a hardlink clone is validated against**: a source rotated wholesale, a subtree silently lost from the walk, a copy that reports failure over a tree both other checks read as whole, and a source identity that resolved at neither end — plus a staging tree that could not be privately owned being discarded rather than published, and the publisher's log showing it waited on the consumer's own reader-lock marker before reclaiming a rotated snapshot | | `publish-snapshot-selftest.sh` | the atomic swap, that a live consumer survives a republish, and the publisher's side of the rotation race: deferred reclamation under a live reader, and its sweep once the reader is gone | | `prune-cache-selftest.sh` | liveness in both its forms — a branch deleted from origin, and one still on it whose tip is already merged — plus protection, locking, eviction order, self-clear, **that a cache a job claims *inside* the check-to-unlink window survives it**, and that a requirement derived from the clone's mutable set evicts exactly enough and then fails rather than under-delivering. Against a real scratch `origin`, including a genuinely shallow clone of it and a `df` that answers from the fixture's own size, since a fixed one cannot show a pass stopping | +| `release-v1-selftest.sh` | that `v1` reaches `main`'s tip only through a gate and never moves backwards: the sweep's no-op, tag and red-gate cases, the stranded-defer trace the sweep exists to recover, and each lost-lease outcome — a newer `v1` skipped cleanly (with a control showing an unleased push steps it back), an older one retried, an unrelated one and a server rejection red. Against a real scratch `origin`; the other writer is sequenced between check and push, not raced | | `restore-mtimes-selftest.sh` | the merge hazard and the watermark that closes it, including the two-jobs-one-namespace case. Needs a real compiler. | Every suite runs the actual script, not a reimplementation of its logic, and diff --git a/scripts/release-v1-selftest.sh b/scripts/release-v1-selftest.sh new file mode 100644 index 0000000..5e85721 --- /dev/null +++ b/scripts/release-v1-selftest.sh @@ -0,0 +1,165 @@ +#!/usr/bin/env bash +# Regression test for release-v1.sh against a scratch bare origin: v1 reaches +# main's tip once it has been gated, never lands on an ungated commit, and +# never moves backwards when two writers race (gitdan-actions#27). +# +# run_sweep mirrors release-sweep.yaml's step order -- check, gate only when +# needed, push the gated tip leased on the v1 the check read -- with the gate +# stood in for by a command, so a failing gate is a failing sweep. +set -euo pipefail +script_dir=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd) +release="$script_dir/release-v1.sh" + +scratch=$(mktemp -d) +trap 'rm -rf "$scratch"' EXIT +pass_count=0 +fail() { echo "ASSERTION FAILED: $*" >&2; exit 1; } +ok() { pass_count=$((pass_count + 1)); echo "PASS: $*"; } + +export GIT_AUTHOR_NAME=t GIT_AUTHOR_EMAIL=t@t GIT_COMMITTER_NAME=t GIT_COMMITTER_EMAIL=t@t +unset GITHUB_OUTPUT + +# A fresh origin with main at one commit and v1 on it; dev pushes to main, +# ci and ci2 are the runners' clones. +fresh() { + rm -rf "$scratch/w"; mkdir -p "$scratch/w" + git init -q --bare "$scratch/w/origin.git" + git clone -q "$scratch/w/origin.git" "$scratch/w/dev" 2>/dev/null + commit_to_main >/dev/null + git -C "$scratch/w/dev" push -q origin HEAD:refs/tags/v1 + git clone -q "$scratch/w/origin.git" "$scratch/w/ci" + git clone -q "$scratch/w/origin.git" "$scratch/w/ci2" +} +commit_to_main() { + git -C "$scratch/w/dev" commit -q --allow-empty -m "c$RANDOM" + git -C "$scratch/w/dev" push -q origin HEAD:refs/heads/main + git -C "$scratch/w/dev" rev-parse HEAD +} +origin_v1() { git -C "$scratch/w/origin.git" rev-parse -q --verify 'refs/tags/v1^{commit}' || true; } +origin_tip() { git -C "$scratch/w/origin.git" rev-parse refs/heads/main; } +in_ci() { (cd "$scratch/w/${CLONE:-ci}" && bash "$release" "$@"); } +field() { sed -n "s/^$1=//p"; } + +run_sweep() { + local gate="$1" out tip v1 + out=$(in_ci sweep-check) + [ "$(field needed <<<"$out")" = true ] || return 0 + tip=$(field tip <<<"$out"); v1=$(field v1 <<<"$out") + "$gate" || return 1 + in_ci push "$tip" "$v1" +} + +echo "=== 1. sweep with v1 at the tip is a no-op ===" +fresh +before=$(origin_v1) +out=$(in_ci sweep-check) +[ "$(field needed <<<"$out")" = false ] || fail "a current v1 was reported as lagging" +run_sweep false || fail "a current v1 ran the gate" +[ "$(origin_v1)" = "$before" ] || fail "a no-op sweep moved v1" +ok "v1 == tip: needed=false, gate not run, v1 unchanged" + +echo +echo "=== 2. sweep with v1 ahead of the tip is a no-op ===" +fresh +ahead=$(git -C "$scratch/w/dev" commit-tree -p HEAD -m ahead 'HEAD^{tree}') +git -C "$scratch/w/dev" push -q -f origin "$ahead:refs/tags/v1" +out=$(in_ci sweep-check) +[ "$(field needed <<<"$out")" = false ] || fail "a v1 descending from the tip was reported as lagging" +ok "v1 descends from tip: needed=false" + +echo +echo "=== 3. sweep with v1 behind and a passing gate tags the tip ===" +fresh +tip=$(commit_to_main) +run_sweep true || fail "a passing sweep failed" +[ "$(origin_v1)" = "$tip" ] || fail "v1 is $(origin_v1), not the gated tip $tip" +ok "v1 behind, gate green: v1 -> tip" + +echo +echo "=== 4. sweep with v1 behind and a failing gate goes red and tags nothing ===" +fresh +before=$(origin_v1) +commit_to_main >/dev/null +if run_sweep false; then fail "a sweep over a failing gate succeeded"; fi +[ "$(origin_v1)" = "$before" ] || fail "a failing gate still moved v1" +ok "v1 behind, gate red: sweep red, v1 unchanged" + +echo +echo "=== 5. a lost lease to a newer writer is a clean skip, never a step back ===" +fresh +t1=$(commit_to_main) +out=$(in_ci sweep-check); v1_read=$(field v1 <<<"$out") +t2=$(commit_to_main) +CLONE=ci2 in_ci merge "$t2" >/dev/null +[ "$(origin_v1)" = "$t2" ] || fail "the merge job did not release its own tip" +in_ci push "$t1" "$v1_read" || fail "a lease lost to a newer v1 went red" +[ "$(origin_v1)" = "$t2" ] || fail "v1 went backwards from $t2 to $(origin_v1)" +ok "older writer lost the lease: exit 0, v1 stays at the newer $t2" +git -C "$scratch/w/ci" push -q -f origin "$t1:refs/tags/v1" +[ "$(origin_v1)" = "$t1" ] || fail "control: an unleased push did not step v1 back" +ok "control: the same push without the lease steps v1 back to $t1" + +echo +echo "=== 6. a lease lost to an older writer retries and lands the newer commit ===" +fresh +t1=$(commit_to_main) +t2=$(commit_to_main) +out=$(in_ci sweep-check); v1_read=$(field v1 <<<"$out") +git -C "$scratch/w/dev" push -q -f origin "$t1:refs/tags/v1" +in_ci push "$t2" "$v1_read" || fail "a lease lost to an older v1 went red" +[ "$(origin_v1)" = "$t2" ] || fail "v1 is $(origin_v1), not $t2" +ok "v1 moved to an ancestor under us: retried, v1 -> $t2" + +echo +echo "=== 7. a lease lost to an unrelated commit goes red ===" +fresh +tip=$(commit_to_main) +out=$(in_ci sweep-check); v1_read=$(field v1 <<<"$out") +stray=$(git -C "$scratch/w/dev" commit-tree -m stray 'HEAD^{tree}') +git -C "$scratch/w/dev" push -q -f origin "$stray:refs/tags/v1" +if in_ci push "$tip" "$v1_read" 2>/dev/null; then fail "a v1 moved sideways was accepted"; fi +[ "$(origin_v1)" = "$stray" ] || fail "the stray v1 was overwritten" +ok "v1 moved to a commit neither ahead nor behind: red, v1 untouched" + +echo +echo "=== 8. a push rejected for another reason goes red ===" +fresh +tip=$(commit_to_main) +mkdir -p "$scratch/w/origin.git/hooks" +printf '#!/bin/sh\nexit 1\n' > "$scratch/w/origin.git/hooks/pre-receive" +chmod +x "$scratch/w/origin.git/hooks/pre-receive" +before=$(origin_v1) +if in_ci merge "$tip" 2>"$scratch/err"; then fail "a rejected push reported success"; fi +[ "$(origin_v1)" = "$before" ] || fail "v1 moved despite the rejection" +grep -q 'not a lost lease' "$scratch/err" || fail "the rejection was not diagnosed as one: $(cat "$scratch/err")" +ok "server rejection with v1 unmoved: red, not a lost lease" + +echo +echo "=== 9. the merge job defers on a moved tip; the next sweep catches up ===" +# The stranded trace: C1's job runs after C2 merged and defers, C2's job was +# cancelled in the concurrency group, and merges stop. +fresh +before=$(origin_v1) +c1=$(commit_to_main) +c2=$(commit_to_main) +in_ci merge "$c1" >/dev/null || fail "the deferring merge job went red" +[ "$(origin_v1)" = "$before" ] || fail "the merge job released a commit that was not the tip" +run_sweep true || fail "the catch-up sweep failed" +[ "$(origin_v1)" = "$c2" ] || fail "v1 is $(origin_v1), not the tip $c2" +ok "C1 deferred, C2 never ran: the sweep moved v1 to $c2" + +echo +echo "=== 10. the merge job releases its own tip, and creates a missing v1 ===" +fresh +tip=$(commit_to_main) +in_ci merge "$tip" >/dev/null +[ "$(origin_v1)" = "$tip" ] || fail "the merge job did not release the tip" +git -C "$scratch/w/dev" push -q origin :refs/tags/v1 +tip=$(commit_to_main) +in_ci merge "$tip" >/dev/null +[ "$(origin_v1)" = "$tip" ] || fail "the merge job did not create an absent v1" +[ "$(origin_tip)" = "$tip" ] || fail "main moved" +ok "tip == gated sha: released, including onto an absent v1" + +echo +echo "release-v1-selftest: all $pass_count assertions passed" diff --git a/scripts/release-v1.sh b/scripts/release-v1.sh new file mode 100644 index 0000000..131edfb --- /dev/null +++ b/scripts/release-v1.sh @@ -0,0 +1,100 @@ +#!/usr/bin/env bash +# Moves the floating `v1` tag forward to a gated commit on `main`, and never +# backwards. Run from a clone whose `origin` is this repository. +# +# release-v1.sh merge merge-triggered job: release +# if it is still main's tip +# release-v1.sh sweep-check scheduled sweep: report whether v1 lags +# main (tip=, v1=, needed= to +# $GITHUB_OUTPUT, or stdout without one) +# release-v1.sh push +# scheduled sweep, after gating the tip +# +# Every push is leased on the v1 value the caller reasoned about. A lost lease +# means another writer moved v1 first: that is a clean skip once v1 is at or +# ahead of , a retry against the new value while v1 is still behind +# it, and a failure otherwise. +set -euo pipefail + +MAX_ATTEMPTS=3 + +fetch_main() { + git fetch -q origin +refs/heads/main:refs/remotes/origin/main + git rev-parse refs/remotes/origin/main +} + +# Prints origin's v1 commit, or nothing when origin has no v1. +fetch_v1() { + if [ -z "$(git ls-remote origin refs/tags/v1)" ]; then + git update-ref -d refs/release-v1/seen 2>/dev/null || true + return 0 + fi + git fetch -q origin +refs/tags/v1:refs/release-v1/seen + git rev-parse 'refs/release-v1/seen^{commit}' +} + +# True when v1 already covers : at it, or a descendant of it. +covers() { + local sha="$1" v1="$2" + [ -n "$v1" ] && git merge-base --is-ancestor "$sha" "$v1" +} + +push_leased() { + local sha="$1" expect="$2" now attempt + for ((attempt = 1; attempt <= MAX_ATTEMPTS; attempt++)); do + if git push -q --force-with-lease="refs/tags/v1:$expect" origin "$sha:refs/tags/v1"; then + echo "v1 moved ${expect:-} -> $sha" + return 0 + fi + now=$(fetch_v1) + if [ "$now" = "$expect" ]; then + echo "ERROR: push of v1 -> $sha rejected while v1 was still ${expect:-} -- not a lost lease" >&2 + return 1 + fi + if covers "$sha" "$now"; then + echo "lost the lease: another writer moved v1 to $now, at or ahead of $sha -- nothing to do" + return 0 + fi + if [ -n "$now" ] && ! git merge-base --is-ancestor "$now" "$sha"; then + echo "ERROR: v1 moved to $now, which is neither behind nor ahead of $sha" >&2 + return 1 + fi + echo "lost the lease: v1 moved to ${now:-}, still behind $sha -- retrying" + expect="$now" + done + echo "ERROR: lost the lease on v1 $MAX_ATTEMPTS times running" >&2 + return 1 +} + +cmd="${1:?usage: release-v1.sh merge | sweep-check | push }" +shift +case "$cmd" in + merge) + SHA="${1:?usage: release-v1.sh merge }" + TIP=$(fetch_main) + if [ "$TIP" != "$SHA" ]; then + echo "main's tip ($TIP) is past this run's gated commit ($SHA) -- deferring; the sweep releases the tip" + exit 0 + fi + V1=$(fetch_v1) + if covers "$SHA" "$V1"; then + echo "v1 ($V1) already at or ahead of $SHA -- nothing to do" + exit 0 + fi + push_leased "$SHA" "$V1" + ;; + sweep-check) + TIP=$(fetch_main) + V1=$(fetch_v1) + if covers "$TIP" "$V1"; then NEEDED=false; else NEEDED=true; fi + echo "main=$TIP v1=${V1:-} release-needed=$NEEDED" + printf 'tip=%s\nv1=%s\nneeded=%s\n' "$TIP" "$V1" "$NEEDED" >> "${GITHUB_OUTPUT:-/dev/stdout}" + ;; + push) + push_leased "${1:?usage: release-v1.sh push }" "${2-}" + ;; + *) + echo "release-v1.sh: unknown command '$cmd'" >&2 + exit 2 + ;; +esac diff --git a/scripts/selftest.sh b/scripts/selftest.sh index 0c7168a..a215799 100755 --- a/scripts/selftest.sh +++ b/scripts/selftest.sh @@ -13,7 +13,7 @@ script_dir=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd) FAST=0 [ "${1:-}" = "--fast" ] && FAST=1 -FIXTURE_TESTS=(cache-root-selftest.sh seed-target-dir-selftest.sh publish-snapshot-selftest.sh prune-cache-selftest.sh) +FIXTURE_TESTS=(cache-root-selftest.sh seed-target-dir-selftest.sh publish-snapshot-selftest.sh prune-cache-selftest.sh release-v1-selftest.sh) CARGO_TESTS=(hardlink-clone-selftest.sh restore-mtimes-selftest.sh) TESTS=("${FIXTURE_TESTS[@]}") From 77cc5917b6de7d350eb18e798b9b3ff1ab836345 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 22 Sep 2026 16:14:41 -0500 Subject: [PATCH 8/8] fix(release): refuse to overwrite an unrelated hand-placed v1 release-v1.sh's push_leased() only detected a lost lease after a push was *rejected* -- but force-with-lease compares the remote ref's raw value against the caller's expected value, not ancestry. If v1 already sat on a hand-placed, unrelated commit when push_leased() was first called (no race, nobody moves it mid-call), the very first push found the ref exactly where it expected, succeeded outright, and silently overwrote the unrelated v1 with -- skipping every ancestry check the function has, since those only run after a rejection. Fix: before the first push attempt, check whether the caller's `expect` is neither an ancestor of `sha` (the ordinary stale-v1 case) nor already covering it (nothing to do) -- and go red naming both SHAs if so. `expect` is always a peeled commit (fetch_v1() reads `refs/tags/v1^{commit}`), so this doesn't add a second failure mode for an annotated v1; that tag form's existing "not a lost lease" behavior on the first rejected push is untouched. Surfaced by PR #28's final review. New selftest scenario 11 in release-v1-selftest.sh, red-proven against the unfixed script (v1 was silently moved off the stray commit); green after the fix, with the full 7-suite gate (shellcheck + selftest.sh) passing. Ride-alongs from the same review: - ci.yaml:120-123 claimed the README's Versioning section documented what's verified about the release token's write access; it said nothing. Added an accurate sentence there (the grant is unobserved until the first merge, capped by repo/owner token-permission maxima and unreadable tag protections) and pointed the comment at it. - Deleted two comment-as-decision-history paragraphs per comments-are-not-exposition: ci.yaml's "no job-level concurrency" rationale (kept one line of intent) and prune-cache-selftest.sh scenario 9's account of how an existence-only assertion used to pass with the pass-2 guard removed (kept a one-line statement of what it checks). - README's release-v1-selftest.sh table row now names the new hand-placed-v1 scenario. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb --- .gitea/workflows/ci.yaml | 17 ++++++++--------- README.md | 9 ++++++++- scripts/prune-cache-selftest.sh | 12 +++--------- scripts/release-v1-selftest.sh | 15 +++++++++++++++ scripts/release-v1.sh | 9 +++++++++ 5 files changed, 43 insertions(+), 19 deletions(-) diff --git a/.gitea/workflows/ci.yaml b/.gitea/workflows/ci.yaml index 886e515..ff8a747 100644 --- a/.gitea/workflows/ci.yaml +++ b/.gitea/workflows/ci.yaml @@ -108,19 +108,18 @@ jobs: # selftest skips this job, so v1 never advances onto a broken build. The # `if:` restricts it to an actual push to main. # - # No job-level `concurrency:`. The lease in release-v1.sh already keeps v1 - # from moving backwards, and a group here only cancelled queued jobs in - # whatever order their selftests finished -- which could leave no job to - # release the newest merge. Anything this job defers or misses, - # release-sweep.yaml picks up. + # No job-level `concurrency:` -- the lease in release-v1.sh already keeps + # v1 from moving backwards, and release-sweep.yaml picks up anything this + # job defers or misses. needs: selftest if: ${{ github.event_name == 'push' && github.ref == 'refs/heads/main' }} runs-on: ubuntu-latest timeout-minutes: 2 - # Requests write access from the run's built-in token (see README's - # Versioning section for what's actually verified about it). Without - # this the checkout below still succeeds -- it's the push that would be - # rejected, which is a red job, not a silent no-op. + # Requests write access from the run's built-in token -- whether that + # grant actually lets it push here is unobserved until the first merge + # (see README's Versioning section). Without this the checkout below + # still succeeds -- it's the push that would be rejected, which is a red + # job, not a silent no-op. permissions: contents: write steps: diff --git a/README.md b/README.md index b21aa46..e34e5bc 100644 --- a/README.md +++ b/README.md @@ -648,6 +648,13 @@ built-in `GITHUB_TOKEN`: comparison. Otherwise it runs the same shellcheck and selftests against the tip and moves `v1` there only if they pass. +Both jobs request `contents: write` on the run's built-in token, but whether +that actually grants a push to this repo is **unobserved until the first +merge** — the grant is capped by the repo's and owner's maximum token +permissions, and branch/tag protections on `v1` can't be read without admin +access. A rejected push is a red job, not a silent no-op, so the first merge +after this lands is the real test. + So `v1` trails a green `main` by at most about one sweep interval plus one selftest run, and **a `main` that fails its gate shows up as a red sweep on every tick until it is fixed** — as does a push the token is not allowed to @@ -743,7 +750,7 @@ change here reaches all of them at once. That is what the gate is for. | `seed-target-dir-selftest.sh` | seed-source preference, lock-file stripping, two jobs racing on one cache key, **and one scenario per check a hardlink clone is validated against**: a source rotated wholesale, a subtree silently lost from the walk, a copy that reports failure over a tree both other checks read as whole, and a source identity that resolved at neither end — plus a staging tree that could not be privately owned being discarded rather than published, and the publisher's log showing it waited on the consumer's own reader-lock marker before reclaiming a rotated snapshot | | `publish-snapshot-selftest.sh` | the atomic swap, that a live consumer survives a republish, and the publisher's side of the rotation race: deferred reclamation under a live reader, and its sweep once the reader is gone | | `prune-cache-selftest.sh` | liveness in both its forms — a branch deleted from origin, and one still on it whose tip is already merged — plus protection, locking, eviction order, self-clear, **that a cache a job claims *inside* the check-to-unlink window survives it**, and that a requirement derived from the clone's mutable set evicts exactly enough and then fails rather than under-delivering. Against a real scratch `origin`, including a genuinely shallow clone of it and a `df` that answers from the fixture's own size, since a fixed one cannot show a pass stopping | -| `release-v1-selftest.sh` | that `v1` reaches `main`'s tip only through a gate and never moves backwards: the sweep's no-op, tag and red-gate cases, the stranded-defer trace the sweep exists to recover, and each lost-lease outcome — a newer `v1` skipped cleanly (with a control showing an unleased push steps it back), an older one retried, an unrelated one and a server rejection red. Against a real scratch `origin`; the other writer is sequenced between check and push, not raced | +| `release-v1-selftest.sh` | that `v1` reaches `main`'s tip only through a gate and never moves backwards: the sweep's no-op, tag and red-gate cases, the stranded-defer trace the sweep exists to recover, each lost-lease outcome — a newer `v1` skipped cleanly (with a control showing an unleased push steps it back), an older one retried, an unrelated one and a server rejection red — and a `v1` hand-placed on an unrelated commit before any push is ever attempted, also red. Against a real scratch `origin`; the other writer is sequenced between check and push, not raced | | `restore-mtimes-selftest.sh` | the merge hazard and the watermark that closes it, including the two-jobs-one-namespace case. Needs a real compiler. | Every suite runs the actual script, not a reimplementation of its logic, and diff --git a/scripts/prune-cache-selftest.sh b/scripts/prune-cache-selftest.sh index bf90c0e..e845746 100755 --- a/scripts/prune-cache-selftest.sh +++ b/scripts/prune-cache-selftest.sh @@ -34,15 +34,9 @@ # warm start. # 8. SELF-CLEAR REPORTS LOUDLY to the job summary, not just a log warning. # 9. OWN CACHE NEVER EVICTED by a sibling pass, genuinely under pressure — -# against a real, shrinking `df` (gitdan-actions#26): the static -# CACHE_DF_OVERRIDE every other scenario uses never reflects an -# eviction, so pass 3's self-clear (`rm -rf "$OWN_DIR"; mkdir -p -# "$OWN_DIR"`) fires regardless and recreates an empty OWN_DIR whether -# pass 2 touched it or not — existence survives either way, which is -# why an existence-only assertion here passed even with the pass-2 -# guard removed. This one checks CONTENTS, and sizes the requirement -# so it is satisfiable without self-clear at all: only pass 2's guard -# decides the outcome. +# against a real, shrinking `df` (gitdan-actions#26). Checks CONTENTS, +# not just existence, so pass 3's self-clear can't mask a missed +# pass-2 guard. # 10. SCOPED TO THE CACHE ROOT — a decoy outside it (standing in for another # project's volume) is never touched. # 11. A LIVE READER MARKER PROTECTS A CACHE the same way a lock file does — a diff --git a/scripts/release-v1-selftest.sh b/scripts/release-v1-selftest.sh index 5e85721..6e677ea 100644 --- a/scripts/release-v1-selftest.sh +++ b/scripts/release-v1-selftest.sh @@ -161,5 +161,20 @@ in_ci merge "$tip" >/dev/null [ "$(origin_tip)" = "$tip" ] || fail "main moved" ok "tip == gated sha: released, including onto an absent v1" +echo +echo "=== 11. a v1 hand-placed on an unrelated commit is never silently overwritten ===" +# Unlike #7, nothing races here -- v1 already sits on the stray commit before +# the very first push attempt, so force-with-lease sees exactly the value it +# expects and would otherwise succeed outright. +fresh +stray=$(git -C "$scratch/w/dev" commit-tree -m stray 'HEAD^{tree}') +git -C "$scratch/w/dev" push -q -f origin "$stray:refs/tags/v1" +tip=$(commit_to_main) +if in_ci merge "$tip" 2>"$scratch/err"; then fail "an unrelated hand-placed v1 was overwritten"; fi +[ "$(origin_v1)" = "$stray" ] || fail "v1 moved off the hand-placed $stray" +grep -q "$stray" "$scratch/err" || fail "the error did not name the stray v1: $(cat "$scratch/err")" +grep -q "$tip" "$scratch/err" || fail "the error did not name the gated sha: $(cat "$scratch/err")" +ok "hand-placed v1, unrelated to tip: red on the first push, v1 untouched" + echo echo "release-v1-selftest: all $pass_count assertions passed" diff --git a/scripts/release-v1.sh b/scripts/release-v1.sh index 131edfb..214d4e3 100644 --- a/scripts/release-v1.sh +++ b/scripts/release-v1.sh @@ -41,6 +41,15 @@ covers() { push_leased() { local sha="$1" expect="$2" now attempt + + # force-with-lease only compares the ref's current value, not ancestry, so + # an unrelated v1 -- neither behind nor covering it -- would + # otherwise be silently overwritten on the very first push. + if [ -n "$expect" ] && ! covers "$sha" "$expect" && ! git merge-base --is-ancestor "$expect" "$sha"; then + echo "ERROR: v1 ($expect) is neither an ancestor of $sha nor at/ahead of it -- refusing to overwrite an unrelated v1" >&2 + return 1 + fi + for ((attempt = 1; attempt <= MAX_ATTEMPTS; attempt++)); do if git push -q --force-with-lease="refs/tags/v1:$expect" origin "$sha:refs/tags/v1"; then echo "v1 moved ${expect:-} -> $sha"