fix(ci): only ever push a commit this run actually gated; fix vacuous scenario-9 guard

## v1 could advance onto an ungated commit

`needs: selftest` gates this run's own commit, but the push targeted
origin/main's freshly-fetched tip with nothing comparing the two.
Trace: M1 merges green; M2 merges while M1's selftest is still
running; M1's release-tag job fetches tip = M2 and pushes v1 = M2,
whose own selftest may be queued, running, or red. If M2 is red, its
own job is skipped, so v1 sits on a red commit across every consuming
project until the next green merge -- with nothing red pointing at
the release itself. ci.yaml:107-109 and README.md:640-641 both
asserted this couldn't happen; ea48c03's own diff established the
precondition (tip "can be minutes stale... behind its own selftest
job") without closing it.

Fix: skip the push unless origin/main's tip IS this run's own
github.sha, checked before the existing v1-monotonicity check
(dc1e631) rather than replacing it -- the two compose (tip-mismatch
first, since it's the coarser reason to defer; ancestor-check second,
for a duplicate run whose own commit is still current). Reverted the
push target from the fetched tip back to `${{ github.sha }}`, now
that the guard makes them provably equal whenever the push fires.

## Convergence trace: does the newest commit's job still run?

Read gitea's source further at the pinned v1.27.2 tag:
PrepareToStartJobWithConcurrency (services/actions/clear_tasks.go)
calls CancelPreviousJobsByJobConcurrency on every job entering the
group, unconditionally cancelling whatever was previously
Waiting/Blocked there -- so at most one job sits queued in the group
at a time; each new arrival supersedes it. Because job-level
concurrency is only evaluated once `needs: selftest` is satisfied
(job_emitter.go re-evaluates readiness there), "arrival order" tracks
each commit's own selftest-completion time, not raw merge order -- an
older commit with a slower selftest can enter the group after a
newer one and cancel its queued slot.

That cancelled job is gone for good; it will never push. But the
commit that's genuinely current at the moment merges stop arriving is
always the one still queued when the running job finishes, because
every subsequent arrival (from every subsequent merge, not just the
"newest" one at any single instant) keeps re-superseding the queue.
So v1 always eventually catches up -- "one merge later" in the common
case, "at the next merge, whenever that happens" in the adversarial
case where a stale survivor runs, finds itself no longer current,
defers, and nothing is left queued. It cannot get stuck forever short
of the repository never receiving another merge, because every future
push re-attempts the same check against whatever's current by then.
Stated this plainly in the comment and README rather than repeating
the false "never" guarantee in softer words.

Re-derived the truth table against the new guard in a scratch
origin+clone, six cases: own commit == tip, no v1 (push); v1 already
== own commit (skip, duplicate run); tip moved past own gated commit
because a newer merge landed (skip, defers); the newer commit's own
run once nothing further has landed (push); v1 already ahead of a
now-stale gated commit (skip, tip-mismatch catches it first); tip ==
own commit but v1 independently ahead via a local-only descendant,
isolating the second (ancestor) check on its own (skip). All six
resolved as intended.

## Scenario 9's pass-3 guard was vacuous

scripts/prune-cache-selftest.sh:230 grepped 'self-clear' against
$scratch/log. summary_line() (cache-lib.sh) writes only to
$GITHUB_STEP_SUMMARY; the self-clear branch (prune-cache.sh:632)
writes 'self-clear' there and 'clearing own' to stdout (:631) --
'self-clear' never appears in $scratch/log at all, so the branch was
unreachable and the `ok` unconditional. Scenario 8 already greps the
right string against the right file (assert_log "clearing own" ...);
scenario 9 now does the same, staying on $scratch/log where it
already was -- the string was wrong, not the file.

Red-proved by capturing a real prune-cache.sh log where self-clear
genuinely fired (scenario 9's own fixture with MIN_FREE_PCT
temporarily raised to 100, in a scratch copy, reverted after) and
running both patterns against it: `grep -q 'self-clear'` -> no match
(the old check's vacuous pass, confirmed); `grep -q 'clearing own'`
-> match (the fix's correct fail). Matches the reviewer's own
measurement exactly. The real prune-cache-selftest.sh was untouched
during this experiment; only the grep string changed in the actual
commit.

bash scripts/selftest.sh: all 6 suites green (67 prune-cache
assertions, unchanged in count -- the fix corrects what scenario 9's
existing check compares, not what it asserts). shellcheck -x
--source-path=scripts scripts/*.sh: clean, as before not covering the
inline ci.yaml shell.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
This commit is contained in:
2026-09-22 14:01:28 -05:00
co-authored by Claude Sonnet 5
parent ea48c03aeb
commit 17d87b0647
3 changed files with 54 additions and 46 deletions
+14 -10
View File
@@ -636,22 +636,26 @@ entry, another permission — is a `v2`, not a `v1` move. Everything else moves
**Moving the tag is automatic, gated on the same build that gates a PR.** A
`release-tag` job in `.gitea/workflows/ci.yaml` runs on every push to `main`,
`needs: selftest`, and force-moves `v1` to `origin/main`'s live tip once
`needs: selftest`, and force-moves `v1` to *that run's own commit* once
selftest succeeds — a broken build never reaches it, so `v1` can't advance
onto one. It pushes with the run's built-in `GITHUB_TOKEN`; if that token
turns out not to have write access, the push step fails and the job goes red
in the Actions UI. That's a loud failure, not the silent one this replaced:
`v1` stays put, and nobody has to notice on their own that it lagged.
Two merges landing close together can start two `release-tag` jobs at once; a
job-level `concurrency` group lets only one push at a time, and every job
targets `origin/main`'s current tip rather than its own trigger commit, so it
makes no difference which one the group lets through. Before pushing, the job
also checks whether `v1` already points at that tip or a descendant of it —
covering the case where two jobs read the tip at genuinely different moments
— and skips as a normal, successful outcome rather than pushing backward. A
run whose log says "nothing to do" did its job correctly; it just found
nothing to move.
Before pushing, the job checks two things and pushes only if both hold: that
`origin/main`'s live tip is still this run's own commit (not some later merge
that landed while this job was queued behind its own selftest), and that `v1`
doesn't already point at that commit or a descendant of it. Either check can
skip the push, as a normal, successful outcome — a run whose log says
"nothing to do" did its job correctly. The first check is what a job-level
`concurrency` group alone can't guarantee: Gitea's wake-one/cancel-rest
handling applies no ordering by commit recency, so an older merge's job can
be the one that survives to run — that job now defers instead of releasing a
commit it never gated. The newer merge's own job releases it once its own
turn comes, which can land `v1` one merge behind the newest during a busy
stretch; it always catches up once merges pause, because every later run
re-checks against whatever is current by then.
This used to be a manual step, treated as a deliberate release decision taken
once, knowingly, after the merge — in practice it was still forgotten