fix(ci): target origin/main's live tip, not the run's own trigger commit
CI / shellcheck + selftests (pull_request) Successful in 1m35s
CI / move v1 to main (pull_request) Skipped

The wake-one-cancel-the-rest mechanism behind the concurrency group
(CancelPreviousJobsByJobConcurrency, models/actions/run_job.go, at
the v1.27.2 tag this instance runs) picks its survivor from
models/actions/run_job_list.go's query, which carries no `ORDER BY`
-- so under three-way contention on a shared 2-slot runner, the
*newest* commit's job can be the one cancelled while an older sibling
survives and, correctly from its own vantage, advances v1 forward
from a stale view. No push regresses v1 (the ancestor check from
dc1e631 already prevented that), but the newest merge goes silently
unreleased behind a cancelled job that reads as benign, not red --
exactly the failure #27 exists to end.

Every job now resolves `origin/main`'s tip fresh, right before the
push, instead of using `${{ github.sha }}`. Re-fetched explicitly
rather than trusted from the checkout step, which can be minutes
stale by this point behind its own selftest job. Every execution that
reaches the push step now converges on the same target regardless of
which job the concurrency group lets through, so which one wins the
wake no longer matters -- the survivor pushes where any of them
would have.

That doesn't make the push safe on its own: two jobs can still read
main at genuinely different moments if it advances between their two
fetches, so whichever read the tip earlier must not overwrite the
other's already-pushed, newer one. The ancestor check from dc1e631 is
kept for exactly this -- its target changed (origin/main's live tip,
not this job's own trigger commit) but its job didn't.

What each guard now protects against, after this change:
- concurrency group: stops two jobs from pushing at the same time --
  wasted work now that a cancelled job costs nothing, not a
  correctness backstop by itself.
- ancestor check: stops a job whose own fetch of the tip is stale
  relative to another job's already-pushed, fresher one from
  regressing v1.

Rewrote both the job-level comments and README's Versioning section,
which described "force-moves v1 to that commit" and said nothing
about the concurrency group or the skip-as-success case.

Re-derived the truth table against the new target (a live origin/main
tip, not a fixed commit) in a scratch origin+clone: no v1 yet (push),
v1 exactly at the tip (skip), the tip moved past v1 because a newer
merge landed (push, to the new tip -- not stuck on any prior commit),
v1 already ahead of the tip (skip, defensive), unrelated histories
(push, defensive). All five resolved as intended; ran the exact
condition and fetch sequence the workflow step uses, not a
simulation of it.

bash scripts/selftest.sh: all 6 suites green (67 prune-cache
assertions, unchanged). shellcheck -x --source-path=scripts
scripts/*.sh: clean -- as before, this does not cover the inline
`run:` shell in ci.yaml.

What remains unverifiable short of a real merge is unchanged from
dc1e631: the workflow step's actual execution inside a real Actions
run, and whether the built-in token has write access at all. Neither
this commit nor the one before it can exercise those.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
This commit is contained in:
2026-09-22 12:49:05 -05:00
co-authored by Claude Sonnet 5
parent dc1e6317c6
commit ea48c03aeb
2 changed files with 58 additions and 30 deletions
+17 -7
View File
@@ -636,18 +636,28 @@ entry, another permission — is a `v2`, not a `v1` move. Everything else moves
**Moving the tag is automatic, gated on the same build that gates a PR.** A
`release-tag` job in `.gitea/workflows/ci.yaml` runs on every push to `main`,
`needs: selftest`, and force-moves `v1` to that commit once selftest
succeeds — a broken build never reaches it, so `v1` can't advance onto one.
It pushes with the run's built-in `GITHUB_TOKEN`; if that token turns out not
to have write access, the push step fails and the job goes red in the Actions
UI. That's a loud failure, not the silent one this replaced: `v1` stays put,
and nobody has to notice on their own that it lagged.
`needs: selftest`, and force-moves `v1` to `origin/main`'s live tip once
selftest succeeds — a broken build never reaches it, so `v1` can't advance
onto one. It pushes with the run's built-in `GITHUB_TOKEN`; if that token
turns out not to have write access, the push step fails and the job goes red
in the Actions UI. That's a loud failure, not the silent one this replaced:
`v1` stays put, and nobody has to notice on their own that it lagged.
Two merges landing close together can start two `release-tag` jobs at once; a
job-level `concurrency` group lets only one push at a time, and every job
targets `origin/main`'s current tip rather than its own trigger commit, so it
makes no difference which one the group lets through. Before pushing, the job
also checks whether `v1` already points at that tip or a descendant of it —
covering the case where two jobs read the tip at genuinely different moments
— and skips as a normal, successful outcome rather than pushing backward. A
run whose log says "nothing to do" did its job correctly; it just found
nothing to move.
This used to be a manual step, treated as a deliberate release decision taken
once, knowingly, after the merge — in practice it was still forgotten
(gitdan-actions#27): PR #25 merged to `main` and `v1` stayed on the previous
release until someone asked whether it had moved. The manual form below is
now the recovery path, for when the automated job can't push:
still the recovery path, for when the automated job can't push:
```bash
git fetch origin