fix(ci): target origin/main's live tip, not the run's own trigger commit
The wake-one-cancel-the-rest mechanism behind the concurrency group (CancelPreviousJobsByJobConcurrency, models/actions/run_job.go, at the v1.27.2 tag this instance runs) picks its survivor from models/actions/run_job_list.go's query, which carries no `ORDER BY` -- so under three-way contention on a shared 2-slot runner, the *newest* commit's job can be the one cancelled while an older sibling survives and, correctly from its own vantage, advances v1 forward from a stale view. No push regresses v1 (the ancestor check fromdc1e631already prevented that), but the newest merge goes silently unreleased behind a cancelled job that reads as benign, not red -- exactly the failure #27 exists to end. Every job now resolves `origin/main`'s tip fresh, right before the push, instead of using `${{ github.sha }}`. Re-fetched explicitly rather than trusted from the checkout step, which can be minutes stale by this point behind its own selftest job. Every execution that reaches the push step now converges on the same target regardless of which job the concurrency group lets through, so which one wins the wake no longer matters -- the survivor pushes where any of them would have. That doesn't make the push safe on its own: two jobs can still read main at genuinely different moments if it advances between their two fetches, so whichever read the tip earlier must not overwrite the other's already-pushed, newer one. The ancestor check fromdc1e631is kept for exactly this -- its target changed (origin/main's live tip, not this job's own trigger commit) but its job didn't. What each guard now protects against, after this change: - concurrency group: stops two jobs from pushing at the same time -- wasted work now that a cancelled job costs nothing, not a correctness backstop by itself. - ancestor check: stops a job whose own fetch of the tip is stale relative to another job's already-pushed, fresher one from regressing v1. Rewrote both the job-level comments and README's Versioning section, which described "force-moves v1 to that commit" and said nothing about the concurrency group or the skip-as-success case. Re-derived the truth table against the new target (a live origin/main tip, not a fixed commit) in a scratch origin+clone: no v1 yet (push), v1 exactly at the tip (skip), the tip moved past v1 because a newer merge landed (push, to the new tip -- not stuck on any prior commit), v1 already ahead of the tip (skip, defensive), unrelated histories (push, defensive). All five resolved as intended; ran the exact condition and fetch sequence the workflow step uses, not a simulation of it. bash scripts/selftest.sh: all 6 suites green (67 prune-cache assertions, unchanged). shellcheck -x --source-path=scripts scripts/*.sh: clean -- as before, this does not cover the inline `run:` shell in ci.yaml. What remains unverifiable short of a real merge is unchanged fromdc1e631: the workflow step's actual execution inside a real Actions run, and whether the built-in token has write access at all. Neither this commit nor the one before it can exercise those. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
This commit is contained in:
@@ -636,18 +636,28 @@ entry, another permission — is a `v2`, not a `v1` move. Everything else moves
|
||||
|
||||
**Moving the tag is automatic, gated on the same build that gates a PR.** A
|
||||
`release-tag` job in `.gitea/workflows/ci.yaml` runs on every push to `main`,
|
||||
`needs: selftest`, and force-moves `v1` to that commit once selftest
|
||||
succeeds — a broken build never reaches it, so `v1` can't advance onto one.
|
||||
It pushes with the run's built-in `GITHUB_TOKEN`; if that token turns out not
|
||||
to have write access, the push step fails and the job goes red in the Actions
|
||||
UI. That's a loud failure, not the silent one this replaced: `v1` stays put,
|
||||
and nobody has to notice on their own that it lagged.
|
||||
`needs: selftest`, and force-moves `v1` to `origin/main`'s live tip once
|
||||
selftest succeeds — a broken build never reaches it, so `v1` can't advance
|
||||
onto one. It pushes with the run's built-in `GITHUB_TOKEN`; if that token
|
||||
turns out not to have write access, the push step fails and the job goes red
|
||||
in the Actions UI. That's a loud failure, not the silent one this replaced:
|
||||
`v1` stays put, and nobody has to notice on their own that it lagged.
|
||||
|
||||
Two merges landing close together can start two `release-tag` jobs at once; a
|
||||
job-level `concurrency` group lets only one push at a time, and every job
|
||||
targets `origin/main`'s current tip rather than its own trigger commit, so it
|
||||
makes no difference which one the group lets through. Before pushing, the job
|
||||
also checks whether `v1` already points at that tip or a descendant of it —
|
||||
covering the case where two jobs read the tip at genuinely different moments
|
||||
— and skips as a normal, successful outcome rather than pushing backward. A
|
||||
run whose log says "nothing to do" did its job correctly; it just found
|
||||
nothing to move.
|
||||
|
||||
This used to be a manual step, treated as a deliberate release decision taken
|
||||
once, knowingly, after the merge — in practice it was still forgotten
|
||||
(gitdan-actions#27): PR #25 merged to `main` and `v1` stayed on the previous
|
||||
release until someone asked whether it had moved. The manual form below is
|
||||
now the recovery path, for when the automated job can't push:
|
||||
still the recovery path, for when the automated job can't push:
|
||||
|
||||
```bash
|
||||
git fetch origin
|
||||
|
||||
Reference in New Issue
Block a user