fix(ci): target origin/main's live tip, not the run's own trigger commit
The wake-one-cancel-the-rest mechanism behind the concurrency group (CancelPreviousJobsByJobConcurrency, models/actions/run_job.go, at the v1.27.2 tag this instance runs) picks its survivor from models/actions/run_job_list.go's query, which carries no `ORDER BY` -- so under three-way contention on a shared 2-slot runner, the *newest* commit's job can be the one cancelled while an older sibling survives and, correctly from its own vantage, advances v1 forward from a stale view. No push regresses v1 (the ancestor check fromdc1e631already prevented that), but the newest merge goes silently unreleased behind a cancelled job that reads as benign, not red -- exactly the failure #27 exists to end. Every job now resolves `origin/main`'s tip fresh, right before the push, instead of using `${{ github.sha }}`. Re-fetched explicitly rather than trusted from the checkout step, which can be minutes stale by this point behind its own selftest job. Every execution that reaches the push step now converges on the same target regardless of which job the concurrency group lets through, so which one wins the wake no longer matters -- the survivor pushes where any of them would have. That doesn't make the push safe on its own: two jobs can still read main at genuinely different moments if it advances between their two fetches, so whichever read the tip earlier must not overwrite the other's already-pushed, newer one. The ancestor check fromdc1e631is kept for exactly this -- its target changed (origin/main's live tip, not this job's own trigger commit) but its job didn't. What each guard now protects against, after this change: - concurrency group: stops two jobs from pushing at the same time -- wasted work now that a cancelled job costs nothing, not a correctness backstop by itself. - ancestor check: stops a job whose own fetch of the tip is stale relative to another job's already-pushed, fresher one from regressing v1. Rewrote both the job-level comments and README's Versioning section, which described "force-moves v1 to that commit" and said nothing about the concurrency group or the skip-as-success case. Re-derived the truth table against the new target (a live origin/main tip, not a fixed commit) in a scratch origin+clone: no v1 yet (push), v1 exactly at the tip (skip), the tip moved past v1 because a newer merge landed (push, to the new tip -- not stuck on any prior commit), v1 already ahead of the tip (skip, defensive), unrelated histories (push, defensive). All five resolved as intended; ran the exact condition and fetch sequence the workflow step uses, not a simulation of it. bash scripts/selftest.sh: all 6 suites green (67 prune-cache assertions, unchanged). shellcheck -x --source-path=scripts scripts/*.sh: clean -- as before, this does not cover the inline `run:` shell in ci.yaml. What remains unverifiable short of a real merge is unchanged fromdc1e631: the workflow step's actual execution inside a real Actions run, and whether the built-in token has write access at all. Neither this commit nor the one before it can exercise those. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
This commit is contained in:
+41
-23
@@ -123,13 +123,18 @@ jobs:
|
||||
# fields, evaluated and enforced by separate functions --
|
||||
# CancelPreviousJobsByRunConcurrency vs CancelPreviousJobsByJobConcurrency
|
||||
# in models/actions/{run,run_job}.go -- not one overriding the other).
|
||||
# This serialises jobs that are actually running or queued AT THE SAME
|
||||
# TIME -- it has no notion of commit order between jobs that never
|
||||
# overlap. Two release-tag runs on a 2-slot, 4-repo runner with a
|
||||
# multi-minute selftest ahead of each can easily not overlap at all,
|
||||
# finishing in either order regardless of push order. The step below is
|
||||
# what makes the OUTCOME order-independent; this group only bounds how
|
||||
# much work is wasted getting there.
|
||||
#
|
||||
# This does NOT decide which of several contending jobs gets to push --
|
||||
# Gitea wakes exactly one Blocked job in the group and cancels the rest
|
||||
# outright, with no ordering on which one it picks (no `ORDER BY` in the
|
||||
# query behind CancelPreviousJobsByJobConcurrency,
|
||||
# models/actions/run_job_list.go). What it buys is cheaper: every
|
||||
# execution that does reach the push step targets origin/main's live tip
|
||||
# (below), never its own trigger commit, so it makes no difference which
|
||||
# one wins -- the survivor pushes where any of them would have, and a
|
||||
# cancelled job costs nothing. This group's only job is to stop more than
|
||||
# one job from pushing AT THE SAME TIME, which is wasted work, not a
|
||||
# correctness risk on its own.
|
||||
concurrency:
|
||||
group: release-tag-v1
|
||||
cancel-in-progress: false
|
||||
@@ -150,24 +155,37 @@ jobs:
|
||||
token: ${{ secrets.GITHUB_TOKEN }}
|
||||
fetch-depth: 0
|
||||
|
||||
# The concurrency group above only excludes a SIMULTANEOUS competing
|
||||
# push; it does nothing for two release-tag jobs that never overlap
|
||||
# and finish in the opposite order from the merges that triggered
|
||||
# them -- which a shared 2-slot runner and a multi-minute selftest
|
||||
# ahead of this job make routine, not rare. So the push itself has to
|
||||
# be safe regardless of completion order: skip whenever v1 already
|
||||
# points at this commit or a descendant of it, rather than trusting
|
||||
# this run to be the last one to finish. `--is-ancestor` treats a
|
||||
# commit as its own ancestor, so "already at" and "already ahead" are
|
||||
# one case. A v1 that doesn't exist yet, or that shares no history
|
||||
# with this commit, falls through to the push below -- the guard is
|
||||
# This run's own trigger commit (${{ github.sha }}) is deliberately not
|
||||
# what gets pushed: whichever job the concurrency group above lets
|
||||
# through is the one that pushes, and that choice carries no relation
|
||||
# to commit recency, so every execution has to converge on the SAME
|
||||
# target regardless of which job it is. `origin/main`'s live tip,
|
||||
# re-fetched here rather than trusted from the checkout above (which
|
||||
# can be minutes stale by this point, behind its own selftest job), is
|
||||
# that common target -- read fresh, every job that reaches this step
|
||||
# resolves to the same commit whenever main hasn't moved between them,
|
||||
# and to whatever's newest when it has.
|
||||
#
|
||||
# A live target doesn't make the push itself safe on its own: two jobs
|
||||
# can still read main at genuinely different moments if it advances
|
||||
# between their two fetches, so the one with the earlier reading must
|
||||
# not overwrite the other's already-pushed, newer one. That's what the
|
||||
# ancestor check below still guards -- not "this job's stale trigger
|
||||
# commit" any more, but "this job's freshly-read tip, which another
|
||||
# job's fresher read may have already superseded." `--is-ancestor`
|
||||
# treats a commit as its own ancestor, so "already at" and "already
|
||||
# ahead" are one case. A v1 that doesn't exist yet, or that shares no
|
||||
# history with this tip, falls through to the push -- the guard is
|
||||
# only ever a reason to skip, never a reason to fail.
|
||||
- name: Skip if v1 already at or ahead of this commit
|
||||
- name: Determine origin/main's tip and whether v1 needs to move
|
||||
id: check
|
||||
run: |
|
||||
git fetch origin main
|
||||
TIP=$(git rev-parse origin/main)
|
||||
echo "tip=$TIP" >> "$GITHUB_OUTPUT"
|
||||
if git rev-parse -q --verify refs/tags/v1 >/dev/null \
|
||||
&& git merge-base --is-ancestor "${{ github.sha }}" refs/tags/v1; then
|
||||
echo "v1 already at or ahead of ${{ github.sha }} -- nothing to do"
|
||||
&& git merge-base --is-ancestor "$TIP" refs/tags/v1; then
|
||||
echo "v1 already at or ahead of origin/main's tip ($TIP) -- nothing to do"
|
||||
echo "skip=true" >> "$GITHUB_OUTPUT"
|
||||
else
|
||||
echo "skip=false" >> "$GITHUB_OUTPUT"
|
||||
@@ -176,8 +194,8 @@ jobs:
|
||||
# Lightweight tag, matching what v1 already is (`git cat-file -t v1`
|
||||
# reports `commit`, not `tag`) -- no identity needed to move it, only
|
||||
# to push it.
|
||||
- name: Force v1 to this commit
|
||||
- name: Force v1 to origin/main's tip
|
||||
if: steps.check.outputs.skip != 'true'
|
||||
run: |
|
||||
git tag -f v1 "${{ github.sha }}"
|
||||
git tag -f v1 "${{ steps.check.outputs.tip }}"
|
||||
git push --force origin v1
|
||||
|
||||
@@ -636,18 +636,28 @@ entry, another permission — is a `v2`, not a `v1` move. Everything else moves
|
||||
|
||||
**Moving the tag is automatic, gated on the same build that gates a PR.** A
|
||||
`release-tag` job in `.gitea/workflows/ci.yaml` runs on every push to `main`,
|
||||
`needs: selftest`, and force-moves `v1` to that commit once selftest
|
||||
succeeds — a broken build never reaches it, so `v1` can't advance onto one.
|
||||
It pushes with the run's built-in `GITHUB_TOKEN`; if that token turns out not
|
||||
to have write access, the push step fails and the job goes red in the Actions
|
||||
UI. That's a loud failure, not the silent one this replaced: `v1` stays put,
|
||||
and nobody has to notice on their own that it lagged.
|
||||
`needs: selftest`, and force-moves `v1` to `origin/main`'s live tip once
|
||||
selftest succeeds — a broken build never reaches it, so `v1` can't advance
|
||||
onto one. It pushes with the run's built-in `GITHUB_TOKEN`; if that token
|
||||
turns out not to have write access, the push step fails and the job goes red
|
||||
in the Actions UI. That's a loud failure, not the silent one this replaced:
|
||||
`v1` stays put, and nobody has to notice on their own that it lagged.
|
||||
|
||||
Two merges landing close together can start two `release-tag` jobs at once; a
|
||||
job-level `concurrency` group lets only one push at a time, and every job
|
||||
targets `origin/main`'s current tip rather than its own trigger commit, so it
|
||||
makes no difference which one the group lets through. Before pushing, the job
|
||||
also checks whether `v1` already points at that tip or a descendant of it —
|
||||
covering the case where two jobs read the tip at genuinely different moments
|
||||
— and skips as a normal, successful outcome rather than pushing backward. A
|
||||
run whose log says "nothing to do" did its job correctly; it just found
|
||||
nothing to move.
|
||||
|
||||
This used to be a manual step, treated as a deliberate release decision taken
|
||||
once, knowingly, after the merge — in practice it was still forgotten
|
||||
(gitdan-actions#27): PR #25 merged to `main` and `v1` stayed on the previous
|
||||
release until someone asked whether it had moved. The manual form below is
|
||||
now the recovery path, for when the automated job can't push:
|
||||
still the recovery path, for when the automated job can't push:
|
||||
|
||||
```bash
|
||||
git fetch origin
|
||||
|
||||
Reference in New Issue
Block a user