fix(ci): target origin/main's live tip, not the run's own trigger commit
CI / shellcheck + selftests (pull_request) Successful in 1m35s
CI / move v1 to main (pull_request) Skipped

The wake-one-cancel-the-rest mechanism behind the concurrency group
(CancelPreviousJobsByJobConcurrency, models/actions/run_job.go, at
the v1.27.2 tag this instance runs) picks its survivor from
models/actions/run_job_list.go's query, which carries no `ORDER BY`
-- so under three-way contention on a shared 2-slot runner, the
*newest* commit's job can be the one cancelled while an older sibling
survives and, correctly from its own vantage, advances v1 forward
from a stale view. No push regresses v1 (the ancestor check from
dc1e631 already prevented that), but the newest merge goes silently
unreleased behind a cancelled job that reads as benign, not red --
exactly the failure #27 exists to end.

Every job now resolves `origin/main`'s tip fresh, right before the
push, instead of using `${{ github.sha }}`. Re-fetched explicitly
rather than trusted from the checkout step, which can be minutes
stale by this point behind its own selftest job. Every execution that
reaches the push step now converges on the same target regardless of
which job the concurrency group lets through, so which one wins the
wake no longer matters -- the survivor pushes where any of them
would have.

That doesn't make the push safe on its own: two jobs can still read
main at genuinely different moments if it advances between their two
fetches, so whichever read the tip earlier must not overwrite the
other's already-pushed, newer one. The ancestor check from dc1e631 is
kept for exactly this -- its target changed (origin/main's live tip,
not this job's own trigger commit) but its job didn't.

What each guard now protects against, after this change:
- concurrency group: stops two jobs from pushing at the same time --
  wasted work now that a cancelled job costs nothing, not a
  correctness backstop by itself.
- ancestor check: stops a job whose own fetch of the tip is stale
  relative to another job's already-pushed, fresher one from
  regressing v1.

Rewrote both the job-level comments and README's Versioning section,
which described "force-moves v1 to that commit" and said nothing
about the concurrency group or the skip-as-success case.

Re-derived the truth table against the new target (a live origin/main
tip, not a fixed commit) in a scratch origin+clone: no v1 yet (push),
v1 exactly at the tip (skip), the tip moved past v1 because a newer
merge landed (push, to the new tip -- not stuck on any prior commit),
v1 already ahead of the tip (skip, defensive), unrelated histories
(push, defensive). All five resolved as intended; ran the exact
condition and fetch sequence the workflow step uses, not a
simulation of it.

bash scripts/selftest.sh: all 6 suites green (67 prune-cache
assertions, unchanged). shellcheck -x --source-path=scripts
scripts/*.sh: clean -- as before, this does not cover the inline
`run:` shell in ci.yaml.

What remains unverifiable short of a real merge is unchanged from
dc1e631: the workflow step's actual execution inside a real Actions
run, and whether the built-in token has write access at all. Neither
this commit nor the one before it can exercise those.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
This commit is contained in:
2026-09-22 12:49:05 -05:00
co-authored by Claude Sonnet 5
parent dc1e6317c6
commit ea48c03aeb
2 changed files with 58 additions and 30 deletions
+41 -23
View File
@@ -123,13 +123,18 @@ jobs:
# fields, evaluated and enforced by separate functions -- # fields, evaluated and enforced by separate functions --
# CancelPreviousJobsByRunConcurrency vs CancelPreviousJobsByJobConcurrency # CancelPreviousJobsByRunConcurrency vs CancelPreviousJobsByJobConcurrency
# in models/actions/{run,run_job}.go -- not one overriding the other). # in models/actions/{run,run_job}.go -- not one overriding the other).
# This serialises jobs that are actually running or queued AT THE SAME #
# TIME -- it has no notion of commit order between jobs that never # This does NOT decide which of several contending jobs gets to push --
# overlap. Two release-tag runs on a 2-slot, 4-repo runner with a # Gitea wakes exactly one Blocked job in the group and cancels the rest
# multi-minute selftest ahead of each can easily not overlap at all, # outright, with no ordering on which one it picks (no `ORDER BY` in the
# finishing in either order regardless of push order. The step below is # query behind CancelPreviousJobsByJobConcurrency,
# what makes the OUTCOME order-independent; this group only bounds how # models/actions/run_job_list.go). What it buys is cheaper: every
# much work is wasted getting there. # execution that does reach the push step targets origin/main's live tip
# (below), never its own trigger commit, so it makes no difference which
# one wins -- the survivor pushes where any of them would have, and a
# cancelled job costs nothing. This group's only job is to stop more than
# one job from pushing AT THE SAME TIME, which is wasted work, not a
# correctness risk on its own.
concurrency: concurrency:
group: release-tag-v1 group: release-tag-v1
cancel-in-progress: false cancel-in-progress: false
@@ -150,24 +155,37 @@ jobs:
token: ${{ secrets.GITHUB_TOKEN }} token: ${{ secrets.GITHUB_TOKEN }}
fetch-depth: 0 fetch-depth: 0
# The concurrency group above only excludes a SIMULTANEOUS competing # This run's own trigger commit (${{ github.sha }}) is deliberately not
# push; it does nothing for two release-tag jobs that never overlap # what gets pushed: whichever job the concurrency group above lets
# and finish in the opposite order from the merges that triggered # through is the one that pushes, and that choice carries no relation
# them -- which a shared 2-slot runner and a multi-minute selftest # to commit recency, so every execution has to converge on the SAME
# ahead of this job make routine, not rare. So the push itself has to # target regardless of which job it is. `origin/main`'s live tip,
# be safe regardless of completion order: skip whenever v1 already # re-fetched here rather than trusted from the checkout above (which
# points at this commit or a descendant of it, rather than trusting # can be minutes stale by this point, behind its own selftest job), is
# this run to be the last one to finish. `--is-ancestor` treats a # that common target -- read fresh, every job that reaches this step
# commit as its own ancestor, so "already at" and "already ahead" are # resolves to the same commit whenever main hasn't moved between them,
# one case. A v1 that doesn't exist yet, or that shares no history # and to whatever's newest when it has.
# with this commit, falls through to the push below -- the guard is #
# A live target doesn't make the push itself safe on its own: two jobs
# can still read main at genuinely different moments if it advances
# between their two fetches, so the one with the earlier reading must
# not overwrite the other's already-pushed, newer one. That's what the
# ancestor check below still guards -- not "this job's stale trigger
# commit" any more, but "this job's freshly-read tip, which another
# job's fresher read may have already superseded." `--is-ancestor`
# treats a commit as its own ancestor, so "already at" and "already
# ahead" are one case. A v1 that doesn't exist yet, or that shares no
# history with this tip, falls through to the push -- the guard is
# only ever a reason to skip, never a reason to fail. # only ever a reason to skip, never a reason to fail.
- name: Skip if v1 already at or ahead of this commit - name: Determine origin/main's tip and whether v1 needs to move
id: check id: check
run: | run: |
git fetch origin main
TIP=$(git rev-parse origin/main)
echo "tip=$TIP" >> "$GITHUB_OUTPUT"
if git rev-parse -q --verify refs/tags/v1 >/dev/null \ if git rev-parse -q --verify refs/tags/v1 >/dev/null \
&& git merge-base --is-ancestor "${{ github.sha }}" refs/tags/v1; then && git merge-base --is-ancestor "$TIP" refs/tags/v1; then
echo "v1 already at or ahead of ${{ github.sha }} -- nothing to do" echo "v1 already at or ahead of origin/main's tip ($TIP) -- nothing to do"
echo "skip=true" >> "$GITHUB_OUTPUT" echo "skip=true" >> "$GITHUB_OUTPUT"
else else
echo "skip=false" >> "$GITHUB_OUTPUT" echo "skip=false" >> "$GITHUB_OUTPUT"
@@ -176,8 +194,8 @@ jobs:
# Lightweight tag, matching what v1 already is (`git cat-file -t v1` # Lightweight tag, matching what v1 already is (`git cat-file -t v1`
# reports `commit`, not `tag`) -- no identity needed to move it, only # reports `commit`, not `tag`) -- no identity needed to move it, only
# to push it. # to push it.
- name: Force v1 to this commit - name: Force v1 to origin/main's tip
if: steps.check.outputs.skip != 'true' if: steps.check.outputs.skip != 'true'
run: | run: |
git tag -f v1 "${{ github.sha }}" git tag -f v1 "${{ steps.check.outputs.tip }}"
git push --force origin v1 git push --force origin v1
+17 -7
View File
@@ -636,18 +636,28 @@ entry, another permission — is a `v2`, not a `v1` move. Everything else moves
**Moving the tag is automatic, gated on the same build that gates a PR.** A **Moving the tag is automatic, gated on the same build that gates a PR.** A
`release-tag` job in `.gitea/workflows/ci.yaml` runs on every push to `main`, `release-tag` job in `.gitea/workflows/ci.yaml` runs on every push to `main`,
`needs: selftest`, and force-moves `v1` to that commit once selftest `needs: selftest`, and force-moves `v1` to `origin/main`'s live tip once
succeeds — a broken build never reaches it, so `v1` can't advance onto one. selftest succeeds — a broken build never reaches it, so `v1` can't advance
It pushes with the run's built-in `GITHUB_TOKEN`; if that token turns out not onto one. It pushes with the run's built-in `GITHUB_TOKEN`; if that token
to have write access, the push step fails and the job goes red in the Actions turns out not to have write access, the push step fails and the job goes red
UI. That's a loud failure, not the silent one this replaced: `v1` stays put, in the Actions UI. That's a loud failure, not the silent one this replaced:
and nobody has to notice on their own that it lagged. `v1` stays put, and nobody has to notice on their own that it lagged.
Two merges landing close together can start two `release-tag` jobs at once; a
job-level `concurrency` group lets only one push at a time, and every job
targets `origin/main`'s current tip rather than its own trigger commit, so it
makes no difference which one the group lets through. Before pushing, the job
also checks whether `v1` already points at that tip or a descendant of it —
covering the case where two jobs read the tip at genuinely different moments
— and skips as a normal, successful outcome rather than pushing backward. A
run whose log says "nothing to do" did its job correctly; it just found
nothing to move.
This used to be a manual step, treated as a deliberate release decision taken This used to be a manual step, treated as a deliberate release decision taken
once, knowingly, after the merge — in practice it was still forgotten once, knowingly, after the merge — in practice it was still forgotten
(gitdan-actions#27): PR #25 merged to `main` and `v1` stayed on the previous (gitdan-actions#27): PR #25 merged to `main` and `v1` stayed on the previous
release until someone asked whether it had moved. The manual form below is release until someone asked whether it had moved. The manual form below is
now the recovery path, for when the automated job can't push: still the recovery path, for when the automated job can't push:
```bash ```bash
git fetch origin git fetch origin