## v1 could advance onto an ungated commit
`needs: selftest` gates this run's own commit, but the push targeted
origin/main's freshly-fetched tip with nothing comparing the two.
Trace: M1 merges green; M2 merges while M1's selftest is still
running; M1's release-tag job fetches tip = M2 and pushes v1 = M2,
whose own selftest may be queued, running, or red. If M2 is red, its
own job is skipped, so v1 sits on a red commit across every consuming
project until the next green merge -- with nothing red pointing at
the release itself. ci.yaml:107-109 and README.md:640-641 both
asserted this couldn't happen; ea48c03's own diff established the
precondition (tip "can be minutes stale... behind its own selftest
job") without closing it.
Fix: skip the push unless origin/main's tip IS this run's own
github.sha, checked before the existing v1-monotonicity check
(dc1e631) rather than replacing it -- the two compose (tip-mismatch
first, since it's the coarser reason to defer; ancestor-check second,
for a duplicate run whose own commit is still current). Reverted the
push target from the fetched tip back to `${{ github.sha }}`, now
that the guard makes them provably equal whenever the push fires.
## Convergence trace: does the newest commit's job still run?
Read gitea's source further at the pinned v1.27.2 tag:
PrepareToStartJobWithConcurrency (services/actions/clear_tasks.go)
calls CancelPreviousJobsByJobConcurrency on every job entering the
group, unconditionally cancelling whatever was previously
Waiting/Blocked there -- so at most one job sits queued in the group
at a time; each new arrival supersedes it. Because job-level
concurrency is only evaluated once `needs: selftest` is satisfied
(job_emitter.go re-evaluates readiness there), "arrival order" tracks
each commit's own selftest-completion time, not raw merge order -- an
older commit with a slower selftest can enter the group after a
newer one and cancel its queued slot.
That cancelled job is gone for good; it will never push. But the
commit that's genuinely current at the moment merges stop arriving is
always the one still queued when the running job finishes, because
every subsequent arrival (from every subsequent merge, not just the
"newest" one at any single instant) keeps re-superseding the queue.
So v1 always eventually catches up -- "one merge later" in the common
case, "at the next merge, whenever that happens" in the adversarial
case where a stale survivor runs, finds itself no longer current,
defers, and nothing is left queued. It cannot get stuck forever short
of the repository never receiving another merge, because every future
push re-attempts the same check against whatever's current by then.
Stated this plainly in the comment and README rather than repeating
the false "never" guarantee in softer words.
Re-derived the truth table against the new guard in a scratch
origin+clone, six cases: own commit == tip, no v1 (push); v1 already
== own commit (skip, duplicate run); tip moved past own gated commit
because a newer merge landed (skip, defers); the newer commit's own
run once nothing further has landed (push); v1 already ahead of a
now-stale gated commit (skip, tip-mismatch catches it first); tip ==
own commit but v1 independently ahead via a local-only descendant,
isolating the second (ancestor) check on its own (skip). All six
resolved as intended.
## Scenario 9's pass-3 guard was vacuous
scripts/prune-cache-selftest.sh:230 grepped 'self-clear' against
$scratch/log. summary_line() (cache-lib.sh) writes only to
$GITHUB_STEP_SUMMARY; the self-clear branch (prune-cache.sh:632)
writes 'self-clear' there and 'clearing own' to stdout (:631) --
'self-clear' never appears in $scratch/log at all, so the branch was
unreachable and the `ok` unconditional. Scenario 8 already greps the
right string against the right file (assert_log "clearing own" ...);
scenario 9 now does the same, staying on $scratch/log where it
already was -- the string was wrong, not the file.
Red-proved by capturing a real prune-cache.sh log where self-clear
genuinely fired (scenario 9's own fixture with MIN_FREE_PCT
temporarily raised to 100, in a scratch copy, reverted after) and
running both patterns against it: `grep -q 'self-clear'` -> no match
(the old check's vacuous pass, confirmed); `grep -q 'clearing own'`
-> match (the fix's correct fail). Matches the reviewer's own
measurement exactly. The real prune-cache-selftest.sh was untouched
during this experiment; only the grep string changed in the actual
commit.
bash scripts/selftest.sh: all 6 suites green (67 prune-cache
assertions, unchanged in count -- the fix corrects what scenario 9's
existing check compares, not what it asserts). shellcheck -x
--source-path=scripts scripts/*.sh: clean, as before not covering the
inline ci.yaml shell.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
206 lines
11 KiB
YAML
206 lines
11 KiB
YAML
name: CI
|
|
|
|
on:
|
|
push:
|
|
branches: [main]
|
|
pull_request:
|
|
branches: [main]
|
|
# Spelled out only to keep `edited` in the list — naming any type replaces
|
|
# the whole default set, so the other three have to be restated.
|
|
#
|
|
# `edited`, not `ready_for_review`: draft here is the `WIP:` title prefix,
|
|
# so un-drafting is a title edit and no ready-for-review action is ever
|
|
# raised. That edit is what creates the run which lifts the `if:` skip
|
|
# below (decided once, when a run is CREATED). Do not swap it back and do
|
|
# not drop this to the bare default — either restores the bug. The accepted
|
|
# cost and the evidence are in README's Development section.
|
|
types: [opened, synchronize, reopened, edited]
|
|
|
|
# gitdan-ci runs four repos' CI on two capacity slots, and the compiler-backed
|
|
# suites below are multi-minute. A superseded run costs a slot in front of
|
|
# somebody's build, so drop it.
|
|
#
|
|
# `push` groups on `github.sha` rather than `github.ref`: a constant per-branch
|
|
# group is what let this Gitea cancel two of daniel/gitdan's merge runs
|
|
# outright while `cancel-in-progress` was gated away from `push` entirely —
|
|
# see the long note in that repo's ci.yaml. Observed on 1.26.0; the instance
|
|
# is 1.27.2 now and whether it persists is unverified, which is why every
|
|
# commit gets its own group instead of trusting the flag.
|
|
concurrency:
|
|
group: ${{ github.workflow }}-${{ github.event_name == 'pull_request' && github.ref || github.sha }}
|
|
cancel-in-progress: true
|
|
|
|
jobs:
|
|
selftest:
|
|
name: shellcheck + selftests
|
|
# Two clauses, both load-bearing. The second skips draft PRs: Gitea sets
|
|
# draft:true when the title starts with `WIP:`, so work-in-progress pushes
|
|
# cost the shared runner nothing until the PR is un-WIP'd. The first is
|
|
# what keeps that from also skipping pushes to `main` — a `push` event has
|
|
# no `pull_request` context, so `github.event.pull_request.draft` is empty
|
|
# there and the negation alone would be unreliable. Never drop it.
|
|
if: ${{ github.event_name != 'pull_request' || !github.event.pull_request.draft }}
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 20
|
|
# Deliberately no `container.volumes:` entry, unlike every repo that
|
|
# CONSUMES this action. The suites below build throwaway workspaces under
|
|
# `mktemp -d` and want a cold target dir every time — a persistent cache
|
|
# would make "did this run rebuild?" unanswerable, which is the question
|
|
# restore-mtimes-selftest.sh exists to ask. So this job takes no share of
|
|
# the shared CI cache disk budget.
|
|
steps:
|
|
- name: Checkout sources
|
|
uses: actions/checkout@v4
|
|
|
|
- name: Install shellcheck
|
|
uses: taiki-e/install-action@v2
|
|
with:
|
|
tool: shellcheck
|
|
|
|
# `hardlink-clone-selftest.sh` and `restore-mtimes-selftest.sh` drive a
|
|
# real Cargo against a real scratch workspace — they are the only things
|
|
# here that verify the hardlink-aliasing and mtime-freshness behaviour
|
|
# against the compiler rather than against a fixture, and selftest.sh's
|
|
# own header says `--fast` is for iterating, not for signing off a
|
|
# change. So CI installs a toolchain and runs the full set.
|
|
#
|
|
# The scratch workspaces use path dependencies only, so nothing here
|
|
# reaches crates.io.
|
|
# Nightly first, stable second, so stable ends up the default and
|
|
# nightly is only reachable through an explicit `+nightly`.
|
|
#
|
|
# `hardlink-clone-selftest.sh`'s last scenario needs a Cargo that
|
|
# resolves freshness by CONTENT — the mode where the dep-info file
|
|
# carries per-source checksums, which is the mutation that turns a
|
|
# hardlink clone into silent stale-artifact reuse rather than a slow
|
|
# build. This step is what supplies it, and as of 2026-08-26 it does:
|
|
# the scenario ran and passed against 1.100.0-nightly.
|
|
#
|
|
# It briefly did not. Cargo PR #17382 (2026-08-22) demoted
|
|
# `-Z checksum-freshness` to a gate and gave `build.fingerprint` the
|
|
# choice, defaulting to `mtime`, so the suite — which set only the gate —
|
|
# measured a genuine INACTIVE and skipped its strongest scenario. That
|
|
# read as "upstream withdrew content freshness" and was written up here
|
|
# as this step buying nothing. It was a moved switch, not a withdrawal;
|
|
# the suite now exports both and the coverage is back. See
|
|
# daniel/gitdan#62 for the investigation.
|
|
- name: Install Rust nightly
|
|
uses: dtolnay/rust-toolchain@nightly
|
|
- name: Install Rust toolchain
|
|
uses: dtolnay/rust-toolchain@stable
|
|
|
|
# Every script, including the suites themselves. `-x` follows the
|
|
# `. cache-lib.sh` each one sources, which is where most of the logic
|
|
# being checked actually lives; without it shellcheck reports SC1091 and
|
|
# analyses each file with a hole in it.
|
|
- name: shellcheck
|
|
run: shellcheck -x --source-path=scripts scripts/*.sh
|
|
|
|
# One command, not six: selftest.sh is the entry point a developer runs,
|
|
# so a suite added there is gated here without a matching edit in this
|
|
# file.
|
|
- name: Selftests
|
|
run: bash scripts/selftest.sh
|
|
|
|
release-tag:
|
|
name: move v1 to main
|
|
# `needs:` is what makes this "after the gate is green" rather than
|
|
# merely "after a push": a failed selftest skips this job outright, so
|
|
# v1 can never advance onto a broken build. The `if:` restricts it to an
|
|
# actual push to main -- a pull_request run targeting main shares this
|
|
# workflow but has no ref worth tagging.
|
|
needs: selftest
|
|
if: ${{ github.event_name == 'push' && github.ref == 'refs/heads/main' }}
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 2
|
|
# The workflow-level group above is keyed per-commit (github.sha) so
|
|
# unrelated commits' CI never blocks each other -- which also means two
|
|
# merges landing close together can run two concurrent release-tag jobs.
|
|
# A job-level `concurrency:` is a second, independent group scoped to
|
|
# this job alone -- it does not replace the workflow-level one, it adds
|
|
# to it (confirmed by reading gitea's source at the v1.27.2 tag this
|
|
# instance runs: run-level and job-level concurrency are separate model
|
|
# fields, evaluated and enforced by separate functions --
|
|
# CancelPreviousJobsByRunConcurrency vs CancelPreviousJobsByJobConcurrency
|
|
# in models/actions/{run,run_job}.go -- not one overriding the other).
|
|
#
|
|
# This does NOT decide which of several contending jobs gets to push --
|
|
# Gitea wakes exactly one Blocked job in the group and cancels the rest
|
|
# outright, with no ordering on which one it picks (no `ORDER BY` in the
|
|
# query behind CancelPreviousJobsByJobConcurrency,
|
|
# models/actions/run_job_list.go). Every execution still only ever pushes
|
|
# its OWN gated commit (below), never another job's, so a job that gets
|
|
# cancelled here costs nothing beyond its own wasted run -- it was never
|
|
# going to push anyone else's commit either. This group's only job is to
|
|
# stop more than one job from pushing AT THE SAME TIME, which is wasted
|
|
# work, not a correctness risk on its own.
|
|
concurrency:
|
|
group: release-tag-v1
|
|
cancel-in-progress: false
|
|
# Requests write access from the run's built-in token (see README's
|
|
# Versioning section for what's actually verified about it). Without
|
|
# this the checkout below still succeeds -- it's the push that would be
|
|
# rejected, which is a red job, not a silent no-op.
|
|
permissions:
|
|
contents: write
|
|
steps:
|
|
# fetch-depth: 0 fetches full history AND all tags (actions/checkout's
|
|
# own description: "0 indicates all history for all branches and
|
|
# tags") -- REQUIRED so refs/tags/v1 and the ancestry behind it are
|
|
# both present locally for the merge-base check below, regardless of
|
|
# which commit this run happens to be built from.
|
|
- uses: actions/checkout@v4
|
|
with:
|
|
token: ${{ secrets.GITHUB_TOKEN }}
|
|
fetch-depth: 0
|
|
|
|
# `needs: selftest` only gates THIS run's own commit -- it says nothing
|
|
# about whether `origin/main` has since moved on to a merge whose own
|
|
# selftest hasn't finished, is still queued, or is red. Two things
|
|
# follow, checked in this order:
|
|
#
|
|
# 1. If `origin/main`'s live tip (re-fetched here, not trusted from the
|
|
# checkout above, which can be minutes stale behind this job's own
|
|
# selftest) is no longer THIS run's own `github.sha`, some other
|
|
# merge has landed since. Pushing it would release a commit this run
|
|
# never gated -- so this run defers instead, unconditionally. The
|
|
# commit that IS the live tip has its own run, and that run's own
|
|
# guard is what releases it once ITS turn to push comes -- possibly
|
|
# only once merges pause for a moment, so v1 can land one merge
|
|
# later than the newest one in a busy stretch. It never lands on a
|
|
# commit that wasn't gated, and it always catches up once things go
|
|
# quiet, because every future merge re-attempts the same check
|
|
# against whatever is current by then.
|
|
# 2. Only once (1) confirms this IS the live tip does it matter whether
|
|
# v1 already covers it -- a duplicate run, or a manual push already
|
|
# having done this. `--is-ancestor` treats a commit as its own
|
|
# ancestor, so "already at" and "already ahead" are one case. A v1
|
|
# that doesn't exist yet, or shares no history with this commit,
|
|
# falls through to the push -- both checks are only ever a reason to
|
|
# skip, never a reason to fail.
|
|
- name: Determine whether this commit is still current and needs releasing
|
|
id: check
|
|
run: |
|
|
git fetch origin main
|
|
TIP=$(git rev-parse origin/main)
|
|
SHA="${{ github.sha }}"
|
|
if [ "$TIP" != "$SHA" ]; then
|
|
echo "origin/main's tip ($TIP) has moved past this run's own gated commit ($SHA) -- deferring to whichever run's own commit is now the live tip"
|
|
echo "skip=true" >> "$GITHUB_OUTPUT"
|
|
elif git rev-parse -q --verify refs/tags/v1 >/dev/null \
|
|
&& git merge-base --is-ancestor "$SHA" refs/tags/v1; then
|
|
echo "v1 already at or ahead of $SHA -- nothing to do"
|
|
echo "skip=true" >> "$GITHUB_OUTPUT"
|
|
else
|
|
echo "skip=false" >> "$GITHUB_OUTPUT"
|
|
fi
|
|
|
|
# Lightweight tag, matching what v1 already is (`git cat-file -t v1`
|
|
# reports `commit`, not `tag`) -- no identity needed to move it, only
|
|
# to push it.
|
|
- name: Force v1 to this commit
|
|
if: steps.check.outputs.skip != 'true'
|
|
run: |
|
|
git tag -f v1 "${{ github.sha }}"
|
|
git push --force origin v1
|