fix(ci): add a scheduled v1 sweep and lease every v1 push
CI / shellcheck + selftests (pull_request) Successful in 1m46s
CI / move v1 to main (pull_request) Skipped

The release guard in 17d87b0 was safe but not live. Gitea 1.27.2 calls
CancelPreviousJobsByJobConcurrency whenever a job's `needs` resolve
(services/actions/clear_tasks.go:91, models/actions/run_job.go:641), so
a job's place in the `release-tag-v1` group followed when its own
selftest finished, not merge order. A newer merge C2 finishing selftest
first queued behind the older C1, C1 cancelled it, saw tip = C2, and
deferred: nobody pushed, and if merges then stopped v1 stayed stale
indefinitely behind a Skipped and a Cancelled job. The "always catches
up once merges pause" claim in ci.yaml and README was false.

What now holds:

- release-sweep.yaml runs on `schedule` every 15 minutes, in its own
  workflow and concurrency group, so nothing in ci.yaml can cancel it.
  When v1 already covers main's tip it stops after a checkout and one
  merge-base. Otherwise it checks out the tip, runs the same shellcheck
  and selftest.sh as ci.yaml's selftest job, and tags the tip only if
  they pass; a failing main therefore turns the sweep red on every tick
  while v1 lags, which is #27's AC1 loud-failure half. It reads the tip
  itself because a scheduled run's github.sha is the CommitSHA recorded
  when the schedule was registered on the last push to main
  (services/actions/notifier_helper.go:569-580,
  services/actions/schedule_tasks.go:126-141), and ref is the default
  branch: schedules are registered only from it
  (notifier_helper.go:120, :531, :603-604). event_name is "schedule"
  (context.go:71 reads TriggerEvent, set at schedule_tasks.go:136).
  Cron is 5-field robfig in UTC (models/actions/schedule_spec.go:38-41).

- 15 minutes, not 10: the sweep is the fallback, not the release path,
  and every tick is a run on gitdan-ci's shared slots and a row in the
  Actions list. 96 no-op runs a day of a few seconds each is the cost;
  the lag bound it buys is one interval plus one selftest run.

- Both writers go through scripts/release-v1.sh and push with
  --force-with-lease=refs/tags/v1:<v1 as read>, so v1 cannot move
  backwards when the sweep and a merge job race. A lost lease re-reads
  v1: at or ahead of this run's gated commit is a clean skip (the other
  writer released something at least as new); still behind it is a
  retry leased on the new value, up to three attempts, since the other
  writer may have tagged an older commit and giving up there would leave
  v1 short of a commit this run did gate; anything else goes red. A
  rejection with v1 unmoved is diagnosed as a non-lease failure and goes
  red at once.

- release-tag loses its job-level concurrency group. The lease already
  gives the ordering the group was there for, and the group was what
  cancelled the one job that could have released the newest merge.
  Without it each merge's job runs, and the one whose commit is still
  the tip when it checks releases it.

The shell moves out of ci.yaml into scripts/release-v1.sh so shellcheck
and selftest.sh cover it. release-v1-selftest.sh runs it against a
scratch bare origin: sweep no-op at and ahead of the tip, tag on a
green gate, no tag and a failing sweep on a red one, the stranded trace
above followed by a catching-up sweep, and each lost-lease outcome, with
a control showing an unleased push does step v1 back. Red-proved by
seven mutations of release-v1.sh, each failing a named assertion: plain
--force, accepting any lost lease, a merge job that never defers,
ancestry reduced to equality, no non-lease diagnosis, a sweep that never
needs to run, and a retry that does not re-lease.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
This commit is contained in:
2026-09-22 15:01:05 -05:00
co-authored by Claude Opus 5.5
parent 17d87b0647
commit ca0ee132d9
6 changed files with 371 additions and 106 deletions
+21 -22
View File
@@ -634,34 +634,32 @@ entry, another permission — is a `v2`, not a `v1` move. Everything else moves
`v1`: correctness fixes, new optional inputs, and anything internal to
`scripts/`.
**Moving the tag is automatic, gated on the same build that gates a PR.** A
`release-tag` job in `.gitea/workflows/ci.yaml` runs on every push to `main`,
`needs: selftest`, and force-moves `v1` to *that run's own commit* once
selftest succeeds — a broken build never reaches it, so `v1` can't advance
onto one. It pushes with the run's built-in `GITHUB_TOKEN`; if that token
turns out not to have write access, the push step fails and the job goes red
in the Actions UI. That's a loud failure, not the silent one this replaced:
`v1` stays put, and nobody has to notice on their own that it lagged.
**Moving the tag is automatic, gated on the same build that gates a PR.** Two
jobs move it, both through `scripts/release-v1.sh`, both with the run's
built-in `GITHUB_TOKEN`:
Before pushing, the job checks two things and pushes only if both hold: that
`origin/main`'s live tip is still this run's own commit (not some later merge
that landed while this job was queued behind its own selftest), and that `v1`
doesn't already point at that commit or a descendant of it. Either check can
skip the push, as a normal, successful outcome — a run whose log says
"nothing to do" did its job correctly. The first check is what a job-level
`concurrency` group alone can't guarantee: Gitea's wake-one/cancel-rest
handling applies no ordering by commit recency, so an older merge's job can
be the one that survives to run — that job now defers instead of releasing a
commit it never gated. The newer merge's own job releases it once its own
turn comes, which can land `v1` one merge behind the newest during a busy
stretch; it always catches up once merges pause, because every later run
re-checks against whatever is current by then.
- **`release-tag`** in `.gitea/workflows/ci.yaml` runs on every push to
`main`, `needs: selftest`, and moves `v1` to that run's own commit — but only
while that commit is still `main`'s tip. A run whose merge has already been
overtaken defers, as a successful no-op, rather than release a commit it
never gated.
- **`release-sweep.yaml`** runs every 15 minutes. When `v1` already points at
`main`'s tip or a descendant of it, it exits after a checkout and one
comparison. Otherwise it runs the same shellcheck and selftests against the
tip and moves `v1` there only if they pass.
So `v1` trails a green `main` by at most about one sweep interval plus one
selftest run, and **a `main` that fails its gate shows up as a red sweep on every
tick until it is fixed** — as does a push the token is not allowed to
make. Both jobs push with `--force-with-lease` on the `v1` they read, so
neither can move `v1` backwards over the other; a job that loses the lease to
a newer `v1` finishes green.
This used to be a manual step, treated as a deliberate release decision taken
once, knowingly, after the merge — in practice it was still forgotten
(gitdan-actions#27): PR #25 merged to `main` and `v1` stayed on the previous
release until someone asked whether it had moved. The manual form below is
still the recovery path, for when the automated job can't push:
still the recovery path, for when neither job can push:
```bash
git fetch origin
@@ -745,6 +743,7 @@ change here reaches all of them at once. That is what the gate is for.
| `seed-target-dir-selftest.sh` | seed-source preference, lock-file stripping, two jobs racing on one cache key, **and one scenario per check a hardlink clone is validated against**: a source rotated wholesale, a subtree silently lost from the walk, a copy that reports failure over a tree both other checks read as whole, and a source identity that resolved at neither end — plus a staging tree that could not be privately owned being discarded rather than published, and the publisher's log showing it waited on the consumer's own reader-lock marker before reclaiming a rotated snapshot |
| `publish-snapshot-selftest.sh` | the atomic swap, that a live consumer survives a republish, and the publisher's side of the rotation race: deferred reclamation under a live reader, and its sweep once the reader is gone |
| `prune-cache-selftest.sh` | liveness in both its forms — a branch deleted from origin, and one still on it whose tip is already merged — plus protection, locking, eviction order, self-clear, **that a cache a job claims *inside* the check-to-unlink window survives it**, and that a requirement derived from the clone's mutable set evicts exactly enough and then fails rather than under-delivering. Against a real scratch `origin`, including a genuinely shallow clone of it and a `df` that answers from the fixture's own size, since a fixed one cannot show a pass stopping |
| `release-v1-selftest.sh` | that `v1` reaches `main`'s tip only through a gate and never moves backwards: the sweep's no-op, tag and red-gate cases, the stranded-defer trace the sweep exists to recover, and each lost-lease outcome — a newer `v1` skipped cleanly (with a control showing an unleased push steps it back), an older one retried, an unrelated one and a server rejection red. Against a real scratch `origin`; the other writer is sequenced between check and push, not raced |
| `restore-mtimes-selftest.sh` | the merge hazard and the watermark that closes it, including the two-jobs-one-namespace case. Needs a real compiler. |
Every suite runs the actual script, not a reimplementation of its logic, and