22dafe45e22310377251111f04483eebf328cf6a
13
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
77cc5917b6
|
fix(release): refuse to overwrite an unrelated hand-placed v1
release-v1.sh's push_leased() only detected a lost lease after a push
was *rejected* -- but force-with-lease compares the remote ref's raw
value against the caller's expected value, not ancestry. If v1 already
sat on a hand-placed, unrelated commit when push_leased() was first
called (no race, nobody moves it mid-call), the very first push found
the ref exactly where it expected, succeeded outright, and silently
overwrote the unrelated v1 with <sha> -- skipping every ancestry check
the function has, since those only run after a rejection.
Fix: before the first push attempt, check whether the caller's
`expect` is neither an ancestor of `sha` (the ordinary stale-v1 case)
nor already covering it (nothing to do) -- and go red naming both SHAs
if so. `expect` is always a peeled commit (fetch_v1() reads
`refs/tags/v1^{commit}`), so this doesn't add a second failure mode
for an annotated v1; that tag form's existing "not a lost lease"
behavior on the first rejected push is untouched.
Surfaced by PR #28's final review. New selftest scenario 11 in
release-v1-selftest.sh, red-proven against the unfixed script (v1 was
silently moved off the stray commit); green after the fix, with the
full 7-suite gate (shellcheck + selftest.sh) passing.
Ride-alongs from the same review:
- ci.yaml:120-123 claimed the README's Versioning section documented
what's verified about the release token's write access; it said
nothing. Added an accurate sentence there (the grant is unobserved
until the first merge, capped by repo/owner token-permission maxima
and unreadable tag protections) and pointed the comment at it.
- Deleted two comment-as-decision-history paragraphs per
comments-are-not-exposition: ci.yaml's "no job-level concurrency"
rationale (kept one line of intent) and
prune-cache-selftest.sh scenario 9's account of how an
existence-only assertion used to pass with the pass-2 guard removed
(kept a one-line statement of what it checks).
- README's release-v1-selftest.sh table row now names the new
hand-placed-v1 scenario.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
|
||
|
|
ca0ee132d9
|
fix(ci): add a scheduled v1 sweep and lease every v1 push
The release guard in
|
||
|
|
17d87b0647
|
fix(ci): only ever push a commit this run actually gated; fix vacuous scenario-9 guard
## v1 could advance onto an ungated commit
`needs: selftest` gates this run's own commit, but the push targeted
origin/main's freshly-fetched tip with nothing comparing the two.
Trace: M1 merges green; M2 merges while M1's selftest is still
running; M1's release-tag job fetches tip = M2 and pushes v1 = M2,
whose own selftest may be queued, running, or red. If M2 is red, its
own job is skipped, so v1 sits on a red commit across every consuming
project until the next green merge -- with nothing red pointing at
the release itself. ci.yaml:107-109 and README.md:640-641 both
asserted this couldn't happen; ea48c03's own diff established the
precondition (tip "can be minutes stale... behind its own selftest
job") without closing it.
Fix: skip the push unless origin/main's tip IS this run's own
github.sha, checked before the existing v1-monotonicity check
(
|
||
|
|
ea48c03aeb
|
fix(ci): target origin/main's live tip, not the run's own trigger commit
The wake-one-cancel-the-rest mechanism behind the concurrency group (CancelPreviousJobsByJobConcurrency, models/actions/run_job.go, at the v1.27.2 tag this instance runs) picks its survivor from models/actions/run_job_list.go's query, which carries no `ORDER BY` -- so under three-way contention on a shared 2-slot runner, the *newest* commit's job can be the one cancelled while an older sibling survives and, correctly from its own vantage, advances v1 forward from a stale view. No push regresses v1 (the ancestor check from |
||
|
|
dc1e6317c6
|
fix(ci): make the v1 push monotonic, not just mutually exclusive
The concurrency group added in
|
||
|
|
af1233f14a
|
fix(ci): serialise release-tag across concurrent merges
The workflow-level concurrency group is keyed per-commit (github.sha, ci.yaml:19-28) so unrelated pushes never block each other -- but that also means two merges landing close together run two concurrent release-tag jobs, each force-pushing its own commit to v1. If the older commit's job finishes last, v1 regresses to a stale-but-green commit and stays there until the next merge corrects it forward. Bounded blast radius (never a red commit, self-heals on the next merge), but a silently wrong v1 is the exact failure #27 exists to end. Added a job-level `concurrency:` on release-tag with a fixed group name and `cancel-in-progress: false`. Confirmed this is additive to the workflow-level group, not a replacement, by reading gitea's source at the v1.27.2 tag this instance runs (`tea api version`): run-level and job-level concurrency are separate model fields (ActionRunAttempt.ConcurrencyGroup vs ActionRunJob.ConcurrencyGroup), evaluated by separate functions (EvaluateRunConcurrencyFillModel vs EvaluateJobConcurrencyFillModel, services/actions/concurrency.go) and enforced by separate cancellation paths (CancelPreviousJobsByRunConcurrency in models/actions/run.go:364 vs CancelPreviousJobsByJobConcurrency in models/actions/run_job.go:641) -- job_emitter.go's checkRunConcurrency checks both groups independently (services/actions/job_emitter.go:210-236). A fixed, non-sha group name is what serialises release-tag across commits without touching the per-sha grouping every other job still relies on. `cancel-in-progress: false` matters because Gitea's default queue behaviour (documented same as GitHub's: a new queued job supersedes an older *queued* one in the same group, but not a *running* one) already gives the newest commit's push priority once the running job clears; cancelling the running job too would abandon whichever merge is currently pushing mid-flight, the same defect wearing different clothes. Re-verified: YAML parses, shellcheck clean, `bash scripts/selftest.sh` all 6 suites green (67 prune-cache assertions, unchanged by this commit). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb |
||
|
|
7f18cb2436
|
feat(ci): advance v1 automatically once the gate is green on main
Nothing moved v1 when main advanced, so a merged change was inert until someone remembered to retag by hand — it happened on 2026-09-22 (PR #25 merged, v1 stayed on the previous release) and was only caught because a person asked whether the tag had moved. Option 1 from gitdan-actions#27 (automate it) over option 2 (fail loud while it lags): a `release-tag` job, gated with `needs: selftest` so a broken build never reaches it, force-moves v1 to the pushed commit using the run's built-in GITHUB_TOKEN. If that token turns out not to have write access, the push fails and the job goes red in the Actions UI — a loud failure either way, not the silent one this replaces. Whether the token actually has write access here is unverified short of a real merge; that merge is the next step for this branch. README's Versioning section documented the old manual step as a deliberate decision "never something a merge does by itself" — that claim is now false, so it's rewritten to describe the automated job and keeps the manual command as the recovery path for when the job can't push. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb |
||
|
|
5b6acd7b3b
|
docs(ci): date the cancellation observation to 1.26.0, not the current version
CI / shellcheck + selftests (pull_request) Successful in 1m30s
The concurrency note said "this Gitea (1.26.0)" while `GET /version` now returns 1.27.2 — the instance was upgraded on 2026-08-25/26 (daniel/gitdan's runbook, gitdan#46). Swapping the number would have asserted the cancellation was observed on 1.27.2, which nobody has checked: the runs it cites were seen before the upgrade. The comment now dates the observation to 1.26.0, names the current version, and says persistence is unverified — which is why every commit gets its own group rather than trusting `cancel-in-progress`. That reasoning is unchanged; only the implied currency was wrong. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LjbhSqQf3pwnPA6MVaWcWL |
||
|
|
b791896e03
|
fix(ci): trigger on edited so un-drafting actually lifts the draft skip
CI / shellcheck + selftests (pull_request) Successful in 1m27s
`ready_for_review` does not exist as a pull_request action on this Gitea, so it never fired and the draft skip never lifted: a PR opened as `WIP:` carried its skip decision to merge unless a later push happened to create a run. Draft here is not a persisted column — it is the `WIP:` title prefix, derived by `issue.IsWorkInProgress` — so un-drafting is a title edit, which fires a plain `edited`. Ported from daniel/emowheel commit 08da820 (see daniel/emowheel#71) and daniel/gitdan (see daniel/gitdan#92), where the identical change landed first. Accepted cost: Gitea populates no `changes` field for a title-or-body edit, so the workflow cannot tell an un-drafting edit from an ordinary body PATCH, and every body edit on a non-draft PR now starts a real run. That is a heavier trade here than in the sibling repos -- this job installs two Rust toolchains and drives a real Cargo, at `timeout-minutes: 20` -- but the `concurrency:` block groups `pull_request` runs on `github.ref` with `cancel-in-progress: true`, so a burst of edits collapses to one run. Still-draft PRs are unchanged. Docs: the workflow comment is rewritten and shortened, and README's Development section retires the empty-commit workaround it prescribed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LjbhSqQf3pwnPA6MVaWcWL |
||
|
|
554310186f
|
fix(hardlink): content freshness moved switches, it was not withdrawn
CI / shellcheck + selftests (pull_request) Successful in 1m17s
`unshare_mutable_paths`' comment recorded that upstream had stopped rewriting `dep-<target>` in place, on a measurement taken against 1.100.0-nightly. It had not. Two unrelated cargo changes landed within four days of each other and between them moved the switch that turns the behaviour on and the path it writes to: - cargo PR #17382 (2026-08-22) demoted `-Z checksum-freshness` to a gate and gave `build.fingerprint` the choice, defaulting to `mtime`. Setting only the gate is accepted and does nothing, which is exactly the result that was read as a withdrawal. - build-dir layout v2 (cargo PR #17354, stable 1.100.0 on 2026-11-12, nightly default since 1.99) moved the file from `.fingerprint/<unit>/dep-*` to `build/<pkg>/<hash>/fingerprint/dep-*`. Measured 2026-08-26 on 1.100.0-nightly (e8cb624d5): same toolchain, same clone procedure, one env var apart — with the gate alone a `cp -al` clone mutates only the build/ and *.d families; add `CARGO_BUILD_FINGERPRINT=content` and the source's dep-info file is mutated through the shared inode again. The hazard is intact. So the suite now exports both switches, and its strongest scenario runs again on current nightlies — verified passing against both layouts. Its control note learned the v2 path too: it looked for the v1 path only, and so printed "does NOT rewrite ... in place" three lines beneath a listing that showed the rewrite. Every claim these comments make is now dated and cited, because the defect being fixed is a comment that cited one measurement and silently stopped reproducing. Layout v2 also drags the artifacts under `build/`, which collapses this function's real-copy set from 21.3% of a target dir to 99.998%. That is a live cost, not a correctness problem, and it is filed as gitdan-actions#14 rather than fixed here. Part of daniel/gitdan#62. |
||
|
|
eb3f0bd09b
|
docs(ci): say what the nightly step actually buys today
CI / shellcheck + selftests (pull_request) Successful in 1m24s
The step's comment claimed nightly was not optional. CI's own first green run falsified that: 1.100.0-nightly accepts -Z checksum-freshness and resolves freshness by mtime regardless, so the suite's probe reports it and skips the scenario either way. The step stays — twenty seconds, and the coverage returns by itself the day upstream restores the behaviour — but the comment now says that rather than the opposite. The open question about what upstream actually did is daniel/gitdan#62. |
||
|
|
3f97d3d1e7
|
fix(selftest): make the two compiler-backed suites survive a CI runner
CI / shellcheck + selftests (pull_request) Failing after 1m36s
The new gate's first run went red on both suites that drive a real Cargo.
Neither failure was a defect in what they test; both were assumptions about
the machine, which had only ever been a dev box with a nightly installed and
no colour forcing. Fixed here rather than waived — turning a gate on is what
obliges fixing what it finds.
COLOUR. Every assertion in restore-mtimes-selftest.sh, and one in
hardlink-clone-selftest.sh, reads cargo's own words out of a build log
(`Compiling libdep`, `Fresh probe`). gitdan-ci's runner image forces colour, so
cargo wrote `Compiling\e[0m libdep` into the log and `grep -q "Compiling
libdep"` stopped matching. The suite then reported the opposite of what
happened — the failure printed the log, and the log plainly said `Compiling
libdep`. Both suites now pin CARGO_TERM_COLOR=never, which is the format their
assertions are written against.
Red-proven: with the export removed and CARGO_TERM_COLOR=always,
restore-mtimes-selftest reproduces the CI failure verbatim ("expected to find:
Compiling libdep"); with it, 14/14 pass under the same forced colour.
THE NIGHTLY PROBE ASKED THE WRONG QUESTION. `cargo +nightly -V` answers "did a
cargo proxy called with +nightly exit 0", which is not "will this build have
checksum freshness": `-V` short-circuits before `-Z` is validated at all. So
the probe said "on" on a runner where the flag was not in effect, and the
suite asserted the checksum-freshness mutation family against a build that
never had it — failing in the CONTROL, where a failure reads as "the hazard is
gone" rather than "the toolchain is wrong". It now probes the capability:
`cargo +nightly -Z checksum-freshness locate-project`, the narrowest command
that actually parses the flag. It rejects the stable channel and an unknown
flag name alike, needs no network and builds nothing.
Red-proven: the old channel probe, pointed at a cargo without the flag,
reproduces the CI failure exactly.
AND THE OFF PATH DID NOT WORK EITHER. The suite's header claimed that without
checksum freshness "the test still covers the build/ and *.d families". It did
not: the final scenario backdates the source to 2001 and asserts the rebuild
is not Fresh, which is checksum-freshness-only reasoning. Under Cargo's
ordinary mtime freshness a 2001 source IS older than the artifact and Fresh is
the correct answer, so the scenario asserted a bug. It is now gated on the
mode and skipped loudly, like the control's dep-* assertion already was: 5
assertions with a nightly, 3 without.
CI installs a nightly as well as stable (nightly first, so stable stays the
default) — that hazard is the one this whole scheme exists to close, and a CI
that skips it is checking the cheap half.
|
||
|
|
f76789358d
|
ci(gitea): gate scripts/ with shellcheck and the selftests
This repo had eight scripts, six selftest suites and nothing that ran any of
them. It is a composite action consumed at `@v1` — a moving tag — by three
repos' CI, so a bad edit here reaches zemyna, emowheel and lublub at once and
is discovered by whichever of them builds next.
One job: `shellcheck -x --source-path=scripts scripts/*.sh`, then
`bash scripts/selftest.sh`. `-x` follows the `. cache-lib.sh` every script
sources, which is where most of the logic being checked lives; without it
shellcheck reports SC1091 and analyses each file with a hole in it.
The full suite, not `--fast`: `hardlink-clone-selftest.sh` and
`restore-mtimes-selftest.sh` are the two that drive a real Cargo rather than a
fixture, and selftest.sh's own header says `--fast` is for iterating, not for
signing off a change. So the job installs a stable toolchain. The scratch
workspaces they build use path dependencies only, so nothing reaches
crates.io, and the job references no credentials at all.
Modelled on daniel/gitdan's workflow, including the two clauses of the
draft-skip guard (`github.event_name != 'pull_request' ||
!github.event.pull_request.draft` — the first is what stops the second from
also skipping pushes to `main`, where there is no `pull_request` context) and
the per-commit concurrency group for `push`.
No `container.volumes:` entry, deliberately: the suites want a cold target dir
every run, since "did this rebuild?" is exactly what restore-mtimes-selftest
asks. This job takes no share of the shared CI cache disk budget.
Turning the gate on surfaced three pre-existing findings, all fixed here
rather than suppressed or waived:
- `restore-mtimes-selftest.sh` SC2038: `find | xargs touch` -> `-print0 | xargs -0`.
- `hardlink-clone-selftest.sh` SC2295: `${f#$base_fix/}` -> `${f#"$base_fix"/}`.
- `cache-lib.sh` SC2016: a per-line disable with the rationale. The single
quotes are load-bearing — the program is for the inner shell, where `$f` is
its loop variable, `$rc` its accumulator and `$$` its pid — so this is the
documented-false-positive case, not a stopgap.
|