Commit Graph
10 Commits
Author SHA1 Message Date
claudeandClaude Sonnet 5 ea48c03aeb fix(ci): target origin/main's live tip, not the run's own trigger commit
CI / shellcheck + selftests (pull_request) Successful in 1m35s
CI / move v1 to main (pull_request) Skipped
The wake-one-cancel-the-rest mechanism behind the concurrency group
(CancelPreviousJobsByJobConcurrency, models/actions/run_job.go, at
the v1.27.2 tag this instance runs) picks its survivor from
models/actions/run_job_list.go's query, which carries no `ORDER BY`
-- so under three-way contention on a shared 2-slot runner, the
*newest* commit's job can be the one cancelled while an older sibling
survives and, correctly from its own vantage, advances v1 forward
from a stale view. No push regresses v1 (the ancestor check from
dc1e631 already prevented that), but the newest merge goes silently
unreleased behind a cancelled job that reads as benign, not red --
exactly the failure #27 exists to end.

Every job now resolves `origin/main`'s tip fresh, right before the
push, instead of using `${{ github.sha }}`. Re-fetched explicitly
rather than trusted from the checkout step, which can be minutes
stale by this point behind its own selftest job. Every execution that
reaches the push step now converges on the same target regardless of
which job the concurrency group lets through, so which one wins the
wake no longer matters -- the survivor pushes where any of them
would have.

That doesn't make the push safe on its own: two jobs can still read
main at genuinely different moments if it advances between their two
fetches, so whichever read the tip earlier must not overwrite the
other's already-pushed, newer one. The ancestor check from dc1e631 is
kept for exactly this -- its target changed (origin/main's live tip,
not this job's own trigger commit) but its job didn't.

What each guard now protects against, after this change:
- concurrency group: stops two jobs from pushing at the same time --
  wasted work now that a cancelled job costs nothing, not a
  correctness backstop by itself.
- ancestor check: stops a job whose own fetch of the tip is stale
  relative to another job's already-pushed, fresher one from
  regressing v1.

Rewrote both the job-level comments and README's Versioning section,
which described "force-moves v1 to that commit" and said nothing
about the concurrency group or the skip-as-success case.

Re-derived the truth table against the new target (a live origin/main
tip, not a fixed commit) in a scratch origin+clone: no v1 yet (push),
v1 exactly at the tip (skip), the tip moved past v1 because a newer
merge landed (push, to the new tip -- not stuck on any prior commit),
v1 already ahead of the tip (skip, defensive), unrelated histories
(push, defensive). All five resolved as intended; ran the exact
condition and fetch sequence the workflow step uses, not a
simulation of it.

bash scripts/selftest.sh: all 6 suites green (67 prune-cache
assertions, unchanged). shellcheck -x --source-path=scripts
scripts/*.sh: clean -- as before, this does not cover the inline
`run:` shell in ci.yaml.

What remains unverifiable short of a real merge is unchanged from
dc1e631: the workflow step's actual execution inside a real Actions
run, and whether the built-in token has write access at all. Neither
this commit nor the one before it can exercise those.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 12:49:05 -05:00
claudeandClaude Sonnet 5 dc1e6317c6 fix(ci): make the v1 push monotonic, not just mutually exclusive
The concurrency group added in af1233f only excludes two release-tag
jobs that are simultaneously Running/Waiting/Blocked
(services/actions/clear_tasks.go:64-80,
models/actions/run_job.go:641-660 at the v1.27.2 tag this instance
runs) -- it has no notion of commit order between jobs that never
overlap. On a 2-slot runner shared across 4 repos, with a
multi-minute selftest gating each release-tag job, two merges landing
close together routinely finish in the opposite order from the pushes
that triggered them: if the newer commit's job completes and exits
first, the older commit's job later finds no live holder in the
group, is not blocked, and force-pushes v1 backward to itself. The
concurrency comment's "never regress it to an older one" and af1233f's
commit message both asserted the opposite -- true of the simultaneous
case the guard covers, false of the staggered one it doesn't, so
authored-false rather than drift.

Added a merge-base check before the push: skip if v1 already points
at this commit or a descendant of it (`--is-ancestor` treats a commit
as its own ancestor, so "at" and "ahead" are the same branch). A v1
that doesn't exist yet, or shares no history with this commit, falls
through to the push -- the guard only ever skips, never fails. Needs
`fetch-depth: 0` on the checkout: actions/checkout's own description
for that value is "all history for all branches and tags", and its
source (dist/index.js: fetchDepth <= 0 selects
getRefSpecForAllHistory, which includes the tags refspec) confirms
tags are fetched as part of that, not gated behind the separate
fetch-tags input -- so refs/tags/v1 and the history behind it are both
guaranteed present locally without a second fetch step.

Rewrote both false claims: the concurrency comment now says what the
group actually bounds (simultaneous competing pushes, not completion
order), and states plainly that the guard below is what makes the
outcome order-independent.

Verified the check's five cases (no tag yet, tag ahead of this
commit, tag behind this commit, tag equal to this commit, unrelated
history) against a real scratch git repo, extracting the exact
condition used in the workflow step -- all five resolved as intended
(skip only when v1 is already at or ahead). The workflow step itself
cannot be exercised outside a real Actions run.

bash scripts/selftest.sh: all 6 suites green (67 prune-cache
assertions, unchanged). shellcheck -x --source-path=scripts
scripts/*.sh: clean -- note this does not cover the inline `run:`
shell in ci.yaml, which shellcheck was never wired to check in this
repo (verified against the CI job itself, which shellchecks only
scripts/*.sh).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 12:24:27 -05:00
claudeandClaude Sonnet 5 af1233f14a fix(ci): serialise release-tag across concurrent merges
The workflow-level concurrency group is keyed per-commit (github.sha,
ci.yaml:19-28) so unrelated pushes never block each other -- but that
also means two merges landing close together run two concurrent
release-tag jobs, each force-pushing its own commit to v1. If the
older commit's job finishes last, v1 regresses to a stale-but-green
commit and stays there until the next merge corrects it forward.
Bounded blast radius (never a red commit, self-heals on the next
merge), but a silently wrong v1 is the exact failure #27 exists to
end.

Added a job-level `concurrency:` on release-tag with a fixed group
name and `cancel-in-progress: false`. Confirmed this is additive to
the workflow-level group, not a replacement, by reading gitea's source
at the v1.27.2 tag this instance runs (`tea api version`): run-level
and job-level concurrency are separate model fields
(ActionRunAttempt.ConcurrencyGroup vs ActionRunJob.ConcurrencyGroup),
evaluated by separate functions (EvaluateRunConcurrencyFillModel vs
EvaluateJobConcurrencyFillModel, services/actions/concurrency.go) and
enforced by separate cancellation paths (CancelPreviousJobsByRunConcurrency
in models/actions/run.go:364 vs CancelPreviousJobsByJobConcurrency in
models/actions/run_job.go:641) -- job_emitter.go's checkRunConcurrency
checks both groups independently (services/actions/job_emitter.go:210-236).
A fixed, non-sha group name is what serialises release-tag across
commits without touching the per-sha grouping every other job still
relies on. `cancel-in-progress: false` matters because Gitea's default
queue behaviour (documented same as GitHub's: a new queued job
supersedes an older *queued* one in the same group, but not a
*running* one) already gives the newest commit's push priority once
the running job clears; cancelling the running job too would abandon
whichever merge is currently pushing mid-flight, the same defect
wearing different clothes.

Re-verified: YAML parses, shellcheck clean, `bash scripts/selftest.sh`
all 6 suites green (67 prune-cache assertions, unchanged by this
commit).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 12:07:01 -05:00
claudeandClaude Sonnet 5 7f18cb2436 feat(ci): advance v1 automatically once the gate is green on main
Nothing moved v1 when main advanced, so a merged change was inert
until someone remembered to retag by hand — it happened on 2026-09-22
(PR #25 merged, v1 stayed on the previous release) and was only caught
because a person asked whether the tag had moved.

Option 1 from gitdan-actions#27 (automate it) over option 2 (fail loud
while it lags): a `release-tag` job, gated with `needs: selftest` so a
broken build never reaches it, force-moves v1 to the pushed commit
using the run's built-in GITHUB_TOKEN. If that token turns out not to
have write access, the push fails and the job goes red in the Actions
UI — a loud failure either way, not the silent one this replaces.
Whether the token actually has write access here is unverified short
of a real merge; that merge is the next step for this branch.

README's Versioning section documented the old manual step as a
deliberate decision "never something a merge does by itself" — that
claim is now false, so it's rewritten to describe the automated job
and keeps the manual command as the recovery path for when the job
can't push.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UkSjXXtU6JYcN2vPntWhfb
2026-09-22 11:51:51 -05:00
claudeandClaude Opus 5 5b6acd7b3b docs(ci): date the cancellation observation to 1.26.0, not the current version
CI / shellcheck + selftests (pull_request) Successful in 1m30s
The concurrency note said "this Gitea (1.26.0)" while `GET /version` now
returns 1.27.2 — the instance was upgraded on 2026-08-25/26 (daniel/gitdan's
runbook, gitdan#46). Swapping the number would have asserted the cancellation
was observed on 1.27.2, which nobody has checked: the runs it cites were seen
before the upgrade. The comment now dates the observation to 1.26.0, names the
current version, and says persistence is unverified — which is why every commit
gets its own group rather than trusting `cancel-in-progress`. That reasoning is
unchanged; only the implied currency was wrong.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LjbhSqQf3pwnPA6MVaWcWL
2026-09-02 12:49:56 -05:00
claudeandClaude Opus 5 b791896e03 fix(ci): trigger on edited so un-drafting actually lifts the draft skip
CI / shellcheck + selftests (pull_request) Successful in 1m27s
`ready_for_review` does not exist as a pull_request action on this Gitea, so
it never fired and the draft skip never lifted: a PR opened as `WIP:` carried
its skip decision to merge unless a later push happened to create a run. Draft
here is not a persisted column — it is the `WIP:` title prefix, derived by
`issue.IsWorkInProgress` — so un-drafting is a title edit, which fires a plain
`edited`.

Ported from daniel/emowheel commit 08da820 (see daniel/emowheel#71) and
daniel/gitdan (see daniel/gitdan#92), where the identical change landed first.

Accepted cost: Gitea populates no `changes` field for a title-or-body edit, so
the workflow cannot tell an un-drafting edit from an ordinary body PATCH, and
every body edit on a non-draft PR now starts a real run. That is a heavier
trade here than in the sibling repos -- this job installs two Rust toolchains
and drives a real Cargo, at `timeout-minutes: 20` -- but the `concurrency:`
block groups `pull_request` runs on `github.ref` with `cancel-in-progress:
true`, so a burst of edits collapses to one run. Still-draft PRs are unchanged.

Docs: the workflow comment is rewritten and shortened, and README's Development
section retires the empty-commit workaround it prescribed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LjbhSqQf3pwnPA6MVaWcWL
2026-09-02 12:34:24 -05:00
claude 554310186f fix(hardlink): content freshness moved switches, it was not withdrawn
CI / shellcheck + selftests (pull_request) Successful in 1m17s
`unshare_mutable_paths`' comment recorded that upstream had stopped
rewriting `dep-<target>` in place, on a measurement taken against
1.100.0-nightly. It had not. Two unrelated cargo changes landed within
four days of each other and between them moved the switch that turns the
behaviour on and the path it writes to:

  - cargo PR #17382 (2026-08-22) demoted `-Z checksum-freshness` to a
    gate and gave `build.fingerprint` the choice, defaulting to `mtime`.
    Setting only the gate is accepted and does nothing, which is exactly
    the result that was read as a withdrawal.
  - build-dir layout v2 (cargo PR #17354, stable 1.100.0 on 2026-11-12,
    nightly default since 1.99) moved the file from
    `.fingerprint/<unit>/dep-*` to `build/<pkg>/<hash>/fingerprint/dep-*`.

Measured 2026-08-26 on 1.100.0-nightly (e8cb624d5): same toolchain, same
clone procedure, one env var apart — with the gate alone a `cp -al` clone
mutates only the build/ and *.d families; add
`CARGO_BUILD_FINGERPRINT=content` and the source's dep-info file is
mutated through the shared inode again. The hazard is intact.

So the suite now exports both switches, and its strongest scenario runs
again on current nightlies — verified passing against both layouts. Its
control note learned the v2 path too: it looked for the v1 path only, and
so printed "does NOT rewrite ... in place" three lines beneath a listing
that showed the rewrite.

Every claim these comments make is now dated and cited, because the
defect being fixed is a comment that cited one measurement and silently
stopped reproducing.

Layout v2 also drags the artifacts under `build/`, which collapses this
function's real-copy set from 21.3% of a target dir to 99.998%. That is a
live cost, not a correctness problem, and it is filed as gitdan-actions#14
rather than fixed here.

Part of daniel/gitdan#62.
2026-08-26 19:48:07 -05:00
claude eb3f0bd09b docs(ci): say what the nightly step actually buys today
CI / shellcheck + selftests (pull_request) Successful in 1m24s
The step's comment claimed nightly was not optional. CI's own first green run
falsified that: 1.100.0-nightly accepts -Z checksum-freshness and resolves
freshness by mtime regardless, so the suite's probe reports it and skips the
scenario either way.

The step stays — twenty seconds, and the coverage returns by itself the day
upstream restores the behaviour — but the comment now says that rather than
the opposite. The open question about what upstream actually did is
daniel/gitdan#62.
2026-08-26 12:57:15 -05:00
claude 3f97d3d1e7 fix(selftest): make the two compiler-backed suites survive a CI runner
CI / shellcheck + selftests (pull_request) Failing after 1m36s
The new gate's first run went red on both suites that drive a real Cargo.
Neither failure was a defect in what they test; both were assumptions about
the machine, which had only ever been a dev box with a nightly installed and
no colour forcing. Fixed here rather than waived — turning a gate on is what
obliges fixing what it finds.

COLOUR. Every assertion in restore-mtimes-selftest.sh, and one in
hardlink-clone-selftest.sh, reads cargo's own words out of a build log
(`Compiling libdep`, `Fresh probe`). gitdan-ci's runner image forces colour, so
cargo wrote `Compiling\e[0m libdep` into the log and `grep -q "Compiling
libdep"` stopped matching. The suite then reported the opposite of what
happened — the failure printed the log, and the log plainly said `Compiling
libdep`. Both suites now pin CARGO_TERM_COLOR=never, which is the format their
assertions are written against.

Red-proven: with the export removed and CARGO_TERM_COLOR=always,
restore-mtimes-selftest reproduces the CI failure verbatim ("expected to find:
Compiling libdep"); with it, 14/14 pass under the same forced colour.

THE NIGHTLY PROBE ASKED THE WRONG QUESTION. `cargo +nightly -V` answers "did a
cargo proxy called with +nightly exit 0", which is not "will this build have
checksum freshness": `-V` short-circuits before `-Z` is validated at all. So
the probe said "on" on a runner where the flag was not in effect, and the
suite asserted the checksum-freshness mutation family against a build that
never had it — failing in the CONTROL, where a failure reads as "the hazard is
gone" rather than "the toolchain is wrong". It now probes the capability:
`cargo +nightly -Z checksum-freshness locate-project`, the narrowest command
that actually parses the flag. It rejects the stable channel and an unknown
flag name alike, needs no network and builds nothing.

Red-proven: the old channel probe, pointed at a cargo without the flag,
reproduces the CI failure exactly.

AND THE OFF PATH DID NOT WORK EITHER. The suite's header claimed that without
checksum freshness "the test still covers the build/ and *.d families". It did
not: the final scenario backdates the source to 2001 and asserts the rebuild
is not Fresh, which is checksum-freshness-only reasoning. Under Cargo's
ordinary mtime freshness a 2001 source IS older than the artifact and Fresh is
the correct answer, so the scenario asserted a bug. It is now gated on the
mode and skipped loudly, like the control's dep-* assertion already was: 5
assertions with a nightly, 3 without.

CI installs a nightly as well as stable (nightly first, so stable stays the
default) — that hazard is the one this whole scheme exists to close, and a CI
that skips it is checking the cheap half.
2026-08-26 12:46:03 -05:00
claude f76789358d ci(gitea): gate scripts/ with shellcheck and the selftests
This repo had eight scripts, six selftest suites and nothing that ran any of
them. It is a composite action consumed at `@v1` — a moving tag — by three
repos' CI, so a bad edit here reaches zemyna, emowheel and lublub at once and
is discovered by whichever of them builds next.

One job: `shellcheck -x --source-path=scripts scripts/*.sh`, then
`bash scripts/selftest.sh`. `-x` follows the `. cache-lib.sh` every script
sources, which is where most of the logic being checked lives; without it
shellcheck reports SC1091 and analyses each file with a hole in it.

The full suite, not `--fast`: `hardlink-clone-selftest.sh` and
`restore-mtimes-selftest.sh` are the two that drive a real Cargo rather than a
fixture, and selftest.sh's own header says `--fast` is for iterating, not for
signing off a change. So the job installs a stable toolchain. The scratch
workspaces they build use path dependencies only, so nothing reaches
crates.io, and the job references no credentials at all.

Modelled on daniel/gitdan's workflow, including the two clauses of the
draft-skip guard (`github.event_name != 'pull_request' ||
!github.event.pull_request.draft` — the first is what stops the second from
also skipping pushes to `main`, where there is no `pull_request` context) and
the per-commit concurrency group for `push`.

No `container.volumes:` entry, deliberately: the suites want a cold target dir
every run, since "did this rebuild?" is exactly what restore-mtimes-selftest
asks. This job takes no share of the shared CI cache disk budget.

Turning the gate on surfaced three pre-existing findings, all fixed here
rather than suppressed or waived:

- `restore-mtimes-selftest.sh` SC2038: `find | xargs touch` -> `-print0 | xargs -0`.
- `hardlink-clone-selftest.sh` SC2295: `${f#$base_fix/}` -> `${f#"$base_fix"/}`.
- `cache-lib.sh` SC2016: a per-line disable with the rationale. The single
  quotes are load-bearing — the program is for the inner shell, where `$f` is
  its loop variable, `$rc` its accumulator and `$$` its pid — so this is the
  documented-false-positive case, not a stopgap.
2026-08-26 12:31:50 -05:00