Commit Graph
34 Commits
Author SHA1 Message Date
JMR-dev 58447d7d12 ci(e2e): gate the emulator on window focus to fix the RootViewPicker flake (#468)
Root cause: intermittently the launched activity window has has-window-focus=false
for the WHOLE instrumented run, so Espresso's RootViewPicker (onView().check(),
Intents.intended(), pressBack(), focus-dependent clipboard) times out after 10s and
fails EVERY focus-dependent test at once while the ~280 pure-Compose semantics tests
(which don't need window focus) pass. A failing E2E (35) leg's logcat (PR #470, run
28985259521) shows has-window-focus=true ZERO times across the whole session and both
the first attempt and the once-retry fail identically -- a persistent environmental
state, not a per-test transient. The prior mitigation, a single fire-and-forget
`adb shell input keyevent 82` (MENU) right after boot, is too weak: MENU no longer
dismisses the modern (API 30+) keyguard and, delivered before SystemUI/keyguard comes
up, is simply dropped -- so the insecure keyguard / non-interactive display persists
and no app window ever takes focus.

Fix: a single shared helper, .github/scripts/emulator_focus_gate.py, invoked
identically by BOTH E2E jobs (the e2e API 29-36 matrix AND e2e-preview API 37) and by
the local preflight runners (local_instrumented.py / api37_e2e.py), so it cannot drift.
It wakes the display (KEYCODE_WAKEUP), dismisses + disables the keyguard
(wm dismiss-keyguard, locksettings set-disabled true), keeps the screen on
(svc power stayon true + max screen_off_timeout), zeroes the animation scales, then
polls dumpsys power/window until the device is interactive AND a real window holds
input focus (mCurrentFocus non-null) -- re-nudging each iteration -- before the suite
runs. Applied uniformly, this also gives e2e-preview the animation-disable the matrix
already had. The gate is soft (bounded wait, then proceeds with a ::warning:: and the
final device state) and non-fatal (`|| true`), preserving #454's guarantee that the
unlock never aborts the boot; it leaves #454's manual boot, #460's path-filter and
#464's wedge-capture untouched.

The pure readiness parser is unit-tested by test_emulator_focus_gate.py (run by the
traffic-control-tests job). Determinism is validated by this PR's own matrix run.

Closes #468
2026-07-08 20:32:16 -05:00
JMR-dev 9407b20914 ci(e2e): restore wedge-capture diagnostics on the matrix E2E legs (#421)
Mirror the e2e-preview job's proven wedge-capture onto the API 29-36 matrix
legs, now that #454 replaced reactivecircus/android-emulator-runner with the
same hand-provisioned manual boot e2e-preview uses.

Each leg now wraps its `./gradlew connectedDebugAndroidTest` in the #404
`timeout -k 30s $WEDGE_TIMEOUT` wrapper; on a hang (exit 124) capture_wedge
dumps the smoking gun (running test via TestRunner logcat, SIGQUIT thread
dumps of the app + instrumentation processes, service list / service check,
dumpsys activity+window, logcat tail) into a per-API-level
wedge-diagnostics-api<level> artifact (if-no-files-found: ignore so healthy
runs stay quiet).

Why it cannot re-hang the legs (the #406/90dfb18 revert reason): the wrapper
wraps ONLY the foreground gradle client, never the backgrounded emulator --
identical in shape to e2e-preview's run_shard. The reverted #404 wrapper wrapped
reactivecircus's emulator boot; #454 removed that. WEDGE_TIMEOUT reuses preview's
1200s: healthy matrix legs run ~8.3-12.0 min (whole job), well under 20 min, which
is itself well under this job's 50-min cap, so a genuine wedge is caught + captured
and a false trip on a healthy run is not possible.
2026-07-08 17:04:54 -05:00
JMR-dev b98ae99abc ci(path-filter): skip E2E for scripts/docs/.claude PRs via negated allow-list (#420)
The #399 E2E path-filter listed its safe paths as POSITIVE globs under
`predicate-quantifier: 'every'`. dorny/paths-filter's `every` makes the
per-file predicate "matches EVERY pattern", so `skippable` required a single
file to be under app/src/test AND docs AND scripts AND .claude AND be *.md at
once -- impossible. `skippable` was therefore always false and the full E2E
matrix ran for every PR, including the docs/scripts-only PRs it was meant to
skip (e.g. scripts-only #419).

Rewrite the filter the way dorny documents `every`: `non_skippable` lists the
same safe allow-globs, each NEGATED, so a file forces E2E iff it matches NONE
of them (i.e. it is outside the allow-list). e2e_needed = non_skippable, keeping
the fail-safe "run E2E unless every changed file is provably irrelevant" rule --
a mixed PR still runs the matrix.

Add an `e2e_harness` override so a change under .claude/skills/preflight/** (the
local instrumented-test harness) still forces E2E even though .claude/** is
otherwise skippable.

Skip: app/src/test/**, **/*.md, docs/**, scripts/**, .claude/** (minus preflight).
Run:  everything else -- app/**, *.gradle*, gradle/**, gradle/libs.versions.toml,
      app/schemas, app/proguard-rules.pro, .github/workflows/**, .github/scripts/**,
      .claude/skills/preflight/**. This ci.yml change itself runs the matrix.

Verified with PyYAML (valid) + a dorny-semantics simulation over representative
changed-file sets (scripts-only/docs/unit-test/.claude skip; app/androidTest/
gradle/ci.yml/preflight/.github-scripts run).

Closes #420
2026-07-08 15:22:03 -05:00
JMR-dev 2b958404bb ci(e2e): converge matrix emulator boot on e2e-preview manual boot
The matrix `e2e` job (API 29-36) relied on reactivecircus/android-emulator-
runner default boot, whose un-guarded, fatal `adb shell input keyevent 82`
races system_server binder republish on snapshot resume ("No service published
for: input") -- an intermittent boot race that flaked the merge queue and hit
BOTH the run and its retry once #446 let runs finally reach boot on API 33.

Replace the android-emulator-runner boot (snapshot-generate + run + retry steps
plus the AVD snapshot cache) with the e2e-preview job proven hand-provisioned
manual boot:
- avdmanager creates the AVD (google_apis/x86_64, pixel_2 -- kept in lockstep
  with testOptions.managedDevices in app/build.gradle.kts);
- a 2-attempt COLD boot loop (-no-snapshot), each with ONE bounded
  `timeout 300 adb wait-for-device shell wait-for-sys.boot_completed` so a stuck
  emulator fails fast instead of hanging to the 50-min job cap;
- a NON-FATAL `adb shell input keyevent 82 || true` unlock (kills the race) plus
  a boot_completed readiness gate before connectedDebugAndroidTest;
- emulator flags mirror e2e-preview; gradle retries once on a test failure.

Cold boot drops the AVD snapshot cache (snapshot resume is the documented root
cause of the race); the shared android-sdk-v1 cache and #446 hardened
pre-install step are untouched. No #404 wedge-capture wrapper (it was reverted
from this matrix job in 90dfb18 for hanging all 8 legs). Streamed logcat +
emulator.log + boot-diagnostics are uploaded for parity diagnosability.

Closes #448
2026-07-08 12:48:54 -05:00
JMR-dev 5c00a8da70 ci(e2e): route matrix emulator + system-image install through the #389-hardened installer
The E2E (33) matrix leg deadlocked the merge queue for ~2h when
android-emulator-runner's un-guarded "Create AVD and generate snapshot" step
died with "Error on ZipFile unknown archive" installing a corrupt Android
Emulator SDK zip. #389 hardened the platform/build-tools install (SHA-verify ->
reject-corrupt -> purge -> re-download) but left the emulator + system-image
install to the action, un-guarded.

Pre-install "emulator" + "system-images;android-<api>;google_apis;x86_64"
through setup_android_sdk.py before the emulator-runner steps, so a corrupt zip
is self-healed here and the action then finds both packages already installed
and skips its fragile fetch. Runs on both AVD-cache hit and miss (the emulator
binary + image live under the SDK root, not the ~/.android AVD-snapshot cache,
so they must be present for even a cached AVD to boot). Not added to the
android-sdk-v1 cache (kept small); re-install is a fast sdkmanager no-op when
already present.

The e2e-preview (API 37) job already routes emulator + system image through the
hardened installer, so no change there. setup_android_sdk.py already handles
these package ids generically; add a unit assertion pinning the matrix's
google_apis/x86_64 id to the correct purge path.

Closes #443
2026-07-08 10:39:38 -05:00
JMR-dev 90dfb189e6 fix(ci): drop matrix E2E wedge-capture that hangs all 8 legs
The `timeout -k 30s` wrapper + `capture_wedge()` added in 2f32657 for the
matrix `e2e` job (API 29-36) reproducibly wedges every leg, while the
manually-provisioned API 37 preview shard running the identical capture
logic passes. Revert the two matrix "Run E2E tests" steps' `script:`
blocks to main's plain script (just the backgrounded logcat stream +
`./gradlew connectedDebugAndroidTest`) and drop the now-dead "Upload wedge
diagnostics" step from the matrix job.

Kept untouched: the job-level `timeout-minutes: 50` backstop added in
27ede55, and the entire `e2e-preview` job (its own capture_wedge/watchdog
and wedge-diagnostics-api37-preview-shard* upload are unaffected).
2026-07-07 11:29:41 -05:00
JMR-dev 27ede55dbd ci(e2e): add job-level timeout to matrix E2E job (#404)
The matrix E2E job (api-level 29-36) had no timeout-minutes, so a wedge
hangs until GitHub's 6-hour default instead of being force-killed. The
sibling e2e-preview job already sets timeout-minutes: 35. A normal
matrix run is ~15-20 min and a retry-inclusive run ~40 min, so set
timeout-minutes: 50 to give headroom above the in-step wedge-capture
timeout (1200s) while still bounding worst-case runtime.
2026-07-07 07:46:33 -05:00
JMR-devandClaude Opus 4.8 2f32657aff ci: capture thread-dump + service state on an E2E wedge to prove the root cause (#404)
E2E legs intermittently WEDGE (hang) with no fast-fail until the job force-kill,
and GitHub's post-force-kill step behavior is unreliable, so #388's diagnostics
don't reliably capture the wedge — and don't capture wedge-specific state anyway.

Wrap the `connectedDebugAndroidTest` run (both the `e2e` matrix first-attempt +
retry, and each `e2e-preview` shard) in an explicit `timeout -k 30s 1200`
(20 min) — comfortably above a normal run (~13-15 min), well below the hard cap —
so a wedge trips the wrapper (exit 124), NOT the force-kill, GUARANTEEING the
capture runs while the emulator is still alive. On 124, capture_wedge grabs the
smoking gun into a `wedge-diagnostics-api<level>` artifact: the running/last test
(logcat TestRunner), SIGQUIT (kill -3) thread dumps of the app + instrumentation
processes (ART -> logcat + /data/anr), dumpsys activity/window, `service list` +
`service check input/window/activity` (the boot-race crux), sys.boot_completed +
init.svc.* state, the snapshot cache-hit note, and accel/kvm/mem/disk. Then it
exits with the real status so #388's diagnostics + the existing retry still fire;
a normal run finishes before the wrapper and is unaffected.

EVIDENCE ONLY — no boot-readiness guard/fix (maintainer: prove the cause first).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 07:46:33 -05:00
JMR-devandClaude Opus 4.8 e668330151 ci: skip the E2E matrix for test-only/docs PRs via a paths-filter + skip-tolerant gate (#399)
Add a cheap `changes` job (dorny/paths-filter v4, pinned SHA) that sets
e2e_needed=false only when EVERY changed file is in a safe allow-list
(app/src/test/**, **/*.md, docs/**, scripts/**, .claude/**); anything
else -- or any non-pull_request event -- defaults to true (conservative,
"err toward running E2E").

Gate `e2e` and `e2e-preview` on needs.changes.outputs.e2e_needed so the
whole matrix runs or skips together, and rewrite the `ci-passed` gate:
it now BLOCKS on changes!=success, any of traffic-control-tests /
static-analysis / debug-build / unit-tests !=success, or e2e/e2e-preview
==failure|cancelled -- while TOLERATING an intentional e2e/e2e-preview
'skipped'. So test-only/docs PRs go green on the fast gate, a real E2E
failure/cancel still blocks, and a broken filter (changes!=success)
still blocks. Branch protection ("CI passed") context is unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 20:55:58 -05:00
JMR-devandClaude Opus 4.8 aa628831be ci: retry + cache Android SDK/emulator setup to survive corrupt-zip sdkmanager failures (#389)
The dominant merge-blocking flake was the "Set up Android SDK" step
(android-actions/setup-android v4.0.1) dying BEFORE the emulator starts:

    Wrong version in preinstalled sdkmanager
    Warning: ... preparing SDK package Android Emulator: Error reading Zip
    content from a SeekableByteChannel.
    Error: The process '.../sdkmanager' failed with exit code 1

Root cause: the action's default cmdline-tools version (20.0) rarely matches the
runner image's preinstalled one, so it logs "Wrong version in preinstalled
sdkmanager" and re-fetches cmdline-tools with NO checksum; it then runs its
default `sdkmanager tools platform-tools` install. Any of those downloads can be
a corrupt/truncated zip, which sdkmanager turns into an un-retried exit 1. v4.0.1
is the latest release, so this is fixed by configuration + hardening, not a bump.

Harden with verify -> reject -> retry, never trusting sdkmanager's exit code
alone, via a new stdlib-only helper .github/scripts/setup_android_sdk.py:

- bootstrap: download the pinned cmdline-tools zip, verify size + SHA-256
  (authoritative pin, cross-checked against Google's published SHA-1), and
  install it to $ANDROID_SDK_ROOT/cmdline-tools/20.0 -- the exact path
  setup-android probes first, so the action reuses the verified tree and never
  does its own unverified "Wrong version" re-download. A mismatch (corrupt OR
  wrong version) deletes the bad zip + any half-extracted dir and re-downloads.
- install: sdkmanager --install with retry + backoff; on a corrupt package zip it
  purges the partial/corrupt package dir (and sdkmanager's temp dirs) before
  retrying, forcing a fresh download instead of a re-read.
- setup-android now runs with packages: "" (no flaky tools/platform-tools
  install) and cmdline-tools-version: "14742923" (reuse the verified bootstrap).
- actions/cache restore + success-gated save so only a verified SDK is ever
  cached (integrity gates the cache); shrinks the re-download/corruption surface.

Applied to every SDK-setup job (debug-build, unit-tests, static-analysis, e2e
matrix, e2e-preview). Emulator BOOT logic, #372 API-37 sharding, and #388
diagnostics are untouched. Pure-logic helpers are unit-tested
(test_setup_android_sdk.py, run by the traffic-control-tests job).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 18:59:05 -05:00
JMR-devandClaude Opus 4.8 187a8effb0 ci(e2e): capture logcat + emulator/system diagnostics across the E2E matrix (#387)
The API 29-36 `e2e` matrix uploaded only its test report, so an emulator
flake or a red leg (e.g. `E2E (31)` dying on a bare `sdkmanager` exit 1)
left nothing to diagnose. Bring the #334 API-37 diagnostics to the matrix,
inline (no changes to `e2e-preview`, which PR #372 is restructuring):

- Stream `adb logcat -v time` to `$RUNNER_TEMP/logcat-api<level>.txt` at the
  top of both the "Run E2E tests" and retry reactivecircus steps (emulator is
  booted there); backgrounded so gradle stays the exit-status-bearing command.
- New `if: failure()` step dumps device + runner state (adb devices, logcat
  tail, emulator -accel-check, /dev/kvm, free -h, df -h) to the step log and a
  diagnostics file; every probe guarded with `|| true`.
- New `if: always()` upload-artifact (same pinned v7 SHA) `e2e-diagnostics-api<level>`
  carries the logcat + diagnostics files, `if-no-files-found: warn`.
- Make "Install SDK platform and build-tools" diagnosable: bounded 3x retry with
  backoff for a transient sdkmanager failure, and print `--list_installed` on a
  hard failure instead of a bare exit 1.

Keeps reactivecircus/android-emulator-runner and the existing boot-race retry.
Additive/diagnostic only; no boot-affecting flags change.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 15:53:54 -05:00
JMR-devandClaude Opus 4.8 08ee9e7abc ci: productionize API 37 preview E2E sharding (adopt N=2)
The spike commits already implemented the N=2 shard matrix (numShards/
shardIndex via -Pandroid.testInstrumentationRunnerArguments.*), per-shard
test-retry parity, adb start-server before the boot loop, and shard-suffixed
artifact names. This drops the SPIKE / DRAFT "do not merge as-is" framing from
the ci.yml comments and reframes docs/perf/api37-e2e-sharding-spike.md from a
feasibility spike into the adopted design, so the change is mergeable as-is.

Also fixes the doc's section 3a example, which showed 1-based shardIndex values
[1, 2]; shardIndex is 0-based (0..numShards-1) and the implementation correctly
uses matrix.shard: [0, 1] -- [1, 2] would run an empty bucket and silently drop
half the suite.

Fan-in unchanged and verified: ci-passed still lists e2e-preview once; GHA
matrix aggregation makes its result `failure` if either shard fails, so both
shards must pass for the gate to go green. Branch protection requires the
"CI passed" context (not the per-leg "E2E (API 37 preview) (N)" check names),
so no branch-protection change is needed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 14:52:48 -05:00
JMR-devandClaude Opus 4.8 0e9f54e686 ci(spike): retry parity + adb start-server for API 37 shard PoC
Folds the #370 root-cause finding into the spike. The stable `e2e` matrix
retries its test run once; `e2e-preview` runs connectedDebugAndroidTest exactly
once, so a flaky test self-heals on API 29-36 but wedges the required gate on
API 37 (e.g. #370's SignaturesScreenTest teardown race).

Doc: adds risk item 9 (retry-parity gap + its sharding interaction — per-test
flake is NOT amplified by sharding unlike boot flake, and a per-shard retry
costs only B + T/N; framed mitigation-not-fix) and two §6 recommendations
(retry parity, mirrored into api37_e2e.py; adb start-server before the boot
loop).

PoC (ci.yml): per-shard single test retry (::warning:: on retried-but-passed)
+ adb start-server before the boot loop. The api37_e2e.py retry mirror stays a
documented recommendation (local path needs a real-emulator validation this
spike did not boot). Still DRAFT, not auto-merged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 14:40:02 -05:00
JMR-devandClaude Opus 4.8 66da643249 ci(spike): PoC shard API 37 preview E2E + feasibility doc
Feasibility spike for sharding the e2e-preview job (the hand-provisioned
API 37 / google_apis_ps16k 16 KB-page emulator), CI's longest leg
(~16.4-17.6 min). docs/perf/api37-e2e-sharding-spike.md breaks the leg into
fixed overhead B ~8.3 min (setup + boot + Gradle daemon/config/compile/install)
vs parallelizable test execution T ~8.8 min, models B + T/N for N=2/3/4, and
recommends N=2 (~17.1 -> ~12.7 min, ~28% off the critical path) capped by the
API 30 matrix wall (~12.0 min) beyond N=3.

DRAFT PoC (do NOT merge as-is): converts e2e-preview to a strategy.matrix.shard
[0, 1] fan-out passing AndroidJUnitRunner numShards/shardIndex through the
existing -Pandroid.testInstrumentationRunnerArguments.* channel (no GMD, no
orchestrator, no Gradle change). Artifact names gain a shard suffix;
ci-passed still lists e2e-preview once (matrix fan-in keeps the single gate).
Local preflight stays single-emulator. Relates to #258.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 14:31:40 -05:00
JMR-devandClaude Opus 4.8 939906986b ci: extract traffic-control into its own workflow file
Move the traffic-control (runner-priority orchestration) job verbatim out of
.github/workflows/ci.yml into a new standalone workflow,
.github/workflows/traffic-control.yml, so the heavy CI jobs no longer depend
on it. The job's YAML (name, runs-on, timeout-minutes, permissions, env,
steps) and its documentation comment move unchanged; the decision core
.github/scripts/traffic_control.py is untouched and still unit-tested by the
traffic-control-tests job in ci.yml.

In ci.yml: removed the traffic-control job, dropped needs: traffic-control
from the five heavy jobs (static-analysis, debug-build, unit-tests, e2e,
e2e-preview) and from traffic-control-tests (its only needs, which would
otherwise dangle at a now-deleted job), and updated the now-stale header and
ci-passed comments to point at the extracted workflow.

The new workflow will be disabled pending a rebuild as a published GitHub
Action.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 14:48:03 -05:00
JMR-devandClaude Opus 4.8 05d06eb45b ci: traffic-controller owns CI triggering (GITHUB_TOKEN updates, priority-ordered dispatch)
End the merge cascade and give the traffic-controller ownership of CI *triggering*.

- autoupdate.yml updates PR branches with the built-in GITHUB_TOKEN instead of a PAT,
  so an update push no longer auto-retriggers CI (GitHub's anti-recursion rule) — the
  cascade (every merge re-runs every PR, cancel-in-progress thrashing them) is gone.
- New scheduler ci-trigger.yml -> traffic_control.py --mode trigger (re-)triggers CI
  for the highest-priority PR(s) whose head SHA has absent/stale checks, a few at a
  time (inflight cap), in the existing P0-P9 / broken-draft priority order — a
  poor-man's merge queue reusing the priority core. It runs after autoupdate finishes
  (workflow_run, race-free) plus a cron backstop plus manual dispatch.
- Triggering uses workflow_dispatch, which is EXEMPT from anti-recursion, so the
  built-in GITHUB_TOKEN (actions: write) starts the run — NO PAT / secret change needed.
- ci.yml gains a workflow_dispatch trigger (pr/head_sha/reason inputs) and a per-PR
  concurrency group unifying pull_request and dispatch runs; its on: pull_request path
  is kept so brand-new PRs, human pushes, and fork PRs always get CI (fail-open).

Pure select_triggers / classify_sha_runs decision core added to traffic_control.py with
24 new unit tests (priority order, oldest-first fairness, inflight cap, fork skip, P0
bypass+preempt, head-SHA needy classification, and a liveness/anti-starvation simulation).

Closes #349

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 01:27:40 -05:00
JMR-devandClaude Opus 4.8 8b89219734 ci: run traffic-controller unit tests as a gate job
Adds a fast traffic-control-tests job (ubuntu, actions/checkout +
actions/setup-python, no emulator/Gradle) that runs the 37 pure-stdlib
unit tests for .github/scripts/traffic_control.py on every PR, and
wires it into ci-passed's needs so a regression blocks merge instead
of only being caught locally.

Closes #346

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 22:35:39 -05:00
JMR-devandClaude Opus 4.8 a34db54b97 fix(ci): let any strictly-higher PR reclaim a broken/draft run (#342)
Restore (and extend to drafts) the old bash's broken-reclaim behaviour that the
initial Python refactor had dropped. runs_to_cancel now cancels an OTHER PR's
active/queued runs when EITHER:
  (a) THIS PR is P0 and that PR is strictly-lower (reclaim every lower runner); OR
  (b) that PR is broken/draft (effective priority 10) and THIS PR is strictly-higher
      (effective priority < 10) — a wasted run any ready PR may reclaim.

P1-P9 still never bump a *normal* (non-broken/draft) lower run; a broken/draft PR
(P10) preempts nothing (nothing is strictly-lower than the bottom, and the
equal-or-higher invariant means a P10 never cancels another P10). Self / main-push /
equal-or-higher invariants unchanged.

Updates the module docstring + ci.yml comments (the "only P0 preempts" wording
becomes: P0 preempts everything strictly-lower; additionally, any strictly-higher PR
preempts a broken/draft run) and the job step/permission/needs comments. Adds unit
tests: P3 reclaims a broken P10 run and a draft P10 run; P3 does not bump a normal P5
run; a P10 self preempts nothing; plus an end-to-end P5-reclaims-draft-then-waits
scenario. 37 unit tests pass; ci.yml parses clean; --dry-run shows a P3 cancelling a
draft (and broken) run while still yielding to a higher P1.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 22:19:01 -05:00
JMR-devandClaude Opus 4.8 8b7f895d6a ci: extract traffic-controller into a testable Python module (#342)
The priority-based runner orchestration ("traffic-control") lived as a large
inline-bash step in ci.yml — a two-pass preemption + hold-back script that was
effectively untestable in YAML. Move it into .github/scripts/traffic_control.py,
structured as a pure decision CORE + a thin gh-I/O SHELL:

* Pure functions (no network/clock/subprocess), unit-testable in isolation:
  - effective_priority(pr): lowest-numbered P0-P9, default P5; broken OR draft => 10.
  - runs_to_cancel(this_pr, all_prs, self_run_id): PASS 1 — run ids to cancel,
    empty unless THIS PR is P0; only strictly-lower running/queued runs; never self
    (by number or run id), never equal-or-higher.
  - wait_blockers(this_pr, all_prs): PASS 2 — yield to any strictly-higher PR with
    an active/queued run, and to same-level peers ordered ahead (running-first,
    then oldest createdAt). Empty => proceed.
* Shell (run_live): gathers the snapshot via gh, applies cancels, runs the bounded
  hold-back poll loop; always exits 0. --dry-run feeds the core a snapshot JSON and
  prints decisions with zero network.
* No jq/bash dependency (cross-platform, per the repo's Python-stdlib convention).

ci.yml's traffic-control job now checks out the repo and runs the module. Job
permissions gain `contents: read` (for checkout) alongside the existing
`actions: write` / `pull-requests: read`; env and downstream `needs:` wiring
unchanged; step stays `continue-on-error`.

33 stdlib unittest cases cover priority resolution, P0-only preemption, the
self/main/equal-or-higher invariants, and the same-level running-first/oldest
ordering.

Behaviour is preserved except the ticket's refinements: (1) drafts now count as
P10 (bottom); (2) an explicit same-level running-first-then-oldest tiebreaker; and
(3) per the ticket's order-of-operations, ONLY P0 preempts — the old bash also let
any higher-priority PR cancel a `broken` target's run, which no longer happens
(a broken/draft run is only cancelled by a P0, via the same strictly-lower rule).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 22:08:13 -05:00
JMR-devandClaude Opus 4.8 9e771a9590 ci(e2e): boot diagnostics + verbose/debug logging on the API 37 preview job
The API 37 preview E2E boot has flaked twice (#285, #333) with only
"did not boot within 300s" and no root-cause signal. Add rich boot
diagnostics by default, kept in parity between CI and the local
hand-provisioning script (api37_e2e.py):

- Launch the emulator with `-verbose -debug init,avd_config,kernel`
  (diagnostics only; no boot-affecting flag changed), still redirecting
  to $EMU_LOG.
- Stream `adb logcat -v time` to a file from the moment the device
  registers (via `adb wait-for-device logcat`, backgrounded).
- On a boot timeout, dump accel-check, /dev/kvm presence, GPU mode,
  free mem/disk, the AVD config.ini and the emulator.log tail; CI writes
  these to a boot-diagnostics file, the local script prints them.
- CI uploads emulator.log + logcat.txt + boot-diagnostics.txt as an
  artifact with `if: always()` so they survive a timeout/cancel, and
  prints a concise summary (accel/KVM status + last 50 lines of
  emulator.log) to the step log.

The existing 2-attempt boot retry + boot-completed wait loop are
unchanged; the diagnostics are additive.

Closes #334

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 21:09:28 -05:00
JMR-devandClaude Opus 4.8 2e0eaef4e9 feat(ci): no-regression JVM coverage gate scoped to the testable surface
Scope :app:jacocoTestReport's denominator to the JVM-testable surface and
add a :app:jacocoTestCoverageVerification no-regression gate that shares the
same classDirectories/executionData/sourceDirectories, wired into both the
`check` lifecycle task and CI's unit-test job (part of the `CI passed` gate).

Excluded from the denominator (structurally unreachable from a JVM unit
test): Compose screen/component render code, Android framework entry points
(*Activity/*Service/Application/*BackupAgent), Hilt DI (**/di/**), and the
src/debug cold-open probe. Kept in scope: ViewModels, repositories, mappers,
DAOs, utils, richtext, mail, reporting logic, and the six WorkManager Workers.

Corrects PR #292, which excluded **/*Worker*: SyncWorker, BackfillWorker,
PruneWorker, SendWorker, ReportPurgeWorker and ReportUploadWorker are all
directly unit-tested, so they stay counted in both numerator and denominator
(only their Hilt wiring, WorkManagerModule, is excluded, via **/di/**).

Baseline: 80.21% line (4838/6032). Floor: 0.79 (~1.2% headroom) so ordinary
noise doesn't red-flag it while a real drop fails. Manual ratchet for now:
bump the floor up in the same PR when coverage rises materially.

Closes #251
Closes #292

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 11:27:17 -05:00
JMR-devandClaude Opus 4.8 249381b249 ci(runners): P0 & broken-target preemption; P1-P9 yield without bumping
Only P0 preempts in-progress runs (emergency reservation). P1-P9 no longer
cancel lower-priority runs; instead traffic-control holds back (bounded poll,
kept under timeout-minutes) while strictly-higher-priority PRs still have
active/queued CI runs, so their heavy jobs reach the runner queue first.

New `broken` label forces effective priority below P9 (sentinel 10): a broken
PR never preempts (even if also labelled P0 -- broken wins) and always yields,
and because its run is wasted, ANY higher-priority PR (not just P0) may cancel
its in-progress run to reclaim the runner. Net rule: a strictly-lower run is
cancelled iff (self is P0) OR (target is broken); otherwise yield.

All existing safety preserved: never main/push runs, never our own run, never
an equal-or-higher-priority PR; PR-controlled strings via env/jq only;
continue-on-error + set +e + always exit 0; traffic-control stays a
non-required best-effort job and ci-passed is unchanged.

Validated with actionlint and a mocked-gh + fake-clock logic harness.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 17:08:37 -05:00
JMR-devandClaude Opus 4.8 0e3ffd58ef ci(runners): priority-based runner orchestration via P0–P9 labels
Add a lightweight `traffic-control` job that runs first (the heavy
build/E2E jobs `needs:` it) and preempts contended runners by PR
priority. It reads the triggering PR's P0–P9 label (P0 = highest,
P9 = lowest; default P5 when unlabeled) and cancels the in-progress /
queued CI runs of strictly-lower-priority OTHER open PRs, freeing their
runners for the higher-priority PR.

Safety: never cancels main/push runs, the PR's own run, or an
equal-or-higher-priority PR — only strictly-lower-priority OTHER open
PRs' active CI runs. The job is best-effort (every gh call guarded,
always exits 0, step is continue-on-error) and is NOT part of the
`CI passed` merge gate. `ci-passed` now also treats a `skipped` heavy
job as a gate failure, so a (should-never-happen) traffic-control
failure blocks the merge fail-safe rather than passing it untested.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 16:02:45 -05:00
JMR-devandClaude Opus 4.8 535cbcb49a build(coverage): finish wiring JaCoCo unit-test coverage reporting
Completes the crash-interrupted #192 WIP (app/build.gradle.kts already had a
jacocoTestReport task and toolVersion pin recovered onto build-192-jacoco):

- Move the JaCoCo tool version into gradle/libs.versions.toml instead of a
  hardcoded string in app/build.gradle.kts, matching how every other plugin
  version in this repo is sourced.
- Fix the generated-code exclusion list against the real compileDebugKotlin
  output (verified by inspecting the compiled class tree): Room's
  KSP-generated `_Impl` DAOs/database and the Compose compiler's per-file
  ComposableSingletons holders are the only generated code that actually
  lands in classDirectories, since Hilt/Dagger's generated Java and AGP's
  BuildConfig/R/Manifest are compiled by a separate javac task this report
  never reads. Drop the blanket `**/*$$*` exclude the WIP had — it was
  silently discarding ~200 real classes' worth of coverage on Kotlin's own
  `$$inlined$` synthetic classes (e.g. Flow.map { ... } transforms in the
  repositories), which is hand-written logic, not generated boilerplate.
- Add Hilt_*/Dagger* prefix patterns so the (currently inert,
  belt-and-suspenders) Hilt exclusions are actually correct if the
  classDirectories scope ever changes.
- Add a minimal CI step to the existing unit-tests job that runs
  jacocoTestReport and uploads the XML+HTML report as a build artifact.
  No coverage threshold gate yet (a jacocoTestCoverageVerification rule is
  a natural follow-up once there's a baseline).
- Document the new :app:jacocoTestReport task in CLAUDE.md.

Verified on JDK 21: fast gate (assembleDebug, testDebugUnitTest,
compileDebugAndroidTestKotlin, lintDebug, ktlintCheck, detekt) plus
jacocoTestReport all pass, from both a warm and a `clean` build. The
report shows real signal (30% instruction / 38% line coverage) with no
generated classes leaking in.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 12:23:40 -05:00
JMR-devandClaude Opus 4.8 68a225ed9c ci(e2e): retry the API-29 emulator boot to absorb the android-emulator-runner keyevent race
The matrix E2E (29) job intermittently fails (~2%, API-29 only) in
reactivecircus/android-emulator-runner's un-guarded, fatal post-boot
`adb shell input keyevent 82`: on snapshot resume sys.boot_completed=1 is
restored before system_server republishes the `input` binder service, so
the job aborts before Gradle runs with "No service published for: input"
(fast-fail ~1m43s). Proven on run 28667366203.

Make the "Run E2E tests" step non-fatal (id + continue-on-error) and add a
guarded second attempt (if steps.e2e.outcome == 'failure'). Two independent
boots drop the race to ~0.04%; a genuine failure on both attempts still
fails the job (outcome, not conclusion). Definitive manual-boot fix: #218.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 10:23:36 -05:00
JMR-devandClaude Fable 5 d0a5ccb10d ci: auto-update armed PRs; drop dead merge_group trigger
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-02 12:48:23 -05:00
JMR-devandClaude Fable 5 c84b0605cf ci: trigger on merge_group so the GitHub merge queue can gate PRs
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-02 10:21:47 -05:00
JMR-devandClaude Opus 4.8 0754ad4ebf ci: enforce ktlint + detekt on pull requests
Add a `static-analysis` job (JDK 21 + Android SDK) that runs
`:app:ktlintCheck :app:detekt` and uploads the reports, and add it to the
`ci-passed` aggregating gate so lint regressions block merges like the
other checks.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 21:44:12 -05:00
JMR-devandClaude Opus 4.8 7987333149 ci: require the API 37 preview E2E job to merge
The custom-provisioned API 37 preview emulator has been stable, so fold its E2E
job into the aggregating "CI passed" gate's needs. Because branch protection
requires only that single check, no settings change is needed.

Drop the now-inaccurate "non-blocking" wording from the job name and comments.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 18:17:15 -05:00
JMR-devandClaude Opus 4.8 fe0593dba1 ci: pin ANDROID_AVD_HOME so the API 37 preview emulator finds its AVD
The preview emulator failed every boot with "Unknown AVD name [api37]" and
exited immediately (no device -> the bounded wait timed out). avdmanager had
created the AVD under $ANDROID_SDK_HOME/.android/avd while the emulator searched
$ANDROID_SDK_HOME/avd and $HOME/.android/avd. Pin ANDROID_AVD_HOME to one path
both tools use, carry it to the boot step via $GITHUB_ENV, and list AVDs after
creation to verify.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 12:59:35 -05:00
JMR-devandClaude Opus 4.8 b8086f89df ci: make API 37 preview emulator boot fail-fast and diagnosable
The custom-provisioned preview job hung in 'Boot emulator and run E2E' until
the 35-min cap: adb wait-for-device had no timeout and the emulator's own
output was never captured, so a failed boot was both invisible and unbounded.
Bound the wait with a single 300s timeout covering device registration + full
boot, retry once, capture the emulator log, and dump it (plus logcat) on
failure.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 12:45:41 -05:00
JMR-devandClaude Opus 4.8 957da0c46f ci: run debug-build and unit-tests on x86_64
Linux-arm64 runners can't set up the SDK here: android-actions/setup-android's
sdkmanager fails (exit 1) on the android-37.0 preview platform, and the emulator
package has no arm64-Linux build. These are pure build/JVM-unit jobs whose
results are host-arch-independent, so x86_64 loses no device coverage (real
arm64 ABI coverage would require arm64 emulators, i.e. macOS hosts).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 12:28:55 -05:00
JMR-devandClaude Opus 4.8 13a32df873 Add IMAP folder navigation drawer + per-API E2E CI matrix
Folder view: a left navigation drawer lists each account's IMAP folders;
tapping one browses and caches that folder's mail. IMAP UIDs are unique only
within a folder, so message identity, the fetch/read/flag/delete paths, sync,
and the Room cache all became folder-aware (id = "accountId:folder:uid"; new
`folder` column; schema v7->v8). Standard folders (Inbox/Sent/Drafts/Spam/Trash/
Archive) surface with friendly names + icons via RFC 6154 SPECIAL-USE attributes
with a case-insensitive name fallback; the multi-account drawer adds an account
switcher and a unified "All Inboxes". INBOX stays the only auto-synced,
IDLE-watched, notifying folder; other folders sync on demand.

Lower minSdk 33 -> 29 for a rolling ~7-year Android support window; guard the
API-33 POST_NOTIFICATIONS runtime request accordingly.

Tests and CI:
- Bump espresso-core 3.6.1 -> 3.7.0 so Compose UI tests run on API 37
  (3.6.1's InputManagerEventInjectionStrategy reflects a removed hidden method).
- New coverage across layers: FolderRoleTest, ImapClientTest folder cases,
  MailboxViewModelTest, MailRepositoryImplTest folder routing, a FolderDrawer
  Compose UI test, and LibreMailDatabaseTest folder DAO/reconcile tests.
- Gradle Managed Devices + a CI E2E matrix over every API 29-36; a single
  "CI passed" gate job fans in all jobs and is required by branch protection.
- Non-blocking, custom-provisioned API 37 (preview) E2E job with image caching.
- Build + unit-test jobs run on arm64 (ubuntu-24.04-arm); emulators stay x86_64.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-30 12:23:00 -05:00
JMR-devandClaude Opus 4.8 669946ea34 Add CI/release workflows and first Compose UI tests
GitHub Actions (every action pinned to its commit SHA, version in a comment):
- ci.yml, on pull_request to main, runs three jobs: a debug build, the
  unit tests, and the instrumented suite on a headless emulator.
- release.yml, on workflow_dispatch, builds the release APK, archives the
  source as zip + tar.gz, and publishes a GitHub release.

Compose UI tests (app/src/androidTest), driving the real screens with
fake-backed view models so they need no network, database, or Hilt graph:
- ManualSetupScreen: submit-button validation, advanced-options toggle,
  and add-account success/failure.
- ComposeScreen: send-button enablement and the send -> close flow.
- LibreMailBottomBar: tab rendering and selection callback.

Mark gradlew executable (100755) so it runs on the Linux CI runners.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 21:36:41 -05:00