Root cause: intermittently the launched activity window has has-window-focus=false
for the WHOLE instrumented run, so Espresso's RootViewPicker (onView().check(),
Intents.intended(), pressBack(), focus-dependent clipboard) times out after 10s and
fails EVERY focus-dependent test at once while the ~280 pure-Compose semantics tests
(which don't need window focus) pass. A failing E2E (35) leg's logcat (PR #470, run
28985259521) shows has-window-focus=true ZERO times across the whole session and both
the first attempt and the once-retry fail identically -- a persistent environmental
state, not a per-test transient. The prior mitigation, a single fire-and-forget
`adb shell input keyevent 82` (MENU) right after boot, is too weak: MENU no longer
dismisses the modern (API 30+) keyguard and, delivered before SystemUI/keyguard comes
up, is simply dropped -- so the insecure keyguard / non-interactive display persists
and no app window ever takes focus.
Fix: a single shared helper, .github/scripts/emulator_focus_gate.py, invoked
identically by BOTH E2E jobs (the e2e API 29-36 matrix AND e2e-preview API 37) and by
the local preflight runners (local_instrumented.py / api37_e2e.py), so it cannot drift.
It wakes the display (KEYCODE_WAKEUP), dismisses + disables the keyguard
(wm dismiss-keyguard, locksettings set-disabled true), keeps the screen on
(svc power stayon true + max screen_off_timeout), zeroes the animation scales, then
polls dumpsys power/window until the device is interactive AND a real window holds
input focus (mCurrentFocus non-null) -- re-nudging each iteration -- before the suite
runs. Applied uniformly, this also gives e2e-preview the animation-disable the matrix
already had. The gate is soft (bounded wait, then proceeds with a ::warning:: and the
final device state) and non-fatal (`|| true`), preserving #454's guarantee that the
unlock never aborts the boot; it leaves #454's manual boot, #460's path-filter and
#464's wedge-capture untouched.
The pure readiness parser is unit-tested by test_emulator_focus_gate.py (run by the
traffic-control-tests job). Determinism is validated by this PR's own matrix run.
Closes#468
Mirror the e2e-preview job's proven wedge-capture onto the API 29-36 matrix
legs, now that #454 replaced reactivecircus/android-emulator-runner with the
same hand-provisioned manual boot e2e-preview uses.
Each leg now wraps its `./gradlew connectedDebugAndroidTest` in the #404
`timeout -k 30s $WEDGE_TIMEOUT` wrapper; on a hang (exit 124) capture_wedge
dumps the smoking gun (running test via TestRunner logcat, SIGQUIT thread
dumps of the app + instrumentation processes, service list / service check,
dumpsys activity+window, logcat tail) into a per-API-level
wedge-diagnostics-api<level> artifact (if-no-files-found: ignore so healthy
runs stay quiet).
Why it cannot re-hang the legs (the #406/90dfb18 revert reason): the wrapper
wraps ONLY the foreground gradle client, never the backgrounded emulator --
identical in shape to e2e-preview's run_shard. The reverted #404 wrapper wrapped
reactivecircus's emulator boot; #454 removed that. WEDGE_TIMEOUT reuses preview's
1200s: healthy matrix legs run ~8.3-12.0 min (whole job), well under 20 min, which
is itself well under this job's 50-min cap, so a genuine wedge is caught + captured
and a false trip on a healthy run is not possible.
The #399 E2E path-filter listed its safe paths as POSITIVE globs under
`predicate-quantifier: 'every'`. dorny/paths-filter's `every` makes the
per-file predicate "matches EVERY pattern", so `skippable` required a single
file to be under app/src/test AND docs AND scripts AND .claude AND be *.md at
once -- impossible. `skippable` was therefore always false and the full E2E
matrix ran for every PR, including the docs/scripts-only PRs it was meant to
skip (e.g. scripts-only #419).
Rewrite the filter the way dorny documents `every`: `non_skippable` lists the
same safe allow-globs, each NEGATED, so a file forces E2E iff it matches NONE
of them (i.e. it is outside the allow-list). e2e_needed = non_skippable, keeping
the fail-safe "run E2E unless every changed file is provably irrelevant" rule --
a mixed PR still runs the matrix.
Add an `e2e_harness` override so a change under .claude/skills/preflight/** (the
local instrumented-test harness) still forces E2E even though .claude/** is
otherwise skippable.
Skip: app/src/test/**, **/*.md, docs/**, scripts/**, .claude/** (minus preflight).
Run: everything else -- app/**, *.gradle*, gradle/**, gradle/libs.versions.toml,
app/schemas, app/proguard-rules.pro, .github/workflows/**, .github/scripts/**,
.claude/skills/preflight/**. This ci.yml change itself runs the matrix.
Verified with PyYAML (valid) + a dorny-semantics simulation over representative
changed-file sets (scripts-only/docs/unit-test/.claude skip; app/androidTest/
gradle/ci.yml/preflight/.github-scripts run).
Closes#420
The matrix `e2e` job (API 29-36) relied on reactivecircus/android-emulator-
runner default boot, whose un-guarded, fatal `adb shell input keyevent 82`
races system_server binder republish on snapshot resume ("No service published
for: input") -- an intermittent boot race that flaked the merge queue and hit
BOTH the run and its retry once #446 let runs finally reach boot on API 33.
Replace the android-emulator-runner boot (snapshot-generate + run + retry steps
plus the AVD snapshot cache) with the e2e-preview job proven hand-provisioned
manual boot:
- avdmanager creates the AVD (google_apis/x86_64, pixel_2 -- kept in lockstep
with testOptions.managedDevices in app/build.gradle.kts);
- a 2-attempt COLD boot loop (-no-snapshot), each with ONE bounded
`timeout 300 adb wait-for-device shell wait-for-sys.boot_completed` so a stuck
emulator fails fast instead of hanging to the 50-min job cap;
- a NON-FATAL `adb shell input keyevent 82 || true` unlock (kills the race) plus
a boot_completed readiness gate before connectedDebugAndroidTest;
- emulator flags mirror e2e-preview; gradle retries once on a test failure.
Cold boot drops the AVD snapshot cache (snapshot resume is the documented root
cause of the race); the shared android-sdk-v1 cache and #446 hardened
pre-install step are untouched. No #404 wedge-capture wrapper (it was reverted
from this matrix job in 90dfb18 for hanging all 8 legs). Streamed logcat +
emulator.log + boot-diagnostics are uploaded for parity diagnosability.
Closes#448
The E2E (33) matrix leg deadlocked the merge queue for ~2h when
android-emulator-runner's un-guarded "Create AVD and generate snapshot" step
died with "Error on ZipFile unknown archive" installing a corrupt Android
Emulator SDK zip. #389 hardened the platform/build-tools install (SHA-verify ->
reject-corrupt -> purge -> re-download) but left the emulator + system-image
install to the action, un-guarded.
Pre-install "emulator" + "system-images;android-<api>;google_apis;x86_64"
through setup_android_sdk.py before the emulator-runner steps, so a corrupt zip
is self-healed here and the action then finds both packages already installed
and skips its fragile fetch. Runs on both AVD-cache hit and miss (the emulator
binary + image live under the SDK root, not the ~/.android AVD-snapshot cache,
so they must be present for even a cached AVD to boot). Not added to the
android-sdk-v1 cache (kept small); re-install is a fast sdkmanager no-op when
already present.
The e2e-preview (API 37) job already routes emulator + system image through the
hardened installer, so no change there. setup_android_sdk.py already handles
these package ids generically; add a unit assertion pinning the matrix's
google_apis/x86_64 id to the correct purge path.
Closes#443
The `timeout -k 30s` wrapper + `capture_wedge()` added in 2f32657 for the
matrix `e2e` job (API 29-36) reproducibly wedges every leg, while the
manually-provisioned API 37 preview shard running the identical capture
logic passes. Revert the two matrix "Run E2E tests" steps' `script:`
blocks to main's plain script (just the backgrounded logcat stream +
`./gradlew connectedDebugAndroidTest`) and drop the now-dead "Upload wedge
diagnostics" step from the matrix job.
Kept untouched: the job-level `timeout-minutes: 50` backstop added in
27ede55, and the entire `e2e-preview` job (its own capture_wedge/watchdog
and wedge-diagnostics-api37-preview-shard* upload are unaffected).
The matrix E2E job (api-level 29-36) had no timeout-minutes, so a wedge
hangs until GitHub's 6-hour default instead of being force-killed. The
sibling e2e-preview job already sets timeout-minutes: 35. A normal
matrix run is ~15-20 min and a retry-inclusive run ~40 min, so set
timeout-minutes: 50 to give headroom above the in-step wedge-capture
timeout (1200s) while still bounding worst-case runtime.
E2E legs intermittently WEDGE (hang) with no fast-fail until the job force-kill,
and GitHub's post-force-kill step behavior is unreliable, so #388's diagnostics
don't reliably capture the wedge — and don't capture wedge-specific state anyway.
Wrap the `connectedDebugAndroidTest` run (both the `e2e` matrix first-attempt +
retry, and each `e2e-preview` shard) in an explicit `timeout -k 30s 1200`
(20 min) — comfortably above a normal run (~13-15 min), well below the hard cap —
so a wedge trips the wrapper (exit 124), NOT the force-kill, GUARANTEEING the
capture runs while the emulator is still alive. On 124, capture_wedge grabs the
smoking gun into a `wedge-diagnostics-api<level>` artifact: the running/last test
(logcat TestRunner), SIGQUIT (kill -3) thread dumps of the app + instrumentation
processes (ART -> logcat + /data/anr), dumpsys activity/window, `service list` +
`service check input/window/activity` (the boot-race crux), sys.boot_completed +
init.svc.* state, the snapshot cache-hit note, and accel/kvm/mem/disk. Then it
exits with the real status so #388's diagnostics + the existing retry still fire;
a normal run finishes before the wrapper and is unaffected.
EVIDENCE ONLY — no boot-readiness guard/fix (maintainer: prove the cause first).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add a cheap `changes` job (dorny/paths-filter v4, pinned SHA) that sets
e2e_needed=false only when EVERY changed file is in a safe allow-list
(app/src/test/**, **/*.md, docs/**, scripts/**, .claude/**); anything
else -- or any non-pull_request event -- defaults to true (conservative,
"err toward running E2E").
Gate `e2e` and `e2e-preview` on needs.changes.outputs.e2e_needed so the
whole matrix runs or skips together, and rewrite the `ci-passed` gate:
it now BLOCKS on changes!=success, any of traffic-control-tests /
static-analysis / debug-build / unit-tests !=success, or e2e/e2e-preview
==failure|cancelled -- while TOLERATING an intentional e2e/e2e-preview
'skipped'. So test-only/docs PRs go green on the fast gate, a real E2E
failure/cancel still blocks, and a broken filter (changes!=success)
still blocks. Branch protection ("CI passed") context is unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The dominant merge-blocking flake was the "Set up Android SDK" step
(android-actions/setup-android v4.0.1) dying BEFORE the emulator starts:
Wrong version in preinstalled sdkmanager
Warning: ... preparing SDK package Android Emulator: Error reading Zip
content from a SeekableByteChannel.
Error: The process '.../sdkmanager' failed with exit code 1
Root cause: the action's default cmdline-tools version (20.0) rarely matches the
runner image's preinstalled one, so it logs "Wrong version in preinstalled
sdkmanager" and re-fetches cmdline-tools with NO checksum; it then runs its
default `sdkmanager tools platform-tools` install. Any of those downloads can be
a corrupt/truncated zip, which sdkmanager turns into an un-retried exit 1. v4.0.1
is the latest release, so this is fixed by configuration + hardening, not a bump.
Harden with verify -> reject -> retry, never trusting sdkmanager's exit code
alone, via a new stdlib-only helper .github/scripts/setup_android_sdk.py:
- bootstrap: download the pinned cmdline-tools zip, verify size + SHA-256
(authoritative pin, cross-checked against Google's published SHA-1), and
install it to $ANDROID_SDK_ROOT/cmdline-tools/20.0 -- the exact path
setup-android probes first, so the action reuses the verified tree and never
does its own unverified "Wrong version" re-download. A mismatch (corrupt OR
wrong version) deletes the bad zip + any half-extracted dir and re-downloads.
- install: sdkmanager --install with retry + backoff; on a corrupt package zip it
purges the partial/corrupt package dir (and sdkmanager's temp dirs) before
retrying, forcing a fresh download instead of a re-read.
- setup-android now runs with packages: "" (no flaky tools/platform-tools
install) and cmdline-tools-version: "14742923" (reuse the verified bootstrap).
- actions/cache restore + success-gated save so only a verified SDK is ever
cached (integrity gates the cache); shrinks the re-download/corruption surface.
Applied to every SDK-setup job (debug-build, unit-tests, static-analysis, e2e
matrix, e2e-preview). Emulator BOOT logic, #372 API-37 sharding, and #388
diagnostics are untouched. Pure-logic helpers are unit-tested
(test_setup_android_sdk.py, run by the traffic-control-tests job).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The API 29-36 `e2e` matrix uploaded only its test report, so an emulator
flake or a red leg (e.g. `E2E (31)` dying on a bare `sdkmanager` exit 1)
left nothing to diagnose. Bring the #334 API-37 diagnostics to the matrix,
inline (no changes to `e2e-preview`, which PR #372 is restructuring):
- Stream `adb logcat -v time` to `$RUNNER_TEMP/logcat-api<level>.txt` at the
top of both the "Run E2E tests" and retry reactivecircus steps (emulator is
booted there); backgrounded so gradle stays the exit-status-bearing command.
- New `if: failure()` step dumps device + runner state (adb devices, logcat
tail, emulator -accel-check, /dev/kvm, free -h, df -h) to the step log and a
diagnostics file; every probe guarded with `|| true`.
- New `if: always()` upload-artifact (same pinned v7 SHA) `e2e-diagnostics-api<level>`
carries the logcat + diagnostics files, `if-no-files-found: warn`.
- Make "Install SDK platform and build-tools" diagnosable: bounded 3x retry with
backoff for a transient sdkmanager failure, and print `--list_installed` on a
hard failure instead of a bare exit 1.
Keeps reactivecircus/android-emulator-runner and the existing boot-race retry.
Additive/diagnostic only; no boot-affecting flags change.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The spike commits already implemented the N=2 shard matrix (numShards/
shardIndex via -Pandroid.testInstrumentationRunnerArguments.*), per-shard
test-retry parity, adb start-server before the boot loop, and shard-suffixed
artifact names. This drops the SPIKE / DRAFT "do not merge as-is" framing from
the ci.yml comments and reframes docs/perf/api37-e2e-sharding-spike.md from a
feasibility spike into the adopted design, so the change is mergeable as-is.
Also fixes the doc's section 3a example, which showed 1-based shardIndex values
[1, 2]; shardIndex is 0-based (0..numShards-1) and the implementation correctly
uses matrix.shard: [0, 1] -- [1, 2] would run an empty bucket and silently drop
half the suite.
Fan-in unchanged and verified: ci-passed still lists e2e-preview once; GHA
matrix aggregation makes its result `failure` if either shard fails, so both
shards must pass for the gate to go green. Branch protection requires the
"CI passed" context (not the per-leg "E2E (API 37 preview) (N)" check names),
so no branch-protection change is needed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Folds the #370 root-cause finding into the spike. The stable `e2e` matrix
retries its test run once; `e2e-preview` runs connectedDebugAndroidTest exactly
once, so a flaky test self-heals on API 29-36 but wedges the required gate on
API 37 (e.g. #370's SignaturesScreenTest teardown race).
Doc: adds risk item 9 (retry-parity gap + its sharding interaction — per-test
flake is NOT amplified by sharding unlike boot flake, and a per-shard retry
costs only B + T/N; framed mitigation-not-fix) and two §6 recommendations
(retry parity, mirrored into api37_e2e.py; adb start-server before the boot
loop).
PoC (ci.yml): per-shard single test retry (::warning:: on retried-but-passed)
+ adb start-server before the boot loop. The api37_e2e.py retry mirror stays a
documented recommendation (local path needs a real-emulator validation this
spike did not boot). Still DRAFT, not auto-merged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Feasibility spike for sharding the e2e-preview job (the hand-provisioned
API 37 / google_apis_ps16k 16 KB-page emulator), CI's longest leg
(~16.4-17.6 min). docs/perf/api37-e2e-sharding-spike.md breaks the leg into
fixed overhead B ~8.3 min (setup + boot + Gradle daemon/config/compile/install)
vs parallelizable test execution T ~8.8 min, models B + T/N for N=2/3/4, and
recommends N=2 (~17.1 -> ~12.7 min, ~28% off the critical path) capped by the
API 30 matrix wall (~12.0 min) beyond N=3.
DRAFT PoC (do NOT merge as-is): converts e2e-preview to a strategy.matrix.shard
[0, 1] fan-out passing AndroidJUnitRunner numShards/shardIndex through the
existing -Pandroid.testInstrumentationRunnerArguments.* channel (no GMD, no
orchestrator, no Gradle change). Artifact names gain a shard suffix;
ci-passed still lists e2e-preview once (matrix fan-in keeps the single gate).
Local preflight stays single-emulator. Relates to #258.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Move the traffic-control (runner-priority orchestration) job verbatim out of
.github/workflows/ci.yml into a new standalone workflow,
.github/workflows/traffic-control.yml, so the heavy CI jobs no longer depend
on it. The job's YAML (name, runs-on, timeout-minutes, permissions, env,
steps) and its documentation comment move unchanged; the decision core
.github/scripts/traffic_control.py is untouched and still unit-tested by the
traffic-control-tests job in ci.yml.
In ci.yml: removed the traffic-control job, dropped needs: traffic-control
from the five heavy jobs (static-analysis, debug-build, unit-tests, e2e,
e2e-preview) and from traffic-control-tests (its only needs, which would
otherwise dangle at a now-deleted job), and updated the now-stale header and
ci-passed comments to point at the extracted workflow.
The new workflow will be disabled pending a rebuild as a published GitHub
Action.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
End the merge cascade and give the traffic-controller ownership of CI *triggering*.
- autoupdate.yml updates PR branches with the built-in GITHUB_TOKEN instead of a PAT,
so an update push no longer auto-retriggers CI (GitHub's anti-recursion rule) — the
cascade (every merge re-runs every PR, cancel-in-progress thrashing them) is gone.
- New scheduler ci-trigger.yml -> traffic_control.py --mode trigger (re-)triggers CI
for the highest-priority PR(s) whose head SHA has absent/stale checks, a few at a
time (inflight cap), in the existing P0-P9 / broken-draft priority order — a
poor-man's merge queue reusing the priority core. It runs after autoupdate finishes
(workflow_run, race-free) plus a cron backstop plus manual dispatch.
- Triggering uses workflow_dispatch, which is EXEMPT from anti-recursion, so the
built-in GITHUB_TOKEN (actions: write) starts the run — NO PAT / secret change needed.
- ci.yml gains a workflow_dispatch trigger (pr/head_sha/reason inputs) and a per-PR
concurrency group unifying pull_request and dispatch runs; its on: pull_request path
is kept so brand-new PRs, human pushes, and fork PRs always get CI (fail-open).
Pure select_triggers / classify_sha_runs decision core added to traffic_control.py with
24 new unit tests (priority order, oldest-first fairness, inflight cap, fork skip, P0
bypass+preempt, head-SHA needy classification, and a liveness/anti-starvation simulation).
Closes#349
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adds a fast traffic-control-tests job (ubuntu, actions/checkout +
actions/setup-python, no emulator/Gradle) that runs the 37 pure-stdlib
unit tests for .github/scripts/traffic_control.py on every PR, and
wires it into ci-passed's needs so a regression blocks merge instead
of only being caught locally.
Closes#346
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Restore (and extend to drafts) the old bash's broken-reclaim behaviour that the
initial Python refactor had dropped. runs_to_cancel now cancels an OTHER PR's
active/queued runs when EITHER:
(a) THIS PR is P0 and that PR is strictly-lower (reclaim every lower runner); OR
(b) that PR is broken/draft (effective priority 10) and THIS PR is strictly-higher
(effective priority < 10) — a wasted run any ready PR may reclaim.
P1-P9 still never bump a *normal* (non-broken/draft) lower run; a broken/draft PR
(P10) preempts nothing (nothing is strictly-lower than the bottom, and the
equal-or-higher invariant means a P10 never cancels another P10). Self / main-push /
equal-or-higher invariants unchanged.
Updates the module docstring + ci.yml comments (the "only P0 preempts" wording
becomes: P0 preempts everything strictly-lower; additionally, any strictly-higher PR
preempts a broken/draft run) and the job step/permission/needs comments. Adds unit
tests: P3 reclaims a broken P10 run and a draft P10 run; P3 does not bump a normal P5
run; a P10 self preempts nothing; plus an end-to-end P5-reclaims-draft-then-waits
scenario. 37 unit tests pass; ci.yml parses clean; --dry-run shows a P3 cancelling a
draft (and broken) run while still yielding to a higher P1.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The priority-based runner orchestration ("traffic-control") lived as a large
inline-bash step in ci.yml — a two-pass preemption + hold-back script that was
effectively untestable in YAML. Move it into .github/scripts/traffic_control.py,
structured as a pure decision CORE + a thin gh-I/O SHELL:
* Pure functions (no network/clock/subprocess), unit-testable in isolation:
- effective_priority(pr): lowest-numbered P0-P9, default P5; broken OR draft => 10.
- runs_to_cancel(this_pr, all_prs, self_run_id): PASS 1 — run ids to cancel,
empty unless THIS PR is P0; only strictly-lower running/queued runs; never self
(by number or run id), never equal-or-higher.
- wait_blockers(this_pr, all_prs): PASS 2 — yield to any strictly-higher PR with
an active/queued run, and to same-level peers ordered ahead (running-first,
then oldest createdAt). Empty => proceed.
* Shell (run_live): gathers the snapshot via gh, applies cancels, runs the bounded
hold-back poll loop; always exits 0. --dry-run feeds the core a snapshot JSON and
prints decisions with zero network.
* No jq/bash dependency (cross-platform, per the repo's Python-stdlib convention).
ci.yml's traffic-control job now checks out the repo and runs the module. Job
permissions gain `contents: read` (for checkout) alongside the existing
`actions: write` / `pull-requests: read`; env and downstream `needs:` wiring
unchanged; step stays `continue-on-error`.
33 stdlib unittest cases cover priority resolution, P0-only preemption, the
self/main/equal-or-higher invariants, and the same-level running-first/oldest
ordering.
Behaviour is preserved except the ticket's refinements: (1) drafts now count as
P10 (bottom); (2) an explicit same-level running-first-then-oldest tiebreaker; and
(3) per the ticket's order-of-operations, ONLY P0 preempts — the old bash also let
any higher-priority PR cancel a `broken` target's run, which no longer happens
(a broken/draft run is only cancelled by a P0, via the same strictly-lower rule).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The API 37 preview E2E boot has flaked twice (#285, #333) with only
"did not boot within 300s" and no root-cause signal. Add rich boot
diagnostics by default, kept in parity between CI and the local
hand-provisioning script (api37_e2e.py):
- Launch the emulator with `-verbose -debug init,avd_config,kernel`
(diagnostics only; no boot-affecting flag changed), still redirecting
to $EMU_LOG.
- Stream `adb logcat -v time` to a file from the moment the device
registers (via `adb wait-for-device logcat`, backgrounded).
- On a boot timeout, dump accel-check, /dev/kvm presence, GPU mode,
free mem/disk, the AVD config.ini and the emulator.log tail; CI writes
these to a boot-diagnostics file, the local script prints them.
- CI uploads emulator.log + logcat.txt + boot-diagnostics.txt as an
artifact with `if: always()` so they survive a timeout/cancel, and
prints a concise summary (accel/KVM status + last 50 lines of
emulator.log) to the step log.
The existing 2-attempt boot retry + boot-completed wait loop are
unchanged; the diagnostics are additive.
Closes#334
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Scope :app:jacocoTestReport's denominator to the JVM-testable surface and
add a :app:jacocoTestCoverageVerification no-regression gate that shares the
same classDirectories/executionData/sourceDirectories, wired into both the
`check` lifecycle task and CI's unit-test job (part of the `CI passed` gate).
Excluded from the denominator (structurally unreachable from a JVM unit
test): Compose screen/component render code, Android framework entry points
(*Activity/*Service/Application/*BackupAgent), Hilt DI (**/di/**), and the
src/debug cold-open probe. Kept in scope: ViewModels, repositories, mappers,
DAOs, utils, richtext, mail, reporting logic, and the six WorkManager Workers.
Corrects PR #292, which excluded **/*Worker*: SyncWorker, BackfillWorker,
PruneWorker, SendWorker, ReportPurgeWorker and ReportUploadWorker are all
directly unit-tested, so they stay counted in both numerator and denominator
(only their Hilt wiring, WorkManagerModule, is excluded, via **/di/**).
Baseline: 80.21% line (4838/6032). Floor: 0.79 (~1.2% headroom) so ordinary
noise doesn't red-flag it while a real drop fails. Manual ratchet for now:
bump the floor up in the same PR when coverage rises materially.
Closes#251Closes#292
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Only P0 preempts in-progress runs (emergency reservation). P1-P9 no longer
cancel lower-priority runs; instead traffic-control holds back (bounded poll,
kept under timeout-minutes) while strictly-higher-priority PRs still have
active/queued CI runs, so their heavy jobs reach the runner queue first.
New `broken` label forces effective priority below P9 (sentinel 10): a broken
PR never preempts (even if also labelled P0 -- broken wins) and always yields,
and because its run is wasted, ANY higher-priority PR (not just P0) may cancel
its in-progress run to reclaim the runner. Net rule: a strictly-lower run is
cancelled iff (self is P0) OR (target is broken); otherwise yield.
All existing safety preserved: never main/push runs, never our own run, never
an equal-or-higher-priority PR; PR-controlled strings via env/jq only;
continue-on-error + set +e + always exit 0; traffic-control stays a
non-required best-effort job and ci-passed is unchanged.
Validated with actionlint and a mocked-gh + fake-clock logic harness.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add a lightweight `traffic-control` job that runs first (the heavy
build/E2E jobs `needs:` it) and preempts contended runners by PR
priority. It reads the triggering PR's P0–P9 label (P0 = highest,
P9 = lowest; default P5 when unlabeled) and cancels the in-progress /
queued CI runs of strictly-lower-priority OTHER open PRs, freeing their
runners for the higher-priority PR.
Safety: never cancels main/push runs, the PR's own run, or an
equal-or-higher-priority PR — only strictly-lower-priority OTHER open
PRs' active CI runs. The job is best-effort (every gh call guarded,
always exits 0, step is continue-on-error) and is NOT part of the
`CI passed` merge gate. `ci-passed` now also treats a `skipped` heavy
job as a gate failure, so a (should-never-happen) traffic-control
failure blocks the merge fail-safe rather than passing it untested.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Completes the crash-interrupted #192 WIP (app/build.gradle.kts already had a
jacocoTestReport task and toolVersion pin recovered onto build-192-jacoco):
- Move the JaCoCo tool version into gradle/libs.versions.toml instead of a
hardcoded string in app/build.gradle.kts, matching how every other plugin
version in this repo is sourced.
- Fix the generated-code exclusion list against the real compileDebugKotlin
output (verified by inspecting the compiled class tree): Room's
KSP-generated `_Impl` DAOs/database and the Compose compiler's per-file
ComposableSingletons holders are the only generated code that actually
lands in classDirectories, since Hilt/Dagger's generated Java and AGP's
BuildConfig/R/Manifest are compiled by a separate javac task this report
never reads. Drop the blanket `**/*$$*` exclude the WIP had — it was
silently discarding ~200 real classes' worth of coverage on Kotlin's own
`$$inlined$` synthetic classes (e.g. Flow.map { ... } transforms in the
repositories), which is hand-written logic, not generated boilerplate.
- Add Hilt_*/Dagger* prefix patterns so the (currently inert,
belt-and-suspenders) Hilt exclusions are actually correct if the
classDirectories scope ever changes.
- Add a minimal CI step to the existing unit-tests job that runs
jacocoTestReport and uploads the XML+HTML report as a build artifact.
No coverage threshold gate yet (a jacocoTestCoverageVerification rule is
a natural follow-up once there's a baseline).
- Document the new :app:jacocoTestReport task in CLAUDE.md.
Verified on JDK 21: fast gate (assembleDebug, testDebugUnitTest,
compileDebugAndroidTestKotlin, lintDebug, ktlintCheck, detekt) plus
jacocoTestReport all pass, from both a warm and a `clean` build. The
report shows real signal (30% instruction / 38% line coverage) with no
generated classes leaking in.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The matrix E2E (29) job intermittently fails (~2%, API-29 only) in
reactivecircus/android-emulator-runner's un-guarded, fatal post-boot
`adb shell input keyevent 82`: on snapshot resume sys.boot_completed=1 is
restored before system_server republishes the `input` binder service, so
the job aborts before Gradle runs with "No service published for: input"
(fast-fail ~1m43s). Proven on run 28667366203.
Make the "Run E2E tests" step non-fatal (id + continue-on-error) and add a
guarded second attempt (if steps.e2e.outcome == 'failure'). Two independent
boots drop the race to ~0.04%; a genuine failure on both attempts still
fails the job (outcome, not conclusion). Definitive manual-boot fix: #218.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add a `static-analysis` job (JDK 21 + Android SDK) that runs
`:app:ktlintCheck :app:detekt` and uploads the reports, and add it to the
`ci-passed` aggregating gate so lint regressions block merges like the
other checks.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The custom-provisioned API 37 preview emulator has been stable, so fold its E2E
job into the aggregating "CI passed" gate's needs. Because branch protection
requires only that single check, no settings change is needed.
Drop the now-inaccurate "non-blocking" wording from the job name and comments.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The preview emulator failed every boot with "Unknown AVD name [api37]" and
exited immediately (no device -> the bounded wait timed out). avdmanager had
created the AVD under $ANDROID_SDK_HOME/.android/avd while the emulator searched
$ANDROID_SDK_HOME/avd and $HOME/.android/avd. Pin ANDROID_AVD_HOME to one path
both tools use, carry it to the boot step via $GITHUB_ENV, and list AVDs after
creation to verify.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The custom-provisioned preview job hung in 'Boot emulator and run E2E' until
the 35-min cap: adb wait-for-device had no timeout and the emulator's own
output was never captured, so a failed boot was both invisible and unbounded.
Bound the wait with a single 300s timeout covering device registration + full
boot, retry once, capture the emulator log, and dump it (plus logcat) on
failure.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Linux-arm64 runners can't set up the SDK here: android-actions/setup-android's
sdkmanager fails (exit 1) on the android-37.0 preview platform, and the emulator
package has no arm64-Linux build. These are pure build/JVM-unit jobs whose
results are host-arch-independent, so x86_64 loses no device coverage (real
arm64 ABI coverage would require arm64 emulators, i.e. macOS hosts).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Folder view: a left navigation drawer lists each account's IMAP folders;
tapping one browses and caches that folder's mail. IMAP UIDs are unique only
within a folder, so message identity, the fetch/read/flag/delete paths, sync,
and the Room cache all became folder-aware (id = "accountId:folder:uid"; new
`folder` column; schema v7->v8). Standard folders (Inbox/Sent/Drafts/Spam/Trash/
Archive) surface with friendly names + icons via RFC 6154 SPECIAL-USE attributes
with a case-insensitive name fallback; the multi-account drawer adds an account
switcher and a unified "All Inboxes". INBOX stays the only auto-synced,
IDLE-watched, notifying folder; other folders sync on demand.
Lower minSdk 33 -> 29 for a rolling ~7-year Android support window; guard the
API-33 POST_NOTIFICATIONS runtime request accordingly.
Tests and CI:
- Bump espresso-core 3.6.1 -> 3.7.0 so Compose UI tests run on API 37
(3.6.1's InputManagerEventInjectionStrategy reflects a removed hidden method).
- New coverage across layers: FolderRoleTest, ImapClientTest folder cases,
MailboxViewModelTest, MailRepositoryImplTest folder routing, a FolderDrawer
Compose UI test, and LibreMailDatabaseTest folder DAO/reconcile tests.
- Gradle Managed Devices + a CI E2E matrix over every API 29-36; a single
"CI passed" gate job fans in all jobs and is required by branch protection.
- Non-blocking, custom-provisioned API 37 (preview) E2E job with image caching.
- Build + unit-test jobs run on arm64 (ubuntu-24.04-arm); emulators stay x86_64.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
GitHub Actions (every action pinned to its commit SHA, version in a comment):
- ci.yml, on pull_request to main, runs three jobs: a debug build, the
unit tests, and the instrumented suite on a headless emulator.
- release.yml, on workflow_dispatch, builds the release APK, archives the
source as zip + tar.gz, and publishes a GitHub release.
Compose UI tests (app/src/androidTest), driving the real screens with
fake-backed view models so they need no network, database, or Hilt graph:
- ManualSetupScreen: submit-button validation, advanced-options toggle,
and add-account success/failure.
- ComposeScreen: send-button enablement and the send -> close flow.
- LibreMailBottomBar: tab rendering and selection callback.
Mark gradlew executable (100755) so it runs on the Linux CI runners.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>