Feasibility spike for sharding the e2e-preview job (the hand-provisioned API 37 / google_apis_ps16k 16 KB-page emulator), CI's longest leg (~16.4-17.6 min). docs/perf/api37-e2e-sharding-spike.md breaks the leg into fixed overhead B ~8.3 min (setup + boot + Gradle daemon/config/compile/install) vs parallelizable test execution T ~8.8 min, models B + T/N for N=2/3/4, and recommends N=2 (~17.1 -> ~12.7 min, ~28% off the critical path) capped by the API 30 matrix wall (~12.0 min) beyond N=3. DRAFT PoC (do NOT merge as-is): converts e2e-preview to a strategy.matrix.shard [0, 1] fan-out passing AndroidJUnitRunner numShards/shardIndex through the existing -Pandroid.testInstrumentationRunnerArguments.* channel (no GMD, no orchestrator, no Gradle change). Artifact names gain a shard suffix; ci-passed still lists e2e-preview once (matrix fan-in keeps the single gate). Local preflight stays single-emulator. Relates to #258. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
16 KiB
API 37 preview E2E sharding — feasibility spike
A time-boxed investigation into whether sharding the e2e-preview job (the hand-provisioned
API 37 / google_apis_ps16k 16 KB-page emulator that runs the whole instrumented suite serially)
would meaningfully cut CI wall-clock. e2e-preview is consistently the single longest leg in CI
(~16.4–17.6 min on recent runs), so it sets the pipeline's critical path.
This is an assessment + recommendation, not a finished feature. It was produced from CI run logs
and repo config only — no emulator was booted locally (avoids the machine-freeze risk noted in the
project memory and runner contention). All numbers below come from real successful e2e-preview runs.
TL;DR / recommendation
GO — shard e2e-preview in CI at N = 2 (a strategy.matrix.shard: [1, 2] fan-out, each shard
provisioning its own API 37 emulator and running one half of the suite via AndroidJUnitRunner's built-in
numShards/shardIndex). Projected: the leg drops from ~17.1 min → ~12.7 min, cutting the whole
pipeline's critical path by ~4.9 min (~28%) for the cost of +1 emulator job (~1.5× this leg's
emulator-minutes).
N = 3is a defensible stretch (~11.2 min) but buys only ~0.7 min more of total-CI time, because at that point the rest of the E2E matrix (API 30, ~12.0 min) becomes the new critical path. The extra 0.7 min costs a second added emulator and +50% more preview-boot flake surface.- Do NOT exceed
N = 3. Below ~12 min the API 29–36 matrix caps the pipeline; further API 37 shards burn emulator-minutes for zero wall-clock gain. - Keep local preflight single-emulator.
api37_e2e.py/local_instrumented.pymust stay one cold boot — a dev box has one free hypervisor (VT-x/WHPX) and the project memory explicitly warns parallel emulators can freeze the machine. Sharding is a CI-only change. - The mechanism is already proven in-repo:
local_instrumented.pypasses-Pandroid.testInstrumentationRunnerArguments.class=…to the sameconnectedDebugAndroidTesttask, sonumShards/shardIndexargs flow the identical way — no GMD, no orchestrator, no Gradle change.
1. Current job anatomy + timing breakdown
e2e-preview (.github/workflows/ci.yml) hand-provisions the emulator because API 37's only image is the
nonstandard system-images;android-37.0;google_apis_ps16k;x86_64 (16 KB page size), which neither
reactivecircus/android-emulator-runner nor AGP's Gradle Managed Device DSL can build. It installs the
SDK + image directly (sdkmanager), creates the AVD (avdmanager), cold-boots headless with a
two-attempt retry loop, then runs ./gradlew connectedDebugAndroidTest. api37_e2e.py mirrors this
locally (identical image/flags except -gpu auto-no-window vs CI's -gpu swiftshader_indirect).
Step timing — reference run 28811804353 (branch fix-359-sqlcipher-16kb, success, total 17.6 min)
| Phase | Step(s) | Time | Re-paid per shard? |
|---|---|---|---|
| Job/checkout/JDK/Gradle/KVM setup | 1–6 | ~0.4 min | yes |
Restore image cache + sdkmanager install (cache hit) |
7–8 | ~0.7 min | yes |
avdmanager create avd |
9 | ~0.03 min | yes |
| Emulator cold boot (1 failed attempt + retry) | 10a | ~1.1 min | yes |
| Gradle daemon start + configure + dependency resolution | 10b | ~4.1 min | yes |
| Compile debug app + androidTest APK + install | 10c | ~2.0 min | yes (per runner) |
Test execution (Starting 263 tests → Finished 263 tests) |
10d | ~8.75 min | NO — parallelizable |
| Teardown, upload report/diagnostics, post-cache/gradle | 12–29 | ~0.4 min | yes |
Step 10 ("Boot emulator and run E2E", 16.05 min) bundles boot + build + test into one shell step; the
sub-splits above were recovered from the step's own timestamped log markers
(##[group]Start API 37 emulator, Starting N tests on emulator-5554, Finished N tests, BUILD SUCCESSFUL).
The boot-vs-test ratio (the crux)
Sharding only parallelizes test execution; every shard re-pays boot and the whole fixed prologue (setup + boot + Gradle daemon/config + compile + install + teardown). The fixed cost is not just boot (~1.1 min) — it is dominated by the ~6 min Gradle daemon-start/configure/compile block. Splitting the job:
T(parallelizable test execution) ≈ 8.8 minB(fixed, re-paid by every shard) ≈ 8.3 min
i.e. B ≈ T — roughly a 50/50 split. That is the single most important number in this spike: because the fixed overhead is about as large as the test time, sharding has a hard floor of ~8.3 min no matter how many shards you add.
Cross-run confirmation (test execution is stable, not a one-off)
| Run | Branch | Total leg | Test-exec window | Boot attempts |
|---|---|---|---|---|
28811804353 |
fix-359-sqlcipher-16kb | 17.6 min | ~8.75 min (263 tests) | 2 |
28763651551 |
feat-125-imap-connection-reuse | 16.9 min | ~9.2 min (262 tests) | 1 |
28760208017 |
fix-359-sqlcipher-16kb | 16.4 min | ~8.4 min (261 tests) | 2 |
Test execution is a tight 8.4–9.2 min; the leg total 16.4–17.6 min. Run-to-run variance (~±0.6 min) is mostly the boot retry (1 vs 2 attempts costs ~1 min) — reinforcing that boot flake, not test time, drives the jitter.
2. Instrumented-test inventory
- ~263 instrumented
@Testmethods across 59 test classes (app/src/androidTest/…; the runner reports "Starting 263 tests"). Confirmed by static count (@Testoccurrences) and the CI log. - Distribution is class-skewed: a few heavy classes dominate —
MessageDaoTest(17),MigrationTest(15),MailboxScreenTest(13),AccountDataMigratorTest(8), several DAO/*Databaseclasses (7 each) — while most classes have 1–6. - The suite is a mix of Room/DAO/migration DB tests (CPU + disk on
google_apis_ps16k) and Compose UI tests (Espresso/createComposeRule). No single test is a multi-minute outlier; 263 tests / ~8.8 min ≈ ~2 s/test average, so runtime is spread reasonably evenly and count-based sharding will be approximately runtime-balanced at N = 2–3.
At ~263 tests, N = 2 (~131/shard) or N = 3 (~88/shard) are the sensible shard counts. Beyond that the per-shard test time shrinks below the fixed boot/build cost and the split stops paying off (see §4).
3. Sharding approaches
3a. AndroidJUnitRunner built-in sharding across a GHA matrix — RECOMMENDED
AndroidJUnitRunner has first-class sharding: it hashes each test name into one of numShards buckets and
runs only shardIndex's bucket. AGP surfaces runner args through Gradle properties, so on the existing
hand-provisioned connectedDebugAndroidTest (no GMD required) each shard runs:
./gradlew connectedDebugAndroidTest \
-Pandroid.testInstrumentationRunnerArguments.numShards=2 \
-Pandroid.testInstrumentationRunnerArguments.shardIndex=${{ matrix.shard-index }}
Wrapped in a strategy.matrix.shard: [1, 2] (or [1, 2, 3]) fan-out where each matrix leg provisions its
own API 37 emulator and runs one shard. This is the same fan-in pattern the repo already uses for
the e2e (API 29–36) matrix, and the same runner-arg channel local_instrumented.py already uses for
…arguments.class= — so it is a known-good mechanism here, needs no orchestrator, no Gradle-side change,
and no branch-protection change (see §5).
Confirmed compatible with the hand-provisioned 16 KB job: runner-arg sharding is a property of
AndroidJUnitRunner + the AGP connected*AndroidTest task, entirely independent of how the AVD was
created. GMD is irrelevant. Count-based, deterministic, and stable across runs.
3b. Alternatives (rejected)
- Manual test-class partitioning (
-Pandroid.testInstrumentationRunnerArguments.class=A,B,Cper shard, orpackage/annotation buckets): lets you hand-balance by runtime, but is brittle — every new test class must be manually assigned or it silently runs in no shard (coverage gap) or a wrong one. Not worth it for a ~0.5 min balance gain over hashing. - Android Test Orchestrator (
clearPackageData): solves per-test process isolation, not wall-clock; it actually slows a run (fresh process per test). Orthogonal to sharding — skip. maxParallelForks(in-JVM): N/A — that is a JVM unit-test knob (and the very thing #258 rejected because GreenMail binds fixed ports). Instrumented tests run on the device, not forked JVMs.
4. Expected wall-clock gain vs. CI cost
Model: an N-shard leg ≈ B + T/N (parallel), costing ≈ N × (B + T/N) emulator-minutes. Using the measured
B ≈ 8.3 min, T ≈ 8.8 min:
| N | Leg time B + T/N |
Δ vs today | Emulator-minutes N × leg |
Cost multiple |
|---|---|---|---|---|
| 1 (today) | 17.1 min | — | 17.1 | 1.0× |
| 2 | 12.7 min | −4.4 min | 25.4 | 1.5× |
| 3 | 11.2 min | −5.9 min | 33.7 | 2.0× |
| 4 | 10.5 min | −6.6 min | 42.0 | 2.5× |
| ∞ | 8.3 min (floor = B) | −8.8 min | — | — |
But the total-CI critical path is max(all jobs), and the API 29–36 E2E matrix already runs ~8.3–12.0
min (API 30 ≈ 12.0 is its ceiling). So the API 37 leg only helps the pipeline until it drops to that wall:
| N | API 37 leg | Total-CI critical path = max(API 37, API 30≈12.0) | Total CI saved |
|---|---|---|---|
| 1 | 17.1 | 17.6 (API 37) | — |
| 2 | 12.7 | 12.7 (API 37, just above the matrix wall) | ~4.9 min (~28%) |
| 3 | 11.2 | 12.0 (now API 30 caps it) | ~5.6 min (~32%) |
| 4 | 10.5 | 12.0 (API 30) — no gain over N=3 | ~5.6 min |
Point of diminishing returns. Two effects compound:
- B ≈ T gives a hard ~8.3 min floor for the leg itself (N = ∞).
- Long before that, the ~12.0 min API 30 matrix wall caps total CI. N = 2 lands the leg at 12.7 (~0.7 above the wall); N = 3 tucks it under the wall (API 30 takes over). N = 3 is the last count that changes anything; N ≥ 4 is pure waste unless the API 29–36 matrix is also sharded (a separate, larger effort).
Verdict: N = 2 captures ~4.9 of the ~5.6 min theoretically achievable — the overwhelming majority — for the least cost and flake. N = 3 collects the last ~0.7 min at +1 more preview emulator.
5. Risks / gotchas
- Flake multiplication (top code-risk).
e2e-previewis already the flakiest leg (16 KB preview image; boot/snapshot races; theinput keyevent 82timing hazard; infra-vs-code ambiguity). N shards = N independent cold boots per pipeline, soP(all shards boot)=p^N. Mitigation is built in: each shard inherits the job's existing two-attempt boot loop (the memory note pegs a single un-retried boot race at ~2% → ~0.04% after two attempts), keeping aggregate boot success ~99.9% even at N = 3. Keep the retry loop per shard; do not collapse it to save time. - Runner concurrency (top practical-risk). The pipeline already fans out 8 E2E matrix jobs +
e2e-preview+ 3 lighter jobs (~12 concurrent). N = 2/3 pushes peak to ~13/14 concurrent runners. If the account's concurrent-job limit is hit, shards queue serially and the entire wall-clock gain evaporates (or inverts). The project already actively manages runner allocation/priority (P0–P9), so this is a real, tracked constraint — confirm headroom before enabling, and give the preview shards the same priority as the rest of the E2E matrix. ci-passedgate must require ALL shards — but needs no branch-protection edit. Convertinge2e-previewto a matrix makes GHA roll all shard legs under the onee2e-previewentry already inci-passed.needs; the gate'scontains(needs.*.result, 'failure' | 'cancelled' | 'skipped')check fails if any shard fails (a matrix job'sresultisfailurewhen any leg fails). Branch protection only requires the singleCI passedcontext, so no protected-check list change — same as the existinge2ematrix. Verified against the current gate wiring.- Artifact-name collision. Each shard uploads its own
app/build/reports/androidTests/connected/HTML report;actions/upload-artifact@v7errors on duplicate names. The report name must gain the shard index (e.g.e2e-test-report-api37-preview-shard${{ matrix.shard }}), mirroring the API-level suffix thee2ematrix already uses. Same for the boot-diagnostics artifact. - No cross-shard coverage merge needed (a simplifier vs #258). #258's blocker was JaCoCo aggregation
across shards. Here, instrumented/E2E coverage is explicitly out of scope (
app/build.gradle.ktsJaCoCo scoping + issue #192), so there is nothing to merge — shard results only need to fan into the pass/ fail gate. - Cold-cache thundering herd. The ~1 GB preview image is cached (
actions/cache, keyed identically for all shards). On a warm cache all shards get a fast hit (steady state). On a cold cache (first run after a key change), all N shards download ~1 GB concurrently — a rare N× bandwidth spike, self-healing on the next run. Acceptable; no change needed. - Shard balance. Count-based hashing is only approximately runtime-balanced; a shard that draws
MigrationTest+ the DAO-heavy classes may run marginally longer. Tolerable at N = 2–3 (§2); revisit only if a shard consistently lags. concurrency: cancel-in-progresschurn. Unchanged semantics (all shards share the PR concurrency group, cancelled together on a new push), but there are now more in-flight jobs to cancel and re-trigger on each push — a small addition to the known CI merge-cascade thrash. No new failure mode.
6. Recommendation
Worth doing: yes. e2e-preview is the pipeline's critical path, so unlike #258 (unit tests, which finish
~8.5 min inside the E2E-bounded gate and were correctly closed as not-worthwhile) sharding this leg
directly shortens total CI.
- Ship
N = 2via astrategy.matrix.shardfan-out using AndroidJUnitRunnernumShards/shardIndex(§3a). Expected total-CI critical path ~17.6 → ~12.7 min (~28% off) for +1 emulator job. Keep the per-shard two-attempt boot loop; suffix shard index on the report/diagnostics artifact names; leavee2e-previewinci-passed.needs(matrix fan-in keeps the single gate). N = 3optional only if the last ~0.7 min matters and runner headroom is comfortable — it moves the critical path onto the API 30 matrix leg (~12.0 min). NeverN ≥ 4without also sharding the API 29–36 matrix.- Local preflight stays single-emulator —
api37_e2e.pyandlocal_instrumented.pyare unchanged. One free hypervisor per dev box; parallel local emulators risk freezing the machine (project memory). The divergence is intentional and should be documented in the job comment.
Relationship to #258 (unit-test sharding)
#258 was closed not-worthwhile because the unit-test job (~6.5 min) isn't on the critical path (E2E
dominates), so sharding it saved ~0 wall-clock, and maxParallelForks collided on GreenMail's fixed ports.
Reused pattern: measure the critical path first, only shard what's on it. API 37 preview is on it,
so the conclusion flips to GO. Distinct from #258: different layer (on-device instrumented vs JVM),
different mechanism (numShards/shardIndex matrix vs maxParallelForks), and no coverage-merge
constraint (instrumented coverage is out of scope), which makes this strictly simpler to land.
PoC
This branch carries a draft, non-auto-merged PoC that converts e2e-preview to an N = 2 shard matrix
(clearly marked as a spike). Because a PR runs the workflow from its own branch, the draft PR's own CI
run is the validation — the maintainer can confirm both shards go green and read the real per-shard
wall-clock straight off that run (this is a CI-infra change, so a green pipeline is the "test", per #258's
DoD note). It is not to be merged as-is: adopt after reviewing the live shard timings and confirming
runner headroom.