e2e-preview (the hand-provisioned API 37 / google_apis_ps16k 16 KB-page emulator that ran the whole instrumented suite serially) is consistently the single longest leg in CI (~16.4–17.6 min), so it sets the pipeline's critical path. This shards it across 2 parallel API 37 emulators at N = 2, cutting the leg from ~17.1 → ~12.7 min — about ~4.9 min / ~28% off the total-CI critical path — for the cost of one extra emulator job.
Design + measurements live in docs/perf/api37-e2e-sharding-spike.md (updated from a feasibility spike to the adopted at N=2 design record).
What changed (.github/workflows/ci.yml, e2e-preview job)
2-shard matrix.strategy: { fail-fast: false, matrix: { shard: [0, 1] } }. Each leg hand-provisions its own API 37 emulator (provisioning unchanged) and runs one shard via AndroidJUnitRunner's built-in sharding, passed through AGP's existing runner-arg channel (the same one local_instrumented.py uses for .class=):
shardIndex is 0-based, so the matrix values are [0, 1]. No GMD, no orchestrator, no Gradle-side change.
Artifact names carry the shard index (…-shard${{ matrix.shard }} on both the boot-diagnostics and the E2E-report uploads) — actions/upload-artifact@v7 errors on duplicate artifact names.
Retry parity (mitigation). Each shard retries its connectedDebugAndroidTestonce on failure — matching the single retry the stable e2e (API 29–36) matrix already has, which e2e-preview lacked (so a flaky test self-heals on API 29–36 but would wedge the required gate on API 37, e.g. #370). Per-shard, so the retry re-runs only that shard's half of the suite. Caveat: a blanket retry can mask genuine regressions, so the retried-but-passed case is surfaced as a ::warning:: and the real flaky-test fix stays test-level (tracked separately). Retry count is one, not more.
adb start-server before the boot loop, mirroring api37_e2e.py's wait_for_boot, to avoid the attempt-1 adb-daemon "Address already in use" bind race seen in #370.
Gate fan-in (both shards must pass — no branch-protection change)
ci-passed keeps e2e-preview as a single entry in its needs:. GHA's matrix aggregation makes needs.e2e-preview.result = failure if either shard fails, and the gate fails on any failure/cancelled/skipped — so both shards must pass for CI passed to go green. This is the same fan-in the API 29–36 e2e matrix already relies on. Branch protection requires the CI passed context (verified: Debug build, Unit tests, CI passed), not the per-leg E2E (API 37 preview) (0/1) check names, so no branch-protection change is needed. fail-fast: false keeps a failing shard from cancelling its sibling, so both reports always upload.
Local preflight unchanged
api37_e2e.py / local_instrumented.py stay single-emulator — one free hypervisor per dev box; parallel local emulators risk freezing the machine. Sharding is CI-only by design. (api37_e2e.py already matches CI on the boot-retry loop and adb start-server; mirroring the single test-retry into it is a small documented follow-up, deliberately not bundled here because it can't be validated without booting a local emulator.)
Validation
This is a CI-infra change, so this PR's own CI run is the validation — both E2E (API 37 preview) (0) and (1) shards must go green and satisfy CI passed, and the real per-shard wall-clock is readable straight off this run. YAML syntax, matrix expansion, the numShards/shardIndex args, artifact-name suffixing, fail-fast: false, and the retry block were sanity-checked locally.
## What & why
`e2e-preview` (the hand-provisioned API 37 / `google_apis_ps16k` 16 KB-page emulator that ran the whole instrumented suite serially) is **consistently the single longest leg in CI** (~16.4–17.6 min), so it sets the pipeline's critical path. This shards it across **2 parallel API 37 emulators** at `N = 2`, cutting the leg from **~17.1 → ~12.7 min** — about **~4.9 min / ~28% off the total-CI critical path** — for the cost of one extra emulator job.
Design + measurements live in `docs/perf/api37-e2e-sharding-spike.md` (updated from a feasibility spike to the **adopted at N=2** design record).
## What changed (`.github/workflows/ci.yml`, `e2e-preview` job)
- **2-shard matrix.** `strategy: { fail-fast: false, matrix: { shard: [0, 1] } }`. Each leg hand-provisions its **own** API 37 emulator (provisioning unchanged) and runs one shard via AndroidJUnitRunner's built-in sharding, passed through AGP's existing runner-arg channel (the same one `local_instrumented.py` uses for `.class=`):
```
./gradlew connectedDebugAndroidTest \
-Pandroid.testInstrumentationRunnerArguments.numShards=2 \
-Pandroid.testInstrumentationRunnerArguments.shardIndex=${{ matrix.shard }}
```
`shardIndex` is **0-based**, so the matrix values are `[0, 1]`. No GMD, no orchestrator, no Gradle-side change.
- **Artifact names carry the shard index** (`…-shard${{ matrix.shard }}` on both the boot-diagnostics and the E2E-report uploads) — `actions/upload-artifact@v7` errors on duplicate artifact names.
- **Retry parity (mitigation).** Each shard retries its `connectedDebugAndroidTest` **once** on failure — matching the single retry the stable `e2e` (API 29–36) matrix already has, which `e2e-preview` lacked (so a flaky test self-heals on API 29–36 but would wedge the required gate on API 37, e.g. #370). Per-shard, so the retry re-runs only that shard's half of the suite. **Caveat:** a blanket retry can mask genuine regressions, so the retried-but-passed case is surfaced as a `::warning::` and the real flaky-test fix stays test-level (tracked separately). Retry count is one, not more.
- **`adb start-server` before the boot loop**, mirroring `api37_e2e.py`'s `wait_for_boot`, to avoid the attempt-1 adb-daemon "Address already in use" bind race seen in #370.
## Gate fan-in (both shards must pass — no branch-protection change)
`ci-passed` keeps `e2e-preview` as a **single** entry in its `needs:`. GHA's matrix aggregation makes `needs.e2e-preview.result` = `failure` if **either** shard fails, and the gate fails on any `failure`/`cancelled`/`skipped` — so **both shards must pass** for `CI passed` to go green. This is the same fan-in the API 29–36 `e2e` matrix already relies on. Branch protection requires the **`CI passed`** context (verified: `Debug build`, `Unit tests`, `CI passed`), **not** the per-leg `E2E (API 37 preview) (0/1)` check names, so **no branch-protection change is needed**. `fail-fast: false` keeps a failing shard from cancelling its sibling, so both reports always upload.
## Local preflight unchanged
`api37_e2e.py` / `local_instrumented.py` stay **single-emulator** — one free hypervisor per dev box; parallel local emulators risk freezing the machine. Sharding is CI-only by design. (`api37_e2e.py` already matches CI on the boot-retry loop and `adb start-server`; mirroring the single test-retry into it is a small documented follow-up, deliberately not bundled here because it can't be validated without booting a local emulator.)
## Validation
This is a CI-infra change, so **this PR's own CI run is the validation** — both `E2E (API 37 preview) (0)` and `(1)` shards must go green and satisfy `CI passed`, and the real per-shard wall-clock is readable straight off this run. YAML syntax, matrix expansion, the `numShards`/`shardIndex` args, artifact-name suffixing, `fail-fast: false`, and the retry block were sanity-checked locally.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
What & why
e2e-preview(the hand-provisioned API 37 /google_apis_ps16k16 KB-page emulator that ran the whole instrumented suite serially) is consistently the single longest leg in CI (~16.4–17.6 min), so it sets the pipeline's critical path. This shards it across 2 parallel API 37 emulators atN = 2, cutting the leg from ~17.1 → ~12.7 min — about ~4.9 min / ~28% off the total-CI critical path — for the cost of one extra emulator job.Design + measurements live in
docs/perf/api37-e2e-sharding-spike.md(updated from a feasibility spike to the adopted at N=2 design record).What changed (
.github/workflows/ci.yml,e2e-previewjob)strategy: { fail-fast: false, matrix: { shard: [0, 1] } }. Each leg hand-provisions its own API 37 emulator (provisioning unchanged) and runs one shard via AndroidJUnitRunner's built-in sharding, passed through AGP's existing runner-arg channel (the same onelocal_instrumented.pyuses for.class=):shardIndexis 0-based, so the matrix values are[0, 1]. No GMD, no orchestrator, no Gradle-side change.…-shard${{ matrix.shard }}on both the boot-diagnostics and the E2E-report uploads) —actions/upload-artifact@v7errors on duplicate artifact names.connectedDebugAndroidTestonce on failure — matching the single retry the stablee2e(API 29–36) matrix already has, whiche2e-previewlacked (so a flaky test self-heals on API 29–36 but would wedge the required gate on API 37, e.g. #370). Per-shard, so the retry re-runs only that shard's half of the suite. Caveat: a blanket retry can mask genuine regressions, so the retried-but-passed case is surfaced as a::warning::and the real flaky-test fix stays test-level (tracked separately). Retry count is one, not more.adb start-serverbefore the boot loop, mirroringapi37_e2e.py'swait_for_boot, to avoid the attempt-1 adb-daemon "Address already in use" bind race seen in #370.Gate fan-in (both shards must pass — no branch-protection change)
ci-passedkeepse2e-previewas a single entry in itsneeds:. GHA's matrix aggregation makesneeds.e2e-preview.result=failureif either shard fails, and the gate fails on anyfailure/cancelled/skipped— so both shards must pass forCI passedto go green. This is the same fan-in the API 29–36e2ematrix already relies on. Branch protection requires theCI passedcontext (verified:Debug build,Unit tests,CI passed), not the per-legE2E (API 37 preview) (0/1)check names, so no branch-protection change is needed.fail-fast: falsekeeps a failing shard from cancelling its sibling, so both reports always upload.Local preflight unchanged
api37_e2e.py/local_instrumented.pystay single-emulator — one free hypervisor per dev box; parallel local emulators risk freezing the machine. Sharding is CI-only by design. (api37_e2e.pyalready matches CI on the boot-retry loop andadb start-server; mirroring the single test-retry into it is a small documented follow-up, deliberately not bundled here because it can't be validated without booting a local emulator.)Validation
This is a CI-infra change, so this PR's own CI run is the validation — both
E2E (API 37 preview) (0)and(1)shards must go green and satisfyCI passed, and the real per-shard wall-clock is readable straight off this run. YAML syntax, matrix expansion, thenumShards/shardIndexargs, artifact-name suffixing,fail-fast: false, and the retry block were sanity-checked locally.🤖 Generated with Claude Code