Add a CI check that runs performance benchmarks and fails a PR if it regresses beyond a threshold from a recorded baseline, so new code can't silently degrade performance.
What to measure
Start with paths already profiled and cared about (see docs/perf/):
Optionally hot data-layer paths (page query + summary mapping).
Candidate approaches (decide during design)
Jetpack Macrobenchmark (androidx.benchmark:benchmark-macro-junit4) — startup + frame timing/jank on real UI; instrumented (emulator/device). Closest to user-perceived perf.
Microbenchmark (androidx.benchmark:benchmark-junit4) — tight hot loops (e.g. summary mapping); instrumented, more stable.
JVM / JMH — pure-Kotlin logic with no Android deps.
(Baseline Profiles are orthogonal — they improve perf, they don't gate it.)
Baseline + gating mechanism
Record a baseline metric set (committed JSON, or pulled from the latest main run) and compare the PR's run against it; fail if a metric regresses beyond a threshold (e.g. >10–15% on the median over N iterations).
Decide baseline storage/refresh (committed-in-repo vs. artifact from main) and how a legitimate change gets re-baselined.
Key challenge — CI noise (do NOT ignore)
Shared GitHub runners have variable CPU and no consistent hardware, so absolute thresholds are flaky. Mitigate: enough iterations + statistical comparison (median / confidence-interval overlap, not single runs); thresholds tuned to observed variance; a dedicated/consistent runner class if possible; gate only on large regressions. Emulator Macrobenchmark on CI is especially noisy — validate signal-to-noise and start the check advisory/non-blocking, promote to required only once it's proven stable (mirrors how e2e-preview was handled).
Deliverable
A CI job (in .github/workflows/ci.yml or a dedicated workflow) that runs the chosen benchmark(s), compares to baseline, and reports/fails on regression, plus docs on updating the baseline. Keep the managed-device / E2E-matrix lockstep note (CLAUDE.md) in mind if it adds instrumented devices.
Add a CI check that runs performance benchmarks and **fails a PR if it regresses beyond a threshold from a recorded baseline**, so new code can't silently degrade performance.
## What to measure
Start with paths already profiled and cared about (see `docs/perf/`):
- **Cold startup** to first mailbox frame.
- **Mailbox list** scroll jank + first-page load (paged unified inbox, #124).
- Optionally hot data-layer paths (page query + summary mapping).
## Candidate approaches (decide during design)
- **Jetpack Macrobenchmark** (`androidx.benchmark:benchmark-macro-junit4`) — startup + frame timing/jank on real UI; instrumented (emulator/device). Closest to user-perceived perf.
- **Microbenchmark** (`androidx.benchmark:benchmark-junit4`) — tight hot loops (e.g. summary mapping); instrumented, more stable.
- **JVM / JMH** — pure-Kotlin logic with no Android deps.
- (Baseline Profiles are orthogonal — they improve perf, they don't gate it.)
## Baseline + gating mechanism
- Record a baseline metric set (committed JSON, or pulled from the latest `main` run) and compare the PR's run against it; fail if a metric regresses beyond a threshold (e.g. >10–15% on the median over N iterations).
- Decide baseline storage/refresh (committed-in-repo vs. artifact from `main`) and how a legitimate change gets re-baselined.
## Key challenge — CI noise (do NOT ignore)
Shared GitHub runners have variable CPU and no consistent hardware, so absolute thresholds are flaky. Mitigate: enough iterations + statistical comparison (median / confidence-interval overlap, not single runs); thresholds tuned to observed variance; a dedicated/consistent runner class if possible; gate only on large regressions. Emulator Macrobenchmark on CI is especially noisy — validate signal-to-noise and start the check **advisory/non-blocking**, promote to required only once it's proven stable (mirrors how `e2e-preview` was handled).
## Deliverable
A CI job (in `.github/workflows/ci.yml` or a dedicated workflow) that runs the chosen benchmark(s), compares to baseline, and reports/fails on regression, plus docs on updating the baseline. Keep the managed-device / E2E-matrix lockstep note (CLAUDE.md) in mind if it adds instrumented devices.
Investigation result: defer — the meaningful gate needs a device, and the cheap non-emulator proxies gate size, not the profiled runtime paths
Recommendation: do not add a perf-regression gate now. Every path #233 wants to protect (cold startup, mailbox scroll jank, data-layer query timing) is a Room/SQLite path that can only be measured on a device/emulator — and this repo's own profiling already established that emulator benchmarks yield only relative A/B signal, not the stable absolute baseline a regression gate requires. A gate built on that would be exactly the "flaky baseline = CI noise" failure the ticket warns against. The stable, non-emulator proxies (APK size / dex-method count) are real safeguards but measure a different thing (binary size, not runtime), so shipping one would mis-close #233 while leaving its actual intent unmet. Details + a "when you pick this up" recipe below, mirroring how #258 was handled.
1. What exists today: nothing to build on
No androidx.benchmark (micro or macro), no JMH, no @Benchmark / MacrobenchmarkRule, no :benchmark/:macrobenchmark module, no baseline profile. Clean slate.
CI (.github/workflows/ci.yml) builds debug only (assembleDebug); no assembleRelease/bundleRelease. The critical path is the E2E emulator matrix (API 29–36 + the API 37 preview job, ~15 min wall-clock per the #258 timings) — everything else fans in well under it.
2. Why the meaningful gate needs a device — and why that's premature now
The paths #233 names have all been profiled, and they are all Room/SQLite, not pure Kotlin:
Mailbox list / data-layer (docs/perf/issue-86-profiling.md): the bottleneck was a whole-messages-table SQL scan + cursor materialization; the fix was SQL-scoping the query. Measured on an AVD via EXPLAIN QUERY PLAN + wall-clock.
Unified inbox paging (docs/perf/issue-124-unified-inbox-paging.md): a Room PagingSource; first-page cost is a function of the SQLite LIMIT/planner. Measured on an AVD (+ a physical Pixel cross-check — not an emulator/AVD, tellingly).
Search is now paged DAO/SQL (MailboxViewModel.pagedMessages → pagedUnifiedSearchMessages / pagedFolderSearchMessages); the old in-memory matchesSearch filter was deleted (see MessageDao.kt — "the old in-memory search filter").
Cold startup to first frame is inherently an on-device macrobenchmark (startupCompose / frame timing).
Both perf docs deliberately rejected androidx.benchmark for two reasons that are fatal to a regression gate:
"on an emulator only the relative A/B result is meaningful (absolute nanos aren't representative of a real device)"
"it would have forced the module's global instrumentation runner to AndroidBenchmarkRunner, changing what CI's E2E jobs run under."
A regression gate is the opposite of an A/B run: it compares a PR's absolute numbers to a committed absolute baseline. Our own measurements say the emulator can't give a trustworthy absolute number — and shared GitHub runners (variable CPU, no fixed hardware — the ticket's own "Key challenge") make it worse. Avoiding the runner-hijack means a separate :macrobenchmark module, which in turn means new Gradle Managed Devices + new CI emulator jobs (and the CLAUDE.md managed-device ↔ CI-matrix lockstep) — i.e. more emulator jobs on an already-emulator-bound critical path, each carrying the noise the ticket flags. High cost, low trust, today.
3. Why the pure-Kotlin / JVM-microbenchmark route is low-signal here
A JVM (JMH-style) microbenchmark would be stable-ish and non-emulator, but there's no hot pure-Kotlin path worth gating:
Snippet.of / HtmlToText.convert are the obvious CPU-ish Kotlin candidates, but Snippet's own doc says derivation "runs once, when a body is fetched and cached … never per mailbox-list row." Same for HtmlToText (fetch/render-time). Gating them would guard a non-bottleneck.
The actual per-row / per-keystroke work is SQL (section 2), which JMH can't touch.
And JMH on shared runners is itself noisy — the ticket says so.
So this route spends real complexity to protect code that isn't the bottleneck, with a measurement whose noise the ticket already calls out. Net negative.
4. Why the stable non-emulator size proxies don't close#233
APK size / dex-method count / dependency count are genuinely stable and cheap — but:
Concern
Detail
Wrong axis
They gate binary size/complexity, not the runtime paths (startup, jank, query timing) #233 profiled and names. Closing #233 with one would be a mis-close.
Debug-only CI
CI builds only the unminified debug APK (Compose tooling + test manifest included) — a weak size proxy. A meaningful size gate wants the minified release APK, which CI doesn't build today, so doing it well isn't "minimal" (adds an assembleRelease job).
Baseline churn
In this high-velocity multi-agent repo (194 main .kt files, ~21 k LOC, 57 catalog libs, many concurrent feature PRs), a size baseline needs re-committing on every legitimate dep/feature bump — or a threshold so loose it's toothless. Friction the maintainer should opt into, not have imposed.
No demonstrated need
No size budget has been set or size problem observed. Adding CI machinery that doesn't yet pay for itself is the exact trap #258 avoided.
If a lightweight size safeguard is wanted, it's worth doing — but as its own small ticket, scoped to the release APK, kept advisory first, with an explicit re-baseline doc. It shouldn't ride in under #233's runtime-perf banner.
5. Recommended approach when this is picked up
When there's appetite for a real runtime gate (and ideally a more consistent runner class):
New :macrobenchmark module (androidx.benchmark:benchmark-macro-junit4), its own AndroidBenchmarkRunner — keeps the :app E2E runner untouched.
Start with StartupTimingMetric (cold start) + FrameTimingMetric (mailbox scroll) — the two the perf docs care about most.
Advisory / non-blocking first — same playbook as e2e-preview: prove signal-to-noise across ~20+ runs before it's ever a merge gate.
Compare median over N iterations with confidence-interval overlap, not single runs; gate only on large regressions (≥15–20% over observed variance).
Baseline: prefer pulling the latest green main run's numbers over a committed JSON, so re-baselining is automatic on merge rather than a manual commit.
Keep the managed-device ↔ CI-matrix lockstep (CLAUDE.md) in mind — the new instrumented device is more emulator surface.
Do-it / re-open trigger: pick this up when (a) there's a consistent runner class (self-hosted or a paid larger runner) to tame CI noise, or (b) a concrete perf regression actually ships and motivates a guard — then implement section 5, advisory-first. Until then a perf gate is complexity + flake-surface for signal we can't yet trust. Happy to spin off the optional release-APK-size safeguard (section 4) as its own ticket if that interim guard is wanted.
## Investigation result: **defer** — the meaningful gate needs a device, and the cheap non-emulator proxies gate *size*, not the profiled runtime paths
Recommendation: **do not add a perf-regression gate now.** Every path #233 wants to protect (cold startup, mailbox scroll jank, data-layer query timing) is a Room/SQLite path that can only be measured on a device/emulator — and this repo's own profiling already established that emulator benchmarks yield only **relative A/B** signal, not the **stable absolute baseline** a regression gate requires. A gate built on that would be exactly the "flaky baseline = CI noise" failure the ticket warns against. The stable, non-emulator proxies (APK size / dex-method count) are real safeguards but measure a *different* thing (binary size, not runtime), so shipping one would mis-close #233 while leaving its actual intent unmet. Details + a "when you pick this up" recipe below, mirroring how #258 was handled.
### 1. What exists today: nothing to build on
- No `androidx.benchmark` (micro or macro), no JMH, no `@Benchmark` / `MacrobenchmarkRule`, no `:benchmark`/`:macrobenchmark` module, no baseline profile. Clean slate.
- CI (`.github/workflows/ci.yml`) builds **debug only** (`assembleDebug`); no `assembleRelease`/`bundleRelease`. The critical path is the **E2E emulator matrix** (API 29–36 + the API 37 preview job, ~15 min wall-clock per the #258 timings) — everything else fans in well under it.
### 2. Why the *meaningful* gate needs a device — and why that's premature now
The paths #233 names have all been profiled, and they are **all Room/SQLite**, not pure Kotlin:
- **Mailbox list / data-layer** (`docs/perf/issue-86-profiling.md`): the bottleneck was a whole-`messages`-table SQL scan + cursor materialization; the fix was SQL-scoping the query. Measured on an AVD via `EXPLAIN QUERY PLAN` + wall-clock.
- **Unified inbox paging** (`docs/perf/issue-124-unified-inbox-paging.md`): a Room `PagingSource`; first-page cost is a function of the SQLite `LIMIT`/planner. Measured on an AVD (+ a physical Pixel cross-check — *not* an emulator/AVD, tellingly).
- **Search** is now paged **DAO/SQL** (`MailboxViewModel.pagedMessages` → `pagedUnifiedSearchMessages` / `pagedFolderSearchMessages`); the old in-memory `matchesSearch` filter was deleted (see `MessageDao.kt` — "the old in-memory search filter").
- **Cold startup to first frame** is inherently an on-device macrobenchmark (`startupCompose` / frame timing).
Both perf docs **deliberately rejected `androidx.benchmark`** for two reasons that are fatal to a regression *gate*:
> "on an emulator only the **relative** A/B result is meaningful (absolute nanos aren't representative of a real device)"
> "it would have forced the module's global instrumentation runner to `AndroidBenchmarkRunner`, changing what CI's E2E jobs run under."
A regression gate is the opposite of an A/B run: it compares a PR's **absolute** numbers to a committed **absolute** baseline. Our own measurements say the emulator can't give a trustworthy absolute number — and shared GitHub runners (variable CPU, no fixed hardware — the ticket's own "Key challenge") make it worse. Avoiding the runner-hijack means a **separate `:macrobenchmark` module**, which in turn means **new Gradle Managed Devices + new CI emulator jobs** (and the CLAUDE.md managed-device ↔ CI-matrix lockstep) — i.e. *more* emulator jobs on an already-emulator-bound critical path, each carrying the noise the ticket flags. High cost, low trust, today.
### 3. Why the pure-Kotlin / JVM-microbenchmark route is low-signal here
A JVM (JMH-style) microbenchmark would be stable-ish and non-emulator, but there's **no hot pure-Kotlin path worth gating**:
- `Snippet.of` / `HtmlToText.convert` are the obvious CPU-ish Kotlin candidates, but `Snippet`'s own doc says derivation "runs **once**, when a body is fetched and cached … **never per mailbox-list row**." Same for `HtmlToText` (fetch/render-time). Gating them would guard a non-bottleneck.
- The actual per-row / per-keystroke work is SQL (section 2), which JMH can't touch.
- And JMH on shared runners is itself noisy — the ticket says so.
So this route spends real complexity to protect code that isn't the bottleneck, with a measurement whose noise the ticket already calls out. Net negative.
### 4. Why the stable non-emulator *size* proxies don't close #233
APK size / dex-method count / dependency count are genuinely stable and cheap — but:
| Concern | Detail |
|---|---|
| **Wrong axis** | They gate **binary size/complexity**, not the **runtime** paths (startup, jank, query timing) #233 profiled and names. Closing #233 with one would be a mis-close. |
| **Debug-only CI** | CI builds only the unminified **debug** APK (Compose tooling + test manifest included) — a weak size proxy. A meaningful size gate wants the **minified release** APK, which CI doesn't build today, so doing it *well* isn't "minimal" (adds an `assembleRelease` job). |
| **Baseline churn** | In this high-velocity multi-agent repo (194 main `.kt` files, ~21 k LOC, 57 catalog libs, many concurrent feature PRs), a size baseline needs re-committing on every legitimate dep/feature bump — or a threshold so loose it's toothless. Friction the maintainer should opt into, not have imposed. |
| **No demonstrated need** | No size budget has been set or size problem observed. Adding CI machinery that doesn't yet pay for itself is the exact trap #258 avoided. |
If a lightweight **size** safeguard is wanted, it's worth doing — but as its **own small ticket**, scoped to the **release** APK, kept **advisory** first, with an explicit re-baseline doc. It shouldn't ride in under #233's runtime-perf banner.
### 5. Recommended approach *when this is picked up*
When there's appetite for a real runtime gate (and ideally a more consistent runner class):
1. New **`:macrobenchmark`** module (`androidx.benchmark:benchmark-macro-junit4`), its own `AndroidBenchmarkRunner` — keeps the `:app` E2E runner untouched.
2. Start with **`StartupTimingMetric`** (cold start) + **`FrameTimingMetric`** (mailbox scroll) — the two the perf docs care about most.
3. **Advisory / non-blocking first** — same playbook as `e2e-preview`: prove signal-to-noise across ~20+ runs before it's ever a merge gate.
4. Compare **median over N iterations** with **confidence-interval overlap**, not single runs; gate only on **large** regressions (≥15–20% over observed variance).
5. Baseline: prefer **pulling the latest green `main` run's numbers** over a committed JSON, so re-baselining is automatic on merge rather than a manual commit.
6. Keep the **managed-device ↔ CI-matrix lockstep** (CLAUDE.md) in mind — the new instrumented device is more emulator surface.
### Bottom line
| Option | Non-emulator? | Measures #233's paths? | Stable enough to gate now? | Verdict |
|---|---|---|---|---|
| Macrobenchmark (startup/jank) on CI emulator | ✗ | ✓ | ✗ (relative-only per our docs; runner noise) | Defer — advisory-first later |
| JVM/JMH microbenchmark | ✓ | ✗ (hot paths are SQL; Kotlin candidates are once-per-fetch) | ~ (noisy on shared runners) | ✗ low-signal |
| APK-size / dex-count ceiling | ✓ | ✗ (size, not runtime) | ✓ | Optional **separate** ticket, not #233 |
| **Do nothing now** | — | — | — | ✓ |
**Do-it / re-open trigger:** pick this up when (a) there's a consistent runner class (self-hosted or a paid larger runner) to tame CI noise, **or** (b) a concrete perf regression actually ships and motivates a guard — then implement section 5, advisory-first. Until then a perf gate is complexity + flake-surface for signal we can't yet trust. Happy to spin off the optional release-APK-size safeguard (section 4) as its own ticket if that interim guard is wanted.
Closing as investigated → deferred (details in prior comment): no benchmark scaffolding exists; the perf-sensitive paths are device-only Room/SQLite where emulator baselines are too noisy for a stable regression gate (the repo's perf docs already rejected androidx.benchmark for this reason); size/dex proxies gate the wrong axis and would churn in this high-velocity repo. Re-open when a dedicated :macrobenchmark module + stable baseline are worth the maintenance (recipe in the comment). Optional interim APK-size-ceiling safeguard noted there if wanted.
Closing as investigated → deferred (details in prior comment): no benchmark scaffolding exists; the perf-sensitive paths are device-only Room/SQLite where emulator baselines are too noisy for a stable regression gate (the repo's perf docs already rejected androidx.benchmark for this reason); size/dex proxies gate the wrong axis and would churn in this high-velocity repo. Re-open when a dedicated :macrobenchmark module + stable baseline are worth the maintenance (recipe in the comment). Optional interim APK-size-ceiling safeguard noted there if wanted.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Add a CI check that runs performance benchmarks and fails a PR if it regresses beyond a threshold from a recorded baseline, so new code can't silently degrade performance.
What to measure
Start with paths already profiled and cared about (see
docs/perf/):Candidate approaches (decide during design)
androidx.benchmark:benchmark-macro-junit4) — startup + frame timing/jank on real UI; instrumented (emulator/device). Closest to user-perceived perf.androidx.benchmark:benchmark-junit4) — tight hot loops (e.g. summary mapping); instrumented, more stable.Baseline + gating mechanism
mainrun) and compare the PR's run against it; fail if a metric regresses beyond a threshold (e.g. >10–15% on the median over N iterations).main) and how a legitimate change gets re-baselined.Key challenge — CI noise (do NOT ignore)
Shared GitHub runners have variable CPU and no consistent hardware, so absolute thresholds are flaky. Mitigate: enough iterations + statistical comparison (median / confidence-interval overlap, not single runs); thresholds tuned to observed variance; a dedicated/consistent runner class if possible; gate only on large regressions. Emulator Macrobenchmark on CI is especially noisy — validate signal-to-noise and start the check advisory/non-blocking, promote to required only once it's proven stable (mirrors how
e2e-previewwas handled).Deliverable
A CI job (in
.github/workflows/ci.ymlor a dedicated workflow) that runs the chosen benchmark(s), compares to baseline, and reports/fails on regression, plus docs on updating the baseline. Keep the managed-device / E2E-matrix lockstep note (CLAUDE.md) in mind if it adds instrumented devices.Investigation result: defer — the meaningful gate needs a device, and the cheap non-emulator proxies gate size, not the profiled runtime paths
Recommendation: do not add a perf-regression gate now. Every path #233 wants to protect (cold startup, mailbox scroll jank, data-layer query timing) is a Room/SQLite path that can only be measured on a device/emulator — and this repo's own profiling already established that emulator benchmarks yield only relative A/B signal, not the stable absolute baseline a regression gate requires. A gate built on that would be exactly the "flaky baseline = CI noise" failure the ticket warns against. The stable, non-emulator proxies (APK size / dex-method count) are real safeguards but measure a different thing (binary size, not runtime), so shipping one would mis-close #233 while leaving its actual intent unmet. Details + a "when you pick this up" recipe below, mirroring how #258 was handled.
1. What exists today: nothing to build on
androidx.benchmark(micro or macro), no JMH, no@Benchmark/MacrobenchmarkRule, no:benchmark/:macrobenchmarkmodule, no baseline profile. Clean slate..github/workflows/ci.yml) builds debug only (assembleDebug); noassembleRelease/bundleRelease. The critical path is the E2E emulator matrix (API 29–36 + the API 37 preview job, ~15 min wall-clock per the #258 timings) — everything else fans in well under it.2. Why the meaningful gate needs a device — and why that's premature now
The paths #233 names have all been profiled, and they are all Room/SQLite, not pure Kotlin:
docs/perf/issue-86-profiling.md): the bottleneck was a whole-messages-table SQL scan + cursor materialization; the fix was SQL-scoping the query. Measured on an AVD viaEXPLAIN QUERY PLAN+ wall-clock.docs/perf/issue-124-unified-inbox-paging.md): a RoomPagingSource; first-page cost is a function of the SQLiteLIMIT/planner. Measured on an AVD (+ a physical Pixel cross-check — not an emulator/AVD, tellingly).MailboxViewModel.pagedMessages→pagedUnifiedSearchMessages/pagedFolderSearchMessages); the old in-memorymatchesSearchfilter was deleted (seeMessageDao.kt— "the old in-memory search filter").startupCompose/ frame timing).Both perf docs deliberately rejected
androidx.benchmarkfor two reasons that are fatal to a regression gate:A regression gate is the opposite of an A/B run: it compares a PR's absolute numbers to a committed absolute baseline. Our own measurements say the emulator can't give a trustworthy absolute number — and shared GitHub runners (variable CPU, no fixed hardware — the ticket's own "Key challenge") make it worse. Avoiding the runner-hijack means a separate
:macrobenchmarkmodule, which in turn means new Gradle Managed Devices + new CI emulator jobs (and the CLAUDE.md managed-device ↔ CI-matrix lockstep) — i.e. more emulator jobs on an already-emulator-bound critical path, each carrying the noise the ticket flags. High cost, low trust, today.3. Why the pure-Kotlin / JVM-microbenchmark route is low-signal here
A JVM (JMH-style) microbenchmark would be stable-ish and non-emulator, but there's no hot pure-Kotlin path worth gating:
Snippet.of/HtmlToText.convertare the obvious CPU-ish Kotlin candidates, butSnippet's own doc says derivation "runs once, when a body is fetched and cached … never per mailbox-list row." Same forHtmlToText(fetch/render-time). Gating them would guard a non-bottleneck.So this route spends real complexity to protect code that isn't the bottleneck, with a measurement whose noise the ticket already calls out. Net negative.
4. Why the stable non-emulator size proxies don't close #233
APK size / dex-method count / dependency count are genuinely stable and cheap — but:
assembleReleasejob)..ktfiles, ~21 k LOC, 57 catalog libs, many concurrent feature PRs), a size baseline needs re-committing on every legitimate dep/feature bump — or a threshold so loose it's toothless. Friction the maintainer should opt into, not have imposed.If a lightweight size safeguard is wanted, it's worth doing — but as its own small ticket, scoped to the release APK, kept advisory first, with an explicit re-baseline doc. It shouldn't ride in under #233's runtime-perf banner.
5. Recommended approach when this is picked up
When there's appetite for a real runtime gate (and ideally a more consistent runner class):
:macrobenchmarkmodule (androidx.benchmark:benchmark-macro-junit4), its ownAndroidBenchmarkRunner— keeps the:appE2E runner untouched.StartupTimingMetric(cold start) +FrameTimingMetric(mailbox scroll) — the two the perf docs care about most.e2e-preview: prove signal-to-noise across ~20+ runs before it's ever a merge gate.mainrun's numbers over a committed JSON, so re-baselining is automatic on merge rather than a manual commit.Bottom line
Do-it / re-open trigger: pick this up when (a) there's a consistent runner class (self-hosted or a paid larger runner) to tame CI noise, or (b) a concrete perf regression actually ships and motivates a guard — then implement section 5, advisory-first. Until then a perf gate is complexity + flake-surface for signal we can't yet trust. Happy to spin off the optional release-APK-size safeguard (section 4) as its own ticket if that interim guard is wanted.
Closing as investigated → deferred (details in prior comment): no benchmark scaffolding exists; the perf-sensitive paths are device-only Room/SQLite where emulator baselines are too noisy for a stable regression gate (the repo's perf docs already rejected androidx.benchmark for this reason); size/dex proxies gate the wrong axis and would churn in this high-velocity repo. Re-open when a dedicated :macrobenchmark module + stable baseline are worth the maintenance (recipe in the comment). Optional interim APK-size-ceiling safeguard noted there if wanted.