Commit Graph
10 Commits
Author SHA1 Message Date
JMR-devandClaude Opus 4.8 08ee9e7abc ci: productionize API 37 preview E2E sharding (adopt N=2)
The spike commits already implemented the N=2 shard matrix (numShards/
shardIndex via -Pandroid.testInstrumentationRunnerArguments.*), per-shard
test-retry parity, adb start-server before the boot loop, and shard-suffixed
artifact names. This drops the SPIKE / DRAFT "do not merge as-is" framing from
the ci.yml comments and reframes docs/perf/api37-e2e-sharding-spike.md from a
feasibility spike into the adopted design, so the change is mergeable as-is.

Also fixes the doc's section 3a example, which showed 1-based shardIndex values
[1, 2]; shardIndex is 0-based (0..numShards-1) and the implementation correctly
uses matrix.shard: [0, 1] -- [1, 2] would run an empty bucket and silently drop
half the suite.

Fan-in unchanged and verified: ci-passed still lists e2e-preview once; GHA
matrix aggregation makes its result `failure` if either shard fails, so both
shards must pass for the gate to go green. Branch protection requires the
"CI passed" context (not the per-leg "E2E (API 37 preview) (N)" check names),
so no branch-protection change is needed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 14:52:48 -05:00
JMR-devandClaude Opus 4.8 0e9f54e686 ci(spike): retry parity + adb start-server for API 37 shard PoC
Folds the #370 root-cause finding into the spike. The stable `e2e` matrix
retries its test run once; `e2e-preview` runs connectedDebugAndroidTest exactly
once, so a flaky test self-heals on API 29-36 but wedges the required gate on
API 37 (e.g. #370's SignaturesScreenTest teardown race).

Doc: adds risk item 9 (retry-parity gap + its sharding interaction — per-test
flake is NOT amplified by sharding unlike boot flake, and a per-shard retry
costs only B + T/N; framed mitigation-not-fix) and two §6 recommendations
(retry parity, mirrored into api37_e2e.py; adb start-server before the boot
loop).

PoC (ci.yml): per-shard single test retry (::warning:: on retried-but-passed)
+ adb start-server before the boot loop. The api37_e2e.py retry mirror stays a
documented recommendation (local path needs a real-emulator validation this
spike did not boot). Still DRAFT, not auto-merged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 14:40:02 -05:00
JMR-devandClaude Opus 4.8 66da643249 ci(spike): PoC shard API 37 preview E2E + feasibility doc
Feasibility spike for sharding the e2e-preview job (the hand-provisioned
API 37 / google_apis_ps16k 16 KB-page emulator), CI's longest leg
(~16.4-17.6 min). docs/perf/api37-e2e-sharding-spike.md breaks the leg into
fixed overhead B ~8.3 min (setup + boot + Gradle daemon/config/compile/install)
vs parallelizable test execution T ~8.8 min, models B + T/N for N=2/3/4, and
recommends N=2 (~17.1 -> ~12.7 min, ~28% off the critical path) capped by the
API 30 matrix wall (~12.0 min) beyond N=3.

DRAFT PoC (do NOT merge as-is): converts e2e-preview to a strategy.matrix.shard
[0, 1] fan-out passing AndroidJUnitRunner numShards/shardIndex through the
existing -Pandroid.testInstrumentationRunnerArguments.* channel (no GMD, no
orchestrator, no Gradle change). Artifact names gain a shard suffix;
ci-passed still lists e2e-preview once (matrix fan-in keeps the single gate).
Local preflight stays single-emulator. Relates to #258.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 14:31:40 -05:00
JMR-devandClaude Opus 4.8 cc067408b8 perf(mail): enable IMAP connection reuse by default with a hardened cache
An on-device drilldown proved Gmail server-side throttles LibreMail's
connect-per-operation IMAP: every op was a fresh CONNECT+TLS+LOGIN, and
full-history backfill's body+attachment prefetch generated ~601 connections in
~22 min, tripping (and sustaining) Gmail's per-account rate/bandwidth clamp
(body download collapsed to ~4 KB/s). The `live` gauge peaked at only 5 (Gmail
allows ~15), so it is connection *volume*, not count. Outlook IMAP on the same
device opened in 2-3 s. Reusing one warm socket per account (~601 -> ~1) removes
the throttle's trigger. This wires the reuse path the #125 spike built and left
OFF (issue #357 Part 2 — connection reuse only; prefetch is a separate PR).

How it is enabled (with a safety switch):
- New `BuildConfig.IMAP_CONNECTION_REUSE` (default true) drives the production
  `ImapClient` no-arg `@Inject` constructor. To disable if a server misbehaves,
  flip it to "false" in app/build.gradle.kts — a build-config change, no Kotlin
  edit. The internal `ImapClient(reuseConnections, reuseIdleTimeoutMillis)`
  constructor stays the test/harness seam.
- Universal: applies to all providers (incl. Outlook). No per-provider caps or
  throttling here — that is a separate effort (#356/#360-#364).

Hardening `ImapConnectionCache` for production (was a spike):
- Transparent stale recovery: broadened drop detection to Angus's own
  `iap.ConnectionException` (and a MessagingException caused by one) — the real
  signal `folder.open()` throws on a server-dropped idle socket, which the
  IOException-only check missed, so the reconnect now actually fires. A dropped
  reused socket is rebuilt once and the op retried, so callers see no spurious
  error; a genuine app error (e.g. message-not-found) is never retried.
- Idle eviction: `evictIdle()` closes a connection unused past the reuse idle
  timeout (default 5 min), swept every 2 min by `IdleService`; skips any
  in-use connection.
- Teardown: `IdleService` also tears down reused connections on the low-battery
  push-teardown path (#88/#89/#90), mirroring the IDLE connection teardown.
- Concurrency: one connection per account behind a per-account mutex; the
  eviction sweep takes the lock non-blockingly so it never stalls or interrupts
  an in-flight op. Coexists with IMAP IDLE (its own separate connection).
- PII-free AppLog on the lifecycle (open / reuse-hit / reconnect-stale / evict /
  teardown) keyed by an opaque per-cache ordinal, plus the #358 ImapPerf
  breadcrumb (connect~=0ms on a reuse hit).

Tests (all via the fast gate, no emulator):
- ImapConnectionCacheTest: reuse, retry-once stale recovery, narrow drop
  detection, deterministic idle eviction (injected clock), teardown.
- ImapFolderOpenLatencyTest (GreenMail + counting proxy): N ops share one
  connection/LOGIN; a force-dropped socket is transparently reconnected; an app
  error does not reconnect; idle eviction LOGS-OUT and the next op reconnects.
- Correctness suites (ImapClientTest/ImapClientBackfillTest/MailBackfillerTest)
  pinned to reuse-off to keep their connect-per-op assertions unchanged.

Fast gate green: assembleDebug, testDebugUnitTest, compileDebugAndroidTestKotlin,
lintDebug, ktlintCheck, detekt.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 19:47:01 -05:00
Jason Ross f8ce83d5cb Merge branch 'main' into spike-imap-connection-reuse 2026-07-02 10:14:11 -05:00
Jason Ross 5994ee965e Merge branch 'main' into perf-unified-inbox-paging 2026-07-02 09:35:11 -05:00
JMR-devandClaude Fable 5 8bfc31f17c spike(imap): prototype flag-gated connection reuse for folder-open
Prototype the per-account connection reuse the #125 investigation recommended
and deferred, behind an OFF-by-default flag so it cannot destabilize `main`.

- ImapConnectionCache: keeps one authenticated Store alive per account, guarded
  by a per-account mutex, keyed by connection identity (not the rotating
  secret), with lazy catch-and-retry-once stale handling. No eviction policy
  yet beyond an explicit closeReusedConnections() hook.
- ImapClient gains a `reuseConnections` flag (default false via the @Inject
  no-arg constructor). With it off, withStore is byte-for-byte the previous
  connect + LOGOUT-per-call; with it on, calls borrow the kept-alive Store.
- ImapFolderOpenLatencyTest flips the flag on: the same real-IMAP operations
  that cost N connections / N LOGINs collapse to 1 connection / 1 LOGIN, with
  the necessary per-open EXAMINE unchanged (proven via CountingImapProxy +
  GreenMail; localhost is ~0 RTT so this proves structure, not wall-clock).
- docs/perf/issue-125-connection-reuse-spike.md: prototype design, the
  flag-off-vs-on proof, per-decision trade-offs, and the refined real-device
  validation plan. References #125; does not close it (needs device validation).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-02 09:29:55 -05:00
JMR-devandClaude Fable 5 f8d03a4343 perf(mailbox): page the unified "All inboxes" list (#124)
The unified inbox query (WHERE folder = ?, no accountId) has no folder-leading
index, so it scans in timestamp order and materializes the whole unified inbox
(~4k rows at a 20k cache) into memory on every emission. Apply Paging 3 to the
unified browse path so query, mapping, and recomposition cost scale with the
visible window, not the total cache.

- MessageDao.pagingUnifiedFolderSummaries: a PagingSource over the folder's
  synced rows (inInbox = 1); unified search keeps the whole-folder query so it
  can still surface transient server-search hits.
- MailRepository.pagedUnifiedFolderMessages: a Pager (pageSize 40, initialLoad
  120, no placeholders) mapping summaries to domain.
- MailboxViewModel.pagedMessages: paged while browsing the unified inbox, else
  empty; the messages list flow stays empty in that state so the whole cache is
  never materialized. Selection captures each row's accountId at tap time, so
  "Move" still resolves the selection's account without an in-memory list.
- MailboxScreen renders the unified browse list via collectAsLazyPagingItems;
  per-account and search views render the flat list unchanged (issue #86 stays
  flat).

Profiling (docs/perf/issue-124-unified-inbox-paging.md) on an api29 emulator:
current whole-inbox first-emit ~24.6 ms at a 20k cache vs. the paged first page
~6.8 ms and flat regardless of cache size (~3.6x). EXPLAIN QUERY PLAN shows the
paged query still stops early on the existing timestamp index, so no
(folder, ...) index and no schema migration are added.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-02 08:49:52 -05:00
JMR-devandClaude Fable 5 b0bb942a02 test(imap): measure folder-open round-trip structure (#125)
Investigate IMAP folder-open latency (follow-up to #86). Localhost GreenMail
has ~0 RTT, so real wall-clock latency can't be measured here; instead this
pins the folder-open round-trip STRUCTURE deterministically.

Finding: ImapClient.withStore wraps every operation in its own short-lived
Store, so each folder-open pays a full CONNECT + TLS + LOGIN + EXAMINE +
FETCH + LOGOUT. Only EXAMINE + FETCH is intrinsic to opening a folder; the
whole connection-setup group is avoidable on the 2nd+ operation if a
connection were reused. Optimistic render-from-cache already exists
(selectFolder renders cached rows; the network sync is a background refresh).

Adds:
- CountingImapProxy: a localhost TCP proxy that forwards a cleartext IMAP
  session to GreenMail while counting TCP connections and parsing IMAP
  command words.
- ImapFolderOpenLatencyTest: asserts the current no-reuse behaviour (N opens
  => N connections and N LOGINs; list+read => 2 connections) against a real
  in-process IMAP server. Doubles as the harness to validate a future
  connection-reuse fix (flip the counts to assert reuse).
- docs/perf/issue-125-imap-folder-open.md: the per-open round-trip sequence,
  avoidable vs. necessary round-trips, and the recommended per-account
  connection-reuse/keep-alive mitigation with its IDLE / thread-safety /
  battery / stale-connection constraints.

Analysis + harness only; the connection-reuse fix is deferred pending
real-network + real-device measurement (see the doc's measurement plan), so
this references #125 without closing it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-02 08:24:18 -05:00
JMR-devandClaude Fable 5 e7bb69d2ac perf(mailbox): scope the message-list query to the viewed folder in SQL
The mailbox list observed the entire `messages` table (observeSummaries, no
WHERE/LIMIT), mapped every cached row to a domain Message, and filtered down to
the visible account+folder in MailboxViewModel — so its cost scaled with the
whole cache and re-ran on every write to `messages` (IDLE delivery, a flag
toggle, a backfill page, any folder sync). On a 20k-row cache that is ~125 ms of
work per unrelated write.

Push the account/folder filter into SQL (observeFolderSummaries /
observeUnifiedFolderSummaries, exposed via observeFolderMessages /
observeUnifiedFolderMessages) and flatMapLatest the ViewModel over the selected
account+folder. The only remaining client-side pass separates the normal list
from an active search over the small folder-scoped set.

Validated on an emulator against 1k/5k/20k-row caches (docs/perf/issue-86-
profiling.md): the account-scoped query is ~1.5 ms flat (~80x faster at 20k) and
is already served by the existing (accountId, folder, uid) index — so NO
composite index and NO schema migration are added. The ticket's proposed
(accountId, folder, inInbox, timestampMillis) index changes timing only within
noise and isn't even preferred by SQLite's planner. The unified "All inboxes"
view stays an O(N) folder scan (still 5.6x better) and is a follow-up for paging;
IMAP latency on folder open is a separate, unmeasured concern.

Closes #86

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-02 07:16:12 -05:00