Feasibility spike for sharding the e2e-preview job (the hand-provisioned
API 37 / google_apis_ps16k 16 KB-page emulator), CI's longest leg
(~16.4-17.6 min). docs/perf/api37-e2e-sharding-spike.md breaks the leg into
fixed overhead B ~8.3 min (setup + boot + Gradle daemon/config/compile/install)
vs parallelizable test execution T ~8.8 min, models B + T/N for N=2/3/4, and
recommends N=2 (~17.1 -> ~12.7 min, ~28% off the critical path) capped by the
API 30 matrix wall (~12.0 min) beyond N=3.
DRAFT PoC (do NOT merge as-is): converts e2e-preview to a strategy.matrix.shard
[0, 1] fan-out passing AndroidJUnitRunner numShards/shardIndex through the
existing -Pandroid.testInstrumentationRunnerArguments.* channel (no GMD, no
orchestrator, no Gradle change). Artifact names gain a shard suffix;
ci-passed still lists e2e-preview once (matrix fan-in keeps the single gate).
Local preflight stays single-emulator. Relates to #258.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Restructure isConnectionDrop as leading guard clauses (definite-drop
types, then a not-MessagingException early return) instead of a when
expression, per maintainer review feedback on PR #368. Behavior is
unchanged; verified by the existing ImapConnectionCacheTest suite
(all 8 cases still pass), including the FolderClosedException /
StoreClosedException cases that depend on the check running before
the MessagingException .cause guard.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Resolve the IdleService.kt conflict as a union of both intents:
- #354 (already on main): foreground-service lifecycle rework —
onStartCommand delegates to the IdleForegroundStarter seam
(START_NOT_STICKY), cap-window skip/degrade.
- #357 Part 2 / #368: reused-connection idle-eviction sweep and
low-battery teardown of reused connections.
In startWatchingIfNeeded(), reconcileWatchers() stays inside the
cache-lock-guarded launch and evictIdleReuseConnectionsLoop() launches as
a sibling coroutine that runs while the service lives (its original #368
placement, independent of the cache-lock guard). No behavior change to
either side.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
An on-device drilldown proved Gmail server-side throttles LibreMail's
connect-per-operation IMAP: every op was a fresh CONNECT+TLS+LOGIN, and
full-history backfill's body+attachment prefetch generated ~601 connections in
~22 min, tripping (and sustaining) Gmail's per-account rate/bandwidth clamp
(body download collapsed to ~4 KB/s). The `live` gauge peaked at only 5 (Gmail
allows ~15), so it is connection *volume*, not count. Outlook IMAP on the same
device opened in 2-3 s. Reusing one warm socket per account (~601 -> ~1) removes
the throttle's trigger. This wires the reuse path the #125 spike built and left
OFF (issue #357 Part 2 — connection reuse only; prefetch is a separate PR).
How it is enabled (with a safety switch):
- New `BuildConfig.IMAP_CONNECTION_REUSE` (default true) drives the production
`ImapClient` no-arg `@Inject` constructor. To disable if a server misbehaves,
flip it to "false" in app/build.gradle.kts — a build-config change, no Kotlin
edit. The internal `ImapClient(reuseConnections, reuseIdleTimeoutMillis)`
constructor stays the test/harness seam.
- Universal: applies to all providers (incl. Outlook). No per-provider caps or
throttling here — that is a separate effort (#356/#360-#364).
Hardening `ImapConnectionCache` for production (was a spike):
- Transparent stale recovery: broadened drop detection to Angus's own
`iap.ConnectionException` (and a MessagingException caused by one) — the real
signal `folder.open()` throws on a server-dropped idle socket, which the
IOException-only check missed, so the reconnect now actually fires. A dropped
reused socket is rebuilt once and the op retried, so callers see no spurious
error; a genuine app error (e.g. message-not-found) is never retried.
- Idle eviction: `evictIdle()` closes a connection unused past the reuse idle
timeout (default 5 min), swept every 2 min by `IdleService`; skips any
in-use connection.
- Teardown: `IdleService` also tears down reused connections on the low-battery
push-teardown path (#88/#89/#90), mirroring the IDLE connection teardown.
- Concurrency: one connection per account behind a per-account mutex; the
eviction sweep takes the lock non-blockingly so it never stalls or interrupts
an in-flight op. Coexists with IMAP IDLE (its own separate connection).
- PII-free AppLog on the lifecycle (open / reuse-hit / reconnect-stale / evict /
teardown) keyed by an opaque per-cache ordinal, plus the #358 ImapPerf
breadcrumb (connect~=0ms on a reuse hit).
Tests (all via the fast gate, no emulator):
- ImapConnectionCacheTest: reuse, retry-once stale recovery, narrow drop
detection, deterministic idle eviction (injected clock), teardown.
- ImapFolderOpenLatencyTest (GreenMail + counting proxy): N ops share one
connection/LOGIN; a force-dropped socket is transparently reconnected; an app
error does not reconnect; idle eviction LOGS-OUT and the next op reconnects.
- Correctness suites (ImapClientTest/ImapClientBackfillTest/MailBackfillerTest)
pinned to reuse-off to keep their connect-per-op assertions unchanged.
Fast gate green: assembleDebug, testDebugUnitTest, compileDebugAndroidTestKotlin,
lintDebug, ktlintCheck, detekt.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Root cause: after #302's runtime-cap fallBackToPeriodicSync() stops the
dataSync foreground service, IdleService was restarted (START_STICKY
null-intent redelivery + explicit startForegroundService) and onStartCommand
unconditionally called startForeground(DATA_SYNC) while the rolling-24h budget
was still exhausted. The platform rejected the start with
ForegroundServiceStartNotAllowedException; it was uncaught, the process
crashed, and START_STICKY restarted straight back into the same rejection -- a
crash loop until the 24h window freed budget (#354).
Fix (IdleService.kt):
- onStartCommand now returns START_NOT_STICKY. Push is app-managed
(LibreMailApplication.ensurePushStarted deterministically restarts it), so the
sticky null-intent auto-restart was redundant and fired exactly when a dataSync
FGS start is illegal.
- Guard the foreground start via a new JVM-testable IdleForegroundStarter seam:
a ForegroundServiceStartNotAllowedException (caught via its IllegalStateException
supertype, so no minSdk-29 class load) degrades like the cap handler --
schedulePeriodicSync(), keep the degraded POLLING notification, stopSelf()
promptly (avoids the "did not call startForeground in time" ANR) -- instead of
propagating.
- Record the cap event (elapsedRealtime); while still inside the cap window,
onStartCommand skips the now-guaranteed-illegal foreground start entirely.
- onTimeout stop path kept fast so ForegroundServiceDidNotStopInTimeException
stays mitigated.
PII-free AppLog.w/i on the degrade paths.
Tests:
- Unit (IdleForegroundStarterTest): onStartCommand returns START_NOT_STICKY; a
rejected start is caught and routed to degrade without propagating; the cap
window skips the attempt; a non-ISE propagates.
- Instrumented (IdleServiceForegroundStartInstrumentedTest): the degrade path on
a real Context -- rejection caught, periodic-sync fallback scheduled, degraded
"instant delivery paused" notification built, watching skipped.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adds AppLog breadcrumbs to the message-open path so a debug report can show
where the reader's spinner time goes:
- ImapClient.withStore: per-op connect vs. work timing plus a live
connect-per-op connection gauge (issue #125's provider-ceiling context).
- fetchBodyMarkingSeen: select/body/flag phase timings plus PII-free size
counts (RFC822 size, body chars, attachment count).
- MailRepositoryImpl.openMessage: end-to-end open latency plus the
cached-vs-fetched branch, keyed by accountLogRef and logSafeFolderLabel.
- ReaderViewModel: spinner-to-ready latency, split success vs. failure.
All breadcrumbs are PII-free: accounts are logged via the existing
accountLogRef hash, folders via the existing logSafeFolderLabel allowlist,
and everything else is sizes/durations/booleans only.
Fixes the 4 unit-test classes that exercise this code without mocking
android.util.Log (a throwing stub under plain JVM tests): mockkStatic(Log)
is now installed in MailRepositoryImplCoverageTest, ImapClientBackfillTest,
ImapFolderOpenLatencyTest, and ReaderViewModelActionsTest, following the
existing MailBackfillerTest/ImapClientTest conventions. detekt.yml gains two
more ForbiddenImport excludes for the newly Log-importing test files.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Move the traffic-control (runner-priority orchestration) job verbatim out of
.github/workflows/ci.yml into a new standalone workflow,
.github/workflows/traffic-control.yml, so the heavy CI jobs no longer depend
on it. The job's YAML (name, runs-on, timeout-minutes, permissions, env,
steps) and its documentation comment move unchanged; the decision core
.github/scripts/traffic_control.py is untouched and still unit-tested by the
traffic-control-tests job in ci.yml.
In ci.yml: removed the traffic-control job, dropped needs: traffic-control
from the five heavy jobs (static-analysis, debug-build, unit-tests, e2e,
e2e-preview) and from traffic-control-tests (its only needs, which would
otherwise dangle at a now-deleted job), and updated the now-stale header and
ci-passed comments to point at the extracted workflow.
The new workflow will be disabled pending a rebuild as a published GitHub
Action.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
#350 made ci-trigger.yml dispatch ci.yml with the built-in GITHUB_TOKEN, on the
claim that a workflow_dispatch is anti-recursion-exempt so no PAT is needed. In
practice a GITHUB_TOKEN-triggered run is held in `action_required` awaiting manual
approval and never runs un-attended, so auto-updated PRs' CI never ran (stalled
#285). The original #349 design was right: dispatch with a PAT so the run executes
as the authorized owner with no approval gate.
- ci-trigger.yml: the trigger step's GH_TOKEN is now
`${{ secrets.AUTOUPDATE_TOKEN || github.token }}` (was `${{ github.token }}`).
AUTOUPDATE_TOKEN (the PAT) is REQUIRED for the scheduler; the `|| github.token`
fallback stays fail-open but only starts CI if repo settings don't gate
GITHUB_TOKEN-triggered runs.
- autoupdate.yml: branch update stays on GITHUB_TOKEN (must NOT retrigger CI --
that would re-introduce the cascade). Clarified that AUTOUPDATE_TOKEN is still
required by the repo (by ci-trigger.yml) so the secret isn't deleted.
- Corrected the now-wrong "no PAT needed / workflow_dispatch anti-recursion-exempt"
comments in ci-trigger.yml and the traffic_control.py docstrings.
updates = GITHUB_TOKEN, triggering = PAT.
Validation: all three workflow YAMLs parse clean; traffic-control unit tests still
pass (59 tests) -- the change is workflow-env only, script logic unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Resolve DatabaseModule conflict from #320: main replaced the explicit .addMigrations(...) chain with .addMigrations(*ALL_MIGRATIONS) plus an introspectable ALL_MIGRATIONS list guarded by databaseModuleRegistersEveryDeclaredMigration (registered == declared). Add MIGRATION_19_20 to ALL_MIGRATIONS so the unified-inbox covering-index migration (cache schema v19->v20) is both registered on the Room builder and satisfies that safety-net test. Schema 20.json, the v20 @Database version, and DatabaseEncryptionTest's schema-version assertion (20) are unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
End the merge cascade and give the traffic-controller ownership of CI *triggering*.
- autoupdate.yml updates PR branches with the built-in GITHUB_TOKEN instead of a PAT,
so an update push no longer auto-retriggers CI (GitHub's anti-recursion rule) — the
cascade (every merge re-runs every PR, cancel-in-progress thrashing them) is gone.
- New scheduler ci-trigger.yml -> traffic_control.py --mode trigger (re-)triggers CI
for the highest-priority PR(s) whose head SHA has absent/stale checks, a few at a
time (inflight cap), in the existing P0-P9 / broken-draft priority order — a
poor-man's merge queue reusing the priority core. It runs after autoupdate finishes
(workflow_run, race-free) plus a cron backstop plus manual dispatch.
- Triggering uses workflow_dispatch, which is EXEMPT from anti-recursion, so the
built-in GITHUB_TOKEN (actions: write) starts the run — NO PAT / secret change needed.
- ci.yml gains a workflow_dispatch trigger (pr/head_sha/reason inputs) and a per-PR
concurrency group unifying pull_request and dispatch runs; its on: pull_request path
is kept so brand-new PRs, human pushes, and fork PRs always get CI (fail-open).
Pure select_triggers / classify_sha_runs decision core added to traffic_control.py with
24 new unit tests (priority order, oldest-first fairness, inflight cap, fork skip, P0
bypass+preempt, head-SHA needy classification, and a liveness/anti-starvation simulation).
Closes#349
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add a detekt style>ForbiddenImport rule that forbids `import android.util.Log`
so all logging flows through org.libremail.reporting.AppLog, which mirrors each
line into the debug-report RingLogBuffer. A raw android.util.Log import writes to
Logcat only and never reaches a user-reviewed DebugReport (epic #324, strangler
final step).
Excludes the AppLog facade itself (the one sanctioned wrapper) and the unit tests
that mockkStatic(Log) to verify forwarding — AppLog forwards to Log, a throwing
stub under plain JVM unit tests, so those tests must mock it; they do not bypass
the facade.
Closes#331
Part of #324
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adds a fast traffic-control-tests job (ubuntu, actions/checkout +
actions/setup-python, no emulator/Gradle) that runs the 37 pure-stdlib
unit tests for .github/scripts/traffic_control.py on every PR, and
wires it into ci-passed's needs so a regression blocks merge instead
of only being caught locally.
Closes#346
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>