Author SHA1 Message Date
Jason Ross 8597a0d4d4 Merge branch 'main' into spike-359-sqlcipher-ondevice 2026-07-07 16:03:10 -05:00
Jason Ross 6be3b41dd1 Merge pull request #424 from JMR-dev/feat-369-encrypt-reports-at-rest
feat(security): encrypt persisted reports at rest when cache encryption is on (#369)
2026-07-07 16:02:03 -05:00
JMR-dev 099474d32a feat(security): encrypt persisted reports at rest when cache encryption is on (#369)
When the opt-in cache encryption (encryptCache) is ON, persist crash and
"Report a problem" reports encrypted at rest, decrypting them on read; when
OFF they stay plaintext exactly as before.

- ReportStore gains a ReportEncryption collaborator (default None = plaintext,
  so existing call sites are unchanged). On write it seals the storage JSON with
  AES-256-GCM and tags it with a marker prefix; on read it sniffs the prefix, so
  pre-toggle plaintext and post-toggle sealed reports coexist. Writes FAIL
  CLOSED: a sealing failure drops the report rather than leaving plaintext on
  disk. Decrypt failures are logged (PII-free) and skipped.
- KeystoreReportEncryption reuses the vetted KeystoreCrypto (non-auth master
  key, so a crash while the app is locked can still seal), and mirrors the
  encryptCache setting into a crash-safe in-memory flag warmed at startup (no
  DataStore read on the crashing thread).
- PII-free AppLog logging at the enable/disable transition and both fallback
  paths; never logs report contents.

Tests: JVM unit tests for the ReportStore branching (seal-on-write, plaintext
when off, crash persistence, fail-closed, mixed files, decrypt-failure skip,
markSurfaced re-seal) and for KeystoreReportEncryption; an instrumented test
proves real Keystore ciphertext on disk + round-trip on device.
2026-07-07 14:55:28 -05:00
JMR-dev c04825a637 spike(security): characterize #359 SQLCipher SDK-37/16KB-page failure on-device
Add SqlCipherOpenSpikeTest (androidTest), a focused on-device harness for #359
that exercises the encrypted-cache open path in three isolated stages so a
failure pinpoints the break site: Stage A loads libsqlcipher.so
(System.loadLibrary), Stage B reaches SQLiteConnection.nativeOpen via a keyed
open, Stage C opens the full Room encrypted cache through the production
SupportOpenHelperFactory. Each stage records the in-process page size
(Os.sysconf _SC_PAGESIZE) and re-raises the full UnsatisfiedLinkError (which
.so, cause chain, stacktrace) on failure. Investigation only; no app/src/main
crypto change.

On-device A/B finding (Pixel 8 Pro, husky, real SDK 37 / Android 17):
- 4 KB pages (PAGE_SIZE=4096): Stages A, B, C ALL PASS.
- 16 KB pages: NOT tested on-device — the Pixel "Boot with 16 KB page size"
  toggle is gated behind an unlocked bootloader ("All user data and settings
  will be wiped when activating 16 KB mode"), i.e. destructive + out of scope.

Static ELF proof (refutes the #359 root-cause hypothesis): every bundled
native library is already 16 KB-aligned (all PT_LOAD p_align = 0x4000),
including arm64-v8a libsqlcipher.so from sqlcipher-android 4.16.0 (unchanged
since the original encrypted-cache commit, so the crashing build shipped the
same aligned lib) plus libandroidx.graphics.path.so and
libdatastore_shared_counter.so. So the "unaligned .so" theory does not hold;
the nativeOpen UnsatisfiedLinkError needs a different root cause (library-load
ordering / a nativeOpen reached without a loaded lib, or an APK-delivery /
device-specific issue).
2026-07-07 14:52:47 -05:00
Jason Ross 41dca1840d Merge pull request #406 from JMR-dev/ci-404-wedge-diagnostics
ci: capture thread-dump + service state on an E2E wedge to prove the root cause (#404)
2026-07-07 11:55:11 -05:00
JMR-dev 90dfb189e6 fix(ci): drop matrix E2E wedge-capture that hangs all 8 legs
The `timeout -k 30s` wrapper + `capture_wedge()` added in 2f32657 for the
matrix `e2e` job (API 29-36) reproducibly wedges every leg, while the
manually-provisioned API 37 preview shard running the identical capture
logic passes. Revert the two matrix "Run E2E tests" steps' `script:`
blocks to main's plain script (just the backgrounded logcat stream +
`./gradlew connectedDebugAndroidTest`) and drop the now-dead "Upload wedge
diagnostics" step from the matrix job.

Kept untouched: the job-level `timeout-minutes: 50` backstop added in
27ede55, and the entire `e2e-preview` job (its own capture_wedge/watchdog
and wedge-diagnostics-api37-preview-shard* upload are unaffected).
2026-07-07 11:29:41 -05:00
JMR-dev 27ede55dbd ci(e2e): add job-level timeout to matrix E2E job (#404)
The matrix E2E job (api-level 29-36) had no timeout-minutes, so a wedge
hangs until GitHub's 6-hour default instead of being force-killed. The
sibling e2e-preview job already sets timeout-minutes: 35. A normal
matrix run is ~15-20 min and a retry-inclusive run ~40 min, so set
timeout-minutes: 50 to give headroom above the in-step wedge-capture
timeout (1200s) while still bounding worst-case runtime.
2026-07-07 07:46:33 -05:00
JMR-devandClaude Opus 4.8 2f32657aff ci: capture thread-dump + service state on an E2E wedge to prove the root cause (#404)
E2E legs intermittently WEDGE (hang) with no fast-fail until the job force-kill,
and GitHub's post-force-kill step behavior is unreliable, so #388's diagnostics
don't reliably capture the wedge — and don't capture wedge-specific state anyway.

Wrap the `connectedDebugAndroidTest` run (both the `e2e` matrix first-attempt +
retry, and each `e2e-preview` shard) in an explicit `timeout -k 30s 1200`
(20 min) — comfortably above a normal run (~13-15 min), well below the hard cap —
so a wedge trips the wrapper (exit 124), NOT the force-kill, GUARANTEEING the
capture runs while the emulator is still alive. On 124, capture_wedge grabs the
smoking gun into a `wedge-diagnostics-api<level>` artifact: the running/last test
(logcat TestRunner), SIGQUIT (kill -3) thread dumps of the app + instrumentation
processes (ART -> logcat + /data/anr), dumpsys activity/window, `service list` +
`service check input/window/activity` (the boot-race crux), sys.boot_completed +
init.svc.* state, the snapshot cache-hit note, and accel/kvm/mem/disk. Then it
exits with the real status so #388's diagnostics + the existing retry still fire;
a normal run finishes before the wrapper and is unaffected.

EVIDENCE ONLY — no boot-readiness guard/fix (maintainer: prove the cause first).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 07:46:33 -05:00
Jason Ross 5656b3cd99 Merge pull request #414 from JMR-dev/mergify/configuration-deprecated-update
ci(mergify): upgrade configuration to current format
2026-07-07 07:41:55 -05:00
Jason Ross d3264db920 Merge branch 'main' into mergify/configuration-deprecated-update 2026-07-07 07:20:22 -05:00
Jason Ross 404c107aef Merge pull request #419 from JMR-dev/refactor-405-harness-coldfetch
refactor(scripts): fold cold-fetch pause-hook A/B into device-testing harness (#405)
2026-07-07 05:29:48 -05:00
JMR-devandClaude Opus 4.8 55f1f59e3d refactor(scripts): fold cold-fetch pause-hook A/B into device-testing harness (#405)
Bakes the proven 2026-07-06 cold-vs-warm pause-hook flow into
scripts/device-testing/ as a first-class, reproducible `cold-fetch-ab`
scenario, upstreaming the scratchpad driver.

- fetchgate.py: FETCH_GATE pause/resume/query helpers through the guarded
  adb wrapper, with ordered-broadcast read-back parsing (paused=[...]).
- scenarios.cold_fetch_ab: pre-arm halt -> detect sign-in (sync all
  breadcrumb) -> confirm halt (prefetch skipped) -> wait for header sync ->
  measure cold opens -> resume -> measure warm opens. ALWAYS resumes on exit
  (finally), even on error -- never leaves fetch paused.
- report.render_cold_fetch_ab: gate summary, cold/warm tables, cold-vs-warm
  delta, connect=0ms reuse proof, throttle signature.
- Portability (subsumes #392): file-based uiautomator dump (not /dev/tty),
  UTF-8 adb decode + PYTHONUTF8/console I/O, openMessage-breadcrumb readiness,
  row-selection hardening (skip non-message rows).

The pause hook is debug-build-only (#393/#395), so the scenario needs a debug
APK. Automated validation: mocked unittest coverage (adb/breadcrumbs/gate) for
the helpers and the A/B scenario incl. restore-on-error, plus a --dry-run path
exercised end-to-end through perf_harness.main. A full on-device run is a
follow-up. Dev-tooling only; no app/src changes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 05:07:11 -05:00
Jason Ross 582d1f3077 Merge pull request #418 from JMR-dev/test-386-ratchet-floor
test(compose): re-ratchet JaCoCo line-coverage floor to 0.84 (#386)
2026-07-07 03:33:12 -05:00
JMR-dev 4464e3f5e4 test(compose): re-ratchet JaCoCo line-coverage floor to 0.84 (#386)
Final step of the Robolectric Compose epic (#373): now that batches
#376-384 have all landed and proven stable, measure the new whole-app
JVM line-coverage baseline and raise the no-regression floor to match.

Measured 87.89% line (7994/9095), up from 80.21% (4838/6032) when the
floor was last set. Floor moves 0.79 -> 0.84, a deliberately wider
~3.9% headroom (vs. the usual ~0.5-1%) for this first post-epic
measurement; the maintainer can tighten it further in a follow-up PR.
Docs (CLAUDE.md, preflight SKILL.md) updated to match.
2026-07-07 03:14:01 -05:00
Jason Ross 4f09efaf84 Merge pull request #395 from JMR-dev/feat-393-debug-fetch-gate
feat(debug): dev-only pause/halt mail-fetch hook for test harnesses (#393)
2026-07-07 02:46:14 -05:00
Jason Ross 85009ee88a Merge branch 'main' into feat-393-debug-fetch-gate 2026-07-07 02:26:28 -05:00
Jason Ross 2887516faa Merge pull request #417 from JMR-dev/test-384-robolectric-app-shell
test(compose): Robolectric JVM tests — app shell & lock gate host (#384)
2026-07-07 01:38:33 -05:00
JMR-devandClaude Opus 4.8 9f7ebdafb9 test(compose): Robolectric JVM tests — app shell & lock gate host (#384)
Batch 9/9 (final) of the Robolectric Compose JVM-test epic (#373).

- AppLockGateHostJvmTest: drives the app-lock gate host on the JVM via the
  v2 createComposeRule under Robolectric, covering the Unlocked / Checking /
  Locked render branches, the "content stays composed after re-lock" latch,
  and the no-FragmentActivity auth-error path. AppLockViewModel is mocked.
  Drops **/AppLockGateHost* from jacocoNonJvmTestableSurface.
- LibreMailAppJvmTest: covers the JVM-tractable parts of LibreMailApp.kt —
  LibreMailBottomBar, StartupCrashPrompt (+ its dialog buttons), and
  LibreMailApp's cold-start "hold until known" guards.
- LibreMailApp itself KEPT excluded (the acceptable exception noted in #384):
  its NavHost start destinations call hiltViewModel() and the graph needs
  owners a plain JVM compose rule can't surface, so graph-level nav stays on
  the instrumented OnboardingFlowTest. Documented in the jacoco list.

Instrumented LibreMailBottomBarTest / StartupCrashPromptTest stay as the
on-device E2E. JaCoCo floor unchanged (0.79); scoped line coverage 0.8426.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 01:15:07 -05:00
Jason Ross 97a51ed17d Merge pull request #416 from JMR-dev/test-383-robolectric-mailbox
test(compose): Robolectric JVM tests — mailbox screen + folder drawer (#383)
2026-07-07 01:11:42 -05:00
Jason Ross a1cf982a84 Merge branch 'main' into test-383-robolectric-mailbox 2026-07-07 00:53:06 -05:00
Jason Ross 5f69082ba3 Merge pull request #415 from JMR-dev/test-382-robolectric-compose-editor
test(compose): Robolectric JVM tests — compose editor screen (#382)
2026-07-07 00:43:33 -05:00
Jason Ross 4b6f206f2d Merge branch 'main' into test-382-robolectric-compose-editor 2026-07-07 00:23:13 -05:00
Jason Ross 190343bd84 Merge pull request #401 from JMR-dev/test-380-robolectric-settings
test(compose): Robolectric JVM tests for settings screens (#380)
2026-07-07 00:19:42 -05:00
JMR-devandClaude Opus 4.8 cd25544e07 test(compose): Robolectric JVM tests — mailbox screen + folder drawer (#383)
Convert the Paging 3 mailbox list + folder drawer to Robolectric JVM Compose
tests (batch 8/9 of umbrella #373) and drop them from jacocoNonJvmTestableSurface.

- MailboxScreenJvmTest drives the real MailboxScreen + MailboxViewModel over
  mocked repositories, feeding Paging via static PagingData.from flows (no real
  Room/Paging source, mirroring MailboxViewModelTest). Covers the no-accounts
  welcome fallback, populated list (sender/subject/snippet, offline badge,
  unified per-account labels + filter chips, drafts/outbox entries), the
  empty/loading/no-results states, search open/close, and the multi-select
  contextual action bar (overflow, archive/spam/delete confirms, move picker,
  archive-hidden-in-archive, disambiguated app-bar title).
- FolderDrawerJvmTest drives the callback-driven FolderDrawer: friendly role
  names, duplicate-name provider disambiguation + account-switch gap, folder
  taps, the multi-account switcher/dropdown, and the unread badge (incl. 99+ cap).
- Remove **/MailboxScreen* and **/FolderDrawer* from jacocoNonJvmTestableSurface;
  overall JVM line coverage 84.69% (floor unchanged at 0.79).

The instrumented MailboxScreenTest/FolderDrawerTest stay as the on-device E2E.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 00:08:04 -05:00
JMR-devandClaude Opus 4.8 c42f9a7e01 test(compose): Robolectric JVM tests for settings screens (#380)
Convert the settings screens/components to Robolectric JVM Compose tests
(umbrella #373, batch 5/9) and drop their globs from
`jacocoNonJvmTestableSurface`, so they count toward JaCoCo's JVM-testable
surface without an emulator.

New `src/test` Robolectric Compose tests (v2 createComposeRule, @Config sdk=36,
NATIVE graphics), mocking each ViewModel where needed:
- SettingsComponentsJvmTest (SectionHeader/SwitchRow/ClickRow/RadioRow/RetentionSection)
- SettingsScreenJvmTest (+ stateless ContactAutocompleteRow)
- AccountSettingsScreenJvmTest
- SignaturesScreenJvmTest
- SignatureEditScreenJvmTest

Line coverage of the newly-included files: SettingsComponents 100%,
SignatureEditScreen 100%, SignaturesScreen 97%, SettingsScreen 95%,
AccountSettingsScreen 84%. Overall scoped line coverage 86.2%. The instrumented
androidTest E2E stay; the JaCoCo floor is unchanged (re-ratchet is #386).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 23:57:05 -05:00
JMR-devandClaude Opus 4.8 b9cadde5bd test(compose): Robolectric JVM tests — compose editor screen (#382)
Port the instrumented ComposeScreenTest to a Robolectric JVM Compose
test (batch 7/9 of umbrella #373) so ComposeScreen's render + interaction
code counts toward JaCoCo's JVM-testable surface, and drop
`**/ComposeScreen*` from `jacocoNonJvmTestableSurface`.

ComposeViewModel is large (7 collaborators, several Context/Room-backed),
so it is mocked — mirroring AccountPickerScreenJvmTest / ManualSetup
ScreenJvmTest — with its state/accounts/finished flows stubbed so every
render/state branch is injectable. A RESUMED lifecycle owner (also the
back-press dispatcher owner) and a no-op ActivityResultRegistryOwner let
`collectAsStateWithLifecycle`, the BackHandler, and the attachment/inline
-image launchers compose on the JVM. The embedded RichTextBodyField renders
live; its toolbar accessibility labels and body-change plumbing are covered.

The instrumented ComposeScreenTest stays as the on-device E2E. JaCoCo floor
unchanged (0.79); overall line coverage 0.86.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 23:54:29 -05:00
Jason Ross 74a2cde7b1 Merge pull request #400 from JMR-dev/test-381-robolectric-reader
test(compose): Robolectric JVM tests for the reader screen (#381)
2026-07-06 23:51:44 -05:00
JMR-devandClaude Opus 4.8 15f4af864f test(compose): Robolectric JVM tests for the reader screen (#381)
Port the reader screen's chrome to a Robolectric JVM Compose test (umbrella

ReaderScreenJvmTest drives the real ReaderViewModel over a mocked
MailRepository/SettingsRepository via the v2 createComposeRule() — no emulator —
covering the top bar, star/delete/reply/reply-all/forward actions, the
attachment accordion + downloaded indicator, the attachment download-failure
snackbar, and the loading/plain-text/empty/error/remote-images-banner branches.

WebView caveat: the HTML body renders through HtmlBody, a hardened WebView that
Robolectric can only present as a non-rendering shadow, so the banner branch is
driven via an HTML message with a blank body (no HtmlBody call) and no
WebView-rendered HTML is asserted. HtmlBody.kt stays in scope, covered by its
existing HtmlBodyTest/InlineImageResolverTest. The instrumented ReaderScreenTest
stays as the on-device E2E. ReaderScreenKt lands at 94.3% line coverage; the
bundle rises to 82.7%, above the unchanged 0.79 floor.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 23:27:34 -05:00
Jason Ross 9259591dc8 Merge pull request #397 from JMR-dev/test-379-robolectric-mail-lists
test(compose): Robolectric JVM tests — mail list screens drafts/outbox/reports (batch 4/9)
2026-07-06 23:15:26 -05:00
mergify[bot] 33a7ca1b7d ci(mergify): upgrade configuration to current format 2026-07-07 04:05:32 +00:00
Jason Ross 37008d622e Merge branch 'main' into test-379-robolectric-mail-lists 2026-07-06 22:57:06 -05:00
Jason Ross b4453f6991 Merge pull request #412 from JMR-dev/ci-409-mergify-phase1
ci(mergify): implement Phase 1 — serial merge queue (require-up-to-date KEPT ON)
2026-07-06 22:50:20 -05:00
JMR-devandClaude Opus 4.8 40b14d51cc ci(mergify): add Phase 1 serial merge queue (.mergify.yml)
Implements issue #409: a serial Mergify merge queue that supersedes the
hand-rolled poor-man's queue (autoupdate.yml + ci-trigger.yml +
traffic-control.yml, all already disabled_manually).

- queue_rules "default": batch_size 1 (serial, no batching), merge_method
  merge (merge commits, never squash/rebase), merge_conditions gate on
  check-success = "CI passed" + -draft + -conflict + label != broken.
- merge_queue.max_parallel_checks 1 (true serial; unambiguously
  require-up-to-date-compatible).
- priority_rules map P0..P9 labels (P0 = 10000 highest .. P9 = 1000).
- pull_request_rules queue action triggers auto-queueing (queue_conditions
  alone do NOT auto-queue per Mergify lifecycle docs).

require-up-to-date STAYS ON (Phase 1 is the only trilemma combo that keeps
the checkbox literally enabled AND preserves merge commits). No batching
(that is Phase 2 / #410). Single required gate stays "CI passed".

Validated: YAML parses and conforms to Mergify's published JSON schema
(negative-control confirmed). Merging this activates Mergify, so NO
auto-merge — must be reviewed first.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 22:28:14 -05:00
Jason Ross 594c6167d3 Merge pull request #408 from JMR-dev/ci-407-mergify-spec
docs(ci): propose Mergify merge-queue integration spec (#407)
2026-07-06 22:08:34 -05:00
Jason Ross 6a5e86bb11 Merge branch 'main' into ci-407-mergify-spec 2026-07-06 22:08:23 -05:00
Jason Ross e4db3cadf6 Merge branch 'main' into test-379-robolectric-mail-lists 2026-07-06 21:59:00 -05:00
JMR-devandClaude Opus 4.8 0307df88a1 docs(ci): propose Mergify merge-queue integration spec (#407)
Investigate Mergify (free-for-OSS merge queue + batching + speculative checks)
as the right-way replacement for the hand-rolled traffic-controller
(autoupdate.yml + ci-trigger.yml + mothballed traffic-control.yml) and the
manual serial-bump grind, now that GitHub's native merge queue is org-only and
unavailable to a user account.

Proposal only — NO live .mergify.yml, nothing activates:
- docs/ci/mergify-integration-spec.md: how the queue coexists with the single
  `CI passed` gate; the require-up-to-date x merge-commits x batching trilemma
  and its resolution (Phase 1 serial keeps the rule literally; Phase 2 merge-batch
  moves the up-to-date GUARANTEE into the queue); P0-P9 -> priority_rules mapping;
  what it replaces; interaction with path-filter/sharding/wedge-diag; risks;
  phased adopt recommendation.
- docs/ci/mergify.yml.proposed: annotated, NOT-active proposed config.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 21:55:51 -05:00
Jason Ross 41aed01544 Merge pull request #402 from JMR-dev/ci-399-e2e-path-filter
ci: skip the E2E matrix for test-only/docs PRs via a paths-filter + skip-tolerant gate (#399)
2026-07-06 21:54:25 -05:00
JMR-devandClaude Opus 4.8 98b90a19c3 test(compose): Robolectric JVM tests for mail list screens (#379)
Convert the VM-driven mail list screens (DraftsScreen, OutboxScreen,
ProblemReportsScreen) to Robolectric JVM Compose tests in the `test`
source set, driving each real ViewModel over a mocked MailRepository /
ReportStore + DiagnosticsCollector via the v2 createComposeRule() — no
emulator. Each test covers the empty/populated render states, item
rendering (subject/recipient/body, queued-vs-failed status, crash/manual
kind labels), and the interactions (open, delete, cancel, retry, create).

Drop the three now-JVM-covered globs from jacocoNonJvmTestableSurface so
the screens count toward the JaCoCo denominator; measured coverage is
DraftsScreen 100%, OutboxScreen 100%, ProblemReportsScreen 97%, and the
bundle line ratio rises to ~82.7% (floor 0.79 unchanged). The instrumented
androidTest E2Es (DraftsScreenTest / OutboxScreenTest /
ProblemReportsScreenTest) stay as the on-device tests.

Part of the Robolectric Compose umbrella (#373); mirrors the #375/#376
pattern (AddAnotherAccountScreenJvmTest, format-control JVM tests).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 21:40:27 -05:00
Jason Ross 7f1fb3ea43 Merge branch 'main' into ci-399-e2e-path-filter 2026-07-06 21:35:38 -05:00
Jason Ross b8e3b97377 Merge pull request #396 from JMR-dev/test-377-robolectric-onboarding-lock
test(compose): Robolectric JVM tests for onboarding & lock screens (#377)
2026-07-06 21:25:50 -05:00
Jason Ross c4607c2df1 Merge branch 'main' into test-377-robolectric-onboarding-lock 2026-07-06 21:05:23 -05:00
JMR-devandClaude Opus 4.8 e668330151 ci: skip the E2E matrix for test-only/docs PRs via a paths-filter + skip-tolerant gate (#399)
Add a cheap `changes` job (dorny/paths-filter v4, pinned SHA) that sets
e2e_needed=false only when EVERY changed file is in a safe allow-list
(app/src/test/**, **/*.md, docs/**, scripts/**, .claude/**); anything
else -- or any non-pull_request event -- defaults to true (conservative,
"err toward running E2E").

Gate `e2e` and `e2e-preview` on needs.changes.outputs.e2e_needed so the
whole matrix runs or skips together, and rewrite the `ci-passed` gate:
it now BLOCKS on changes!=success, any of traffic-control-tests /
static-analysis / debug-build / unit-tests !=success, or e2e/e2e-preview
==failure|cancelled -- while TOLERATING an intentional e2e/e2e-preview
'skipped'. So test-only/docs PRs go green on the fast gate, a real E2E
failure/cancel still blocks, and a broken filter (changes!=success)
still blocks. Branch protection ("CI passed") context is unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 20:55:58 -05:00
Jason Ross d922b9a6ee Merge pull request #398 from JMR-dev/test-378-robolectric-account-setup
test(compose): Robolectric JVM tests for account setup screens (#378)
2026-07-06 20:54:48 -05:00
JMR-devandClaude Opus 4.8 3ea00a1ed9 test(compose): Robolectric JVM tests for account setup screens (#378)
Add Robolectric JVM Compose tests (umbrella #373, batch 3/9) for the
account-setup screens and drop them from `jacocoNonJvmTestableSurface` so
their render/interaction code counts toward the JVM-testable coverage surface:

- AccountPickerScreen (98.9% line)
- AppPasswordSetupScreen (98.7% line)
- ManualSetupScreen (98.5% line)

Each test drives the real screen via the v2 `createComposeRule()` under
RobolectricTestRunner with a mocked ViewModel (their own logic stays covered by
the ViewModel unit tests), a RESUMED LifecycleOwner for
`collectAsStateWithLifecycle`, a no-op ActivityResultRegistry for the Outlook
launcher, and a recording UriHandler for the app-password help links — covering
render, per-provider chrome, field/submit wiring, and the enabled/busy/error/
done branches. The instrumented androidTest E2Es stay as the on-device coverage.

The JaCoCo floor (0.79) is unchanged — the re-ratchet is the final #373 step.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 20:34:21 -05:00
JMR-devandClaude Opus 4.8 799669d6a6 test(compose): Robolectric JVM tests for onboarding & lock screens (#377)
Add Robolectric JVM Compose tests (umbrella #373, batch 2/9) for the
stateless onboarding + lock screens, and drop each from
jacocoNonJvmTestableSurface so its render/interaction code now counts as
JVM-testable surface:

- LockScreen: locked title/body, optional error, unlock callback.
- WelcomeContent + OnboardingWelcomeScreen: render + add-account; the
  wrapper's NotificationPermissionEffect launcher is wired to a no-op
  ActivityResultRegistry so no system dialog is surfaced on the JVM.
- LicenseScreen: real bundled GPL text renders, Agree gated on
  scroll-to-end, Decline.
- ContactsAccessContent (skip/grant/request/rationale) plus the
  ContactsAccessScreen wrapper, driven by a mocked OnboardingViewModel.
- BatteryOptimizationScreen: offered vs. done states; Take me there marks
  the prompt handled and resolves the settings intent; Not now finishes.

Contacts/Battery use a tall @Config qualifier so their centered,
non-scrolling columns fit without the lower controls clipping. The
instrumented androidTest tests are kept (and remain the coverage for the
system back press, which the JVM compose rule cannot drive). JaCoCo floor
unchanged at 0.79 (the re-ratchet is the final #373 step).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 20:13:49 -05:00
JMR-devandClaude Opus 4.8 8f8608a430 feat(debug): dev-only pause/halt mail-fetch hook for test harnesses (#393)
The on-device perf harness cannot force a genuinely uncached body fetch: proactive
backfill (#12) and post-sync body prefetch (#88/#89) warm the cache before a test can
open a message. Add a debug-only, adb-reachable hook to pause proactive fetch so a real
uncached open can be measured.

Components:
- DebugFetchGate (src/main): thread-safe in-memory holder of paused FetchScopes
  (BACKFILL, PREFETCH; `all` alias). Defaults to not-paused; HEADER_SYNC and on-demand
  OPEN are never gateable.
- FetchGateReceiver (src/debug only): BroadcastReceiver registered in the debug manifest,
  driven by `adb shell am broadcast -a org.libremail.debug.FETCH_GATE -n .../FetchGateReceiver
  --es action <pause|resume|query> --es scope <backfill,prefetch|all>`. Returns the state as
  ordered-broadcast result data (paused=[...]) for a synchronous read-back.

Enforcement (each read guarded by BuildConfig.DEBUG so R8 strips it from release):
- BackfillWorker.doWork() entry -> skip-and-reschedule when BACKFILL is paused, mirroring
  the existing cache-lock deferral (covers periodic + backfillNow()).
- MailSyncer/MailBackfiller.prefetchIfEnabled -> early-return when PREFETCH is paused.
  openMessage / fetchBodyMarkingSeen / fetchAttachment are deliberately NOT gated.

Debug-only: receiver + <receiver> live wholly in src/debug; every gate read in main is
behind BuildConfig.DEBUG. Verified on assembleRelease that R8 strips DebugFetchGate /
FetchScope / FetchGateReceiver and the log strings from the release APK, and the merged
release manifest has no FETCH_GATE receiver.

PII-free AppLog breadcrumbs on pause/resume/query and on each gate-triggered defer/skip
(scope names only).

Tests: DebugFetchGateTest, BackfillWorkerTest / MailSyncerTest / MailBackfillerTest
enforcement cases, and FetchGateReceiverInstrumentedTest (ordered-broadcast -> gate ->
read-back; gated worker defers while an un-gated path runs).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 20:13:22 -05:00
Jason Ross ae10a5ff3b Merge pull request #394 from JMR-dev/test-376-robolectric-richtext
test(compose): Robolectric JVM tests for rich-text format controls (#376)
2026-07-06 20:09:40 -05:00
Jason Ross c2ab0e3141 Merge branch 'main' into test-376-robolectric-richtext 2026-07-06 19:54:03 -05:00
JMR-devandClaude Opus 4.8 9bbfa2108a test(compose): Robolectric JVM tests for rich-text format controls (#376)
Port the instrumented ColorSwatchRow / FontPicker / FontSizePicker /
ParagraphAlignmentControl tests to Robolectric JVM Compose tests (v2
createComposeRule, @GraphicsMode NATIVE, @Config sdk=36) in the `test`
source set, and drop their four globs from `jacocoNonJvmTestableSurface`
so they count toward the JVM coverage metric. The instrumented tests stay.

Also fix a latent gap in the #375 infra: the JaCoCo agent skips classes
with no code-source location, which is exactly how Robolectric loads the
classes-under-test through its sandbox classloader — so Robolectric-only
Compose coverage recorded as zero (the PoC AddAnotherAccountScreen
included). `isIncludeNoLocationClasses = true` on the Test tasks makes
that coverage register; scoped bundle line coverage rises ~0.80 -> ~0.82.
Floor left at 0.79 (#386 re-ratchets).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 19:49:59 -05:00
Jason Ross cfd433af6f Merge pull request #367 from JMR-dev/fix-359-sqlcipher-16kb
fix(security): fail closed on cache-encryption load failure + StrongBox-back keys (#359)
2026-07-06 19:41:35 -05:00
Jason Ross 8689885964 Merge branch 'main' into fix-359-sqlcipher-16kb 2026-07-06 19:23:24 -05:00
Jason Ross 6213ee8901 Merge pull request #391 from JMR-dev/ci-389-sdk-setup-hardening
ci: retry + cache Android SDK/emulator setup to survive corrupt-zip sdkmanager failures (#389)
2026-07-06 19:22:02 -05:00
Jason Ross 8ac1d28dc8 Merge branch 'main' into ci-389-sdk-setup-hardening 2026-07-06 19:03:26 -05:00
JMR-devandClaude Opus 4.8 aa628831be ci: retry + cache Android SDK/emulator setup to survive corrupt-zip sdkmanager failures (#389)
The dominant merge-blocking flake was the "Set up Android SDK" step
(android-actions/setup-android v4.0.1) dying BEFORE the emulator starts:

    Wrong version in preinstalled sdkmanager
    Warning: ... preparing SDK package Android Emulator: Error reading Zip
    content from a SeekableByteChannel.
    Error: The process '.../sdkmanager' failed with exit code 1

Root cause: the action's default cmdline-tools version (20.0) rarely matches the
runner image's preinstalled one, so it logs "Wrong version in preinstalled
sdkmanager" and re-fetches cmdline-tools with NO checksum; it then runs its
default `sdkmanager tools platform-tools` install. Any of those downloads can be
a corrupt/truncated zip, which sdkmanager turns into an un-retried exit 1. v4.0.1
is the latest release, so this is fixed by configuration + hardening, not a bump.

Harden with verify -> reject -> retry, never trusting sdkmanager's exit code
alone, via a new stdlib-only helper .github/scripts/setup_android_sdk.py:

- bootstrap: download the pinned cmdline-tools zip, verify size + SHA-256
  (authoritative pin, cross-checked against Google's published SHA-1), and
  install it to $ANDROID_SDK_ROOT/cmdline-tools/20.0 -- the exact path
  setup-android probes first, so the action reuses the verified tree and never
  does its own unverified "Wrong version" re-download. A mismatch (corrupt OR
  wrong version) deletes the bad zip + any half-extracted dir and re-downloads.
- install: sdkmanager --install with retry + backoff; on a corrupt package zip it
  purges the partial/corrupt package dir (and sdkmanager's temp dirs) before
  retrying, forcing a fresh download instead of a re-read.
- setup-android now runs with packages: "" (no flaky tools/platform-tools
  install) and cmdline-tools-version: "14742923" (reuse the verified bootstrap).
- actions/cache restore + success-gated save so only a verified SDK is ever
  cached (integrity gates the cache); shrinks the re-download/corruption surface.

Applied to every SDK-setup job (debug-build, unit-tests, static-analysis, e2e
matrix, e2e-preview). Emulator BOOT logic, #372 API-37 sharding, and #388
diagnostics are untouched. Pure-logic helpers are unit-tested
(test_setup_android_sdk.py, run by the traffic-control-tests job).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 18:59:05 -05:00
Jason Ross 1053fa80f0 Merge pull request #375 from JMR-dev/robolectric-compose-scope
test(compose): Robolectric JVM Compose testing infra + PoC (#373)
2026-07-06 18:36:36 -05:00
Jason Ross 35b6b35d16 Merge branch 'main' into robolectric-compose-scope 2026-07-06 17:44:37 -05:00
Jason Ross 5a8c4e386e Merge branch 'main' into fix-359-sqlcipher-16kb 2026-07-06 16:56:27 -05:00
Jason Ross c04198a158 Merge pull request #388 from JMR-dev/ci-387-e2e-diagnostics
ci(e2e): capture logcat + emulator/system diagnostics across the E2E matrix (#387)
2026-07-06 16:55:52 -05:00
Jason Ross 0b1fb05a90 Merge branch 'main' into ci-387-e2e-diagnostics 2026-07-06 16:42:18 -05:00
Jason Ross e062331e09 Merge pull request #374 from JMR-dev/fix-signatures-test-teardown-race
test(settings): cancel viewModelScope before db.close in SignaturesScreenTest to fix a Room teardown race
2026-07-06 16:21:56 -05:00
JMR-devandClaude Opus 4.8 187a8effb0 ci(e2e): capture logcat + emulator/system diagnostics across the E2E matrix (#387)
The API 29-36 `e2e` matrix uploaded only its test report, so an emulator
flake or a red leg (e.g. `E2E (31)` dying on a bare `sdkmanager` exit 1)
left nothing to diagnose. Bring the #334 API-37 diagnostics to the matrix,
inline (no changes to `e2e-preview`, which PR #372 is restructuring):

- Stream `adb logcat -v time` to `$RUNNER_TEMP/logcat-api<level>.txt` at the
  top of both the "Run E2E tests" and retry reactivecircus steps (emulator is
  booted there); backgrounded so gradle stays the exit-status-bearing command.
- New `if: failure()` step dumps device + runner state (adb devices, logcat
  tail, emulator -accel-check, /dev/kvm, free -h, df -h) to the step log and a
  diagnostics file; every probe guarded with `|| true`.
- New `if: always()` upload-artifact (same pinned v7 SHA) `e2e-diagnostics-api<level>`
  carries the logcat + diagnostics files, `if-no-files-found: warn`.
- Make "Install SDK platform and build-tools" diagnosable: bounded 3x retry with
  backoff for a transient sdkmanager failure, and print `--list_installed` on a
  hard failure instead of a bare exit 1.

Keeps reactivecircus/android-emulator-runner and the existing boot-race retry.
Additive/diagnostic only; no boot-affecting flags change.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 15:53:54 -05:00
Jason Ross 0186fa5ed9 Merge branch 'main' into robolectric-compose-scope 2026-07-06 15:53:48 -05:00
JMR-devandClaude Opus 4.8 324f7c2c51 fix(test): resolve Robolectric android-all via Gradle so the JVM Compose test runs in CI (#373)
Robolectric resolved its android-all-instrumented runtime jar lazily at test
time via its own MavenDependencyResolver/MavenArtifactFetcher, and that
download is unreliable on CI runners: AddAnotherAccountScreenJvmTest failed
with `AssertionError at MavenArtifactFetcher ... IOException` ("Failed to
fetch maven artifact"), though it passed locally where ~/.m2 was warm.

Resolve the jar through Gradle instead (reliable, cached, persisted by the CI
Gradle cache) and hand it to Robolectric in offline mode so it never hits the
network at test time:
- Pin org.robolectric:android-all-instrumented:16-robolectric-13921718-i7
  (exactly what Robolectric 4.16.1 DefaultSdkProvider maps @Config(sdk=36) to)
  in the version catalog.
- Add it to a dedicated resolvable configuration (NOT testImplementation/
  testRuntimeOnly, which would flatten the ~200MB instrumented framework onto
  the JVM test classpath and collide with the stub android.jar).
- syncRobolectricAndroidAll stages the jar under its Maven filename, and
  robolectric.offline + robolectric.dependency.dir point Robolectric's
  LocalDependencyResolver at it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 15:49:07 -05:00
Jason Ross 62797bc8c3 Merge branch 'main' into fix-359-sqlcipher-16kb 2026-07-06 15:46:14 -05:00
Jason Ross c8290ee545 Merge branch 'main' into fix-signatures-test-teardown-race 2026-07-06 15:45:09 -05:00
Jason Ross 9aa83a458c Merge pull request #372 from JMR-dev/ci-api37-e2e-sharding-spike
ci: shard the API 37 preview E2E into 2 parallel shards + retry parity
2026-07-06 15:41:02 -05:00
JMR-devandClaude Opus 4.8 cf1a8f6b83 test(compose): Robolectric JVM Compose testing infra + PoC (#373)
Enables unit-testing Jetpack Compose UI on the JVM via Robolectric, so
render-only screens can leave the jacocoNonJvmTestableSurface exclusion
list and be counted by JaCoCo without an emulator.

- add Robolectric 4.16.1 (test scope) + Compose ui-test-junit4/-manifest
- testOptions.unitTests.isIncludeAndroidResources = true so resources
  (strings, Material3 theme) resolve on the JVM
- src/test/resources/robolectric.properties pins sdk=36 (targetSdk 37 is
  a preview level Robolectric 4.16 has no sandbox for)
- PoC: AddAnotherAccountScreenJvmTest drives the screen with the v2
  createComposeRule under RobolectricTestRunner (3 tests, green on the JVM)
- drop AddAnotherAccountScreen from jacocoNonJvmTestableSurface (now
  JVM-covered); floor stays 0.79 — re-ratchet deferred to end of #373

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 15:37:23 -05:00
Jason Ross bb412e263e Merge branch 'main' into fix-signatures-test-teardown-race 2026-07-06 15:29:59 -05:00
Jason Ross b02ed87313 Merge branch 'main' into ci-api37-e2e-sharding-spike 2026-07-06 15:23:07 -05:00
Jason Ross 32199ef037 Merge branch 'main' into fix-359-sqlcipher-16kb 2026-07-06 15:22:45 -05:00
Jason Ross 1423395454 Merge pull request #370 from JMR-dev/feat-device-testing-harness
feat(scripts): device-testing perf harness
2026-07-06 15:22:20 -05:00
Jason Ross 1a3393da53 Merge branch 'main' into ci-api37-e2e-sharding-spike 2026-07-06 15:12:47 -05:00
JMR-devandClaude Opus 4.8 08ee9e7abc ci: productionize API 37 preview E2E sharding (adopt N=2)
The spike commits already implemented the N=2 shard matrix (numShards/
shardIndex via -Pandroid.testInstrumentationRunnerArguments.*), per-shard
test-retry parity, adb start-server before the boot loop, and shard-suffixed
artifact names. This drops the SPIKE / DRAFT "do not merge as-is" framing from
the ci.yml comments and reframes docs/perf/api37-e2e-sharding-spike.md from a
feasibility spike into the adopted design, so the change is mergeable as-is.

Also fixes the doc's section 3a example, which showed 1-based shardIndex values
[1, 2]; shardIndex is 0-based (0..numShards-1) and the implementation correctly
uses matrix.shard: [0, 1] -- [1, 2] would run an empty bucket and silently drop
half the suite.

Fan-in unchanged and verified: ci-passed still lists e2e-preview once; GHA
matrix aggregation makes its result `failure` if either shard fails, so both
shards must pass for the gate to go green. Branch protection requires the
"CI passed" context (not the per-leg "E2E (API 37 preview) (N)" check names),
so no branch-protection change is needed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 14:52:48 -05:00
Jason Ross 5f77eb7d8b Merge branch 'main' into fix-359-sqlcipher-16kb 2026-07-06 14:51:12 -05:00
Jason Ross d6c694c7e0 Merge branch 'main' into feat-device-testing-harness 2026-07-06 14:47:16 -05:00
JMR-devandClaude Opus 4.8 ca6d90b602 test(settings): cancel viewModelScope before db.close in SignaturesScreenTest to fix a Room teardown race (flaky on API-37 CI)
SignaturesScreenTest built a real SignaturesViewModel by hand but tore down
with a bare `db.close()` that never cancelled viewModelScope. The ViewModel's
`signatures` StateFlow is a Room InvalidationTracker Flow kept alive by
stateIn(WhileSubscribed(5_000)), so the collector could stay live up to 5s
after the UI detached — a re-query then landed on the just-closed in-memory DB
and threw SQLITE_MISUSE ("connection is closed"). Timing-dependent, hence the
intermittent API-37 CI failure in tappingRadioOnNonDefault_makesItTheDefault.

Fix: hold the ViewModel in an androidx.lifecycle.ViewModelStore and, in @After,
call store.clear() (→ ViewModel.onCleared() → cancels viewModelScope) BEFORE
db.close(), so the collector is gone before the DB closes. Behaviour and
assertions are unchanged; the fix removes the race by construction.

Audited the androidTest tree for the same hazard and fixed two siblings the
same way:
- AccountSettingsScreenTest: had the same live-Room-Flow-vs-close race,
  previously worked around by never closing the in-memory DB at all; now
  clears the ViewModel then closes the DB.
- ComposeScreenTest: ComposeViewModel's init launches a viewModelScope
  coroutine that reads the real accountSettings/signature Room repos; clear
  the store before db.close() to avoid the same in-flight-read-vs-close race.

Verified locally: connectedDebugAndroidTest green for all three classes
(12/12) on a cold-booted emulator, plus the JVM fast gate (assembleDebug,
testDebugUnitTest, jacocoTestCoverageVerification, compileDebugAndroidTestKotlin,
lintDebug, ktlintCheck, detekt).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 14:46:40 -05:00
JMR-devandClaude Opus 4.8 0e9f54e686 ci(spike): retry parity + adb start-server for API 37 shard PoC
Folds the #370 root-cause finding into the spike. The stable `e2e` matrix
retries its test run once; `e2e-preview` runs connectedDebugAndroidTest exactly
once, so a flaky test self-heals on API 29-36 but wedges the required gate on
API 37 (e.g. #370's SignaturesScreenTest teardown race).

Doc: adds risk item 9 (retry-parity gap + its sharding interaction — per-test
flake is NOT amplified by sharding unlike boot flake, and a per-shard retry
costs only B + T/N; framed mitigation-not-fix) and two §6 recommendations
(retry parity, mirrored into api37_e2e.py; adb start-server before the boot
loop).

PoC (ci.yml): per-shard single test retry (::warning:: on retried-but-passed)
+ adb start-server before the boot loop. The api37_e2e.py retry mirror stays a
documented recommendation (local path needs a real-emulator validation this
spike did not boot). Still DRAFT, not auto-merged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 14:40:02 -05:00
Jason Ross a84b042af5 Merge pull request #371 from JMR-dev/chore-preflight-coverage-gate
chore(preflight): run jacocoTestCoverageVerification in the fast gate
2026-07-06 14:36:55 -05:00
JMR-devandClaude Opus 4.8 66da643249 ci(spike): PoC shard API 37 preview E2E + feasibility doc
Feasibility spike for sharding the e2e-preview job (the hand-provisioned
API 37 / google_apis_ps16k 16 KB-page emulator), CI's longest leg
(~16.4-17.6 min). docs/perf/api37-e2e-sharding-spike.md breaks the leg into
fixed overhead B ~8.3 min (setup + boot + Gradle daemon/config/compile/install)
vs parallelizable test execution T ~8.8 min, models B + T/N for N=2/3/4, and
recommends N=2 (~17.1 -> ~12.7 min, ~28% off the critical path) capped by the
API 30 matrix wall (~12.0 min) beyond N=3.

DRAFT PoC (do NOT merge as-is): converts e2e-preview to a strategy.matrix.shard
[0, 1] fan-out passing AndroidJUnitRunner numShards/shardIndex through the
existing -Pandroid.testInstrumentationRunnerArguments.* channel (no GMD, no
orchestrator, no Gradle change). Artifact names gain a shard suffix;
ci-passed still lists e2e-preview once (matrix fan-in keeps the single gate).
Local preflight stays single-emulator. Relates to #258.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 14:31:40 -05:00
JMR-devandClaude Opus 4.8 042b50116c build(jacoco): scope the fail-closed encryption UI out of the JVM coverage surface (#359)
CacheEncryptionGate.kt (the gate composable, blank cover, error screen, and ephemeral
report-review screen added for #359) is pure Compose render code, structurally
unreachable from a JVM unit test the same way every other Screen file in
jacocoNonJvmTestableSurface is. Left in scope, it dragged the whole-app line ratio to
0.78, just under the 0.79 no-regression floor.

Excluded it via "**/CacheEncryptionGateKt*" rather than the usual bare
"**/CacheEncryptionGate*" pattern this list otherwise uses, because
CacheEncryptionGateViewModel is named with "CacheEncryptionGate" as a literal
prefix - the bare wildcard would also have swallowed the already JVM-tested,
94%-covered ViewModel and its sealed CacheEncryptionGateState. CacheEncryptionGateViewModel
and CacheEncryptionUnavailableException stay in scope unchanged.

Verified locally: testDebugUnitTest + jacocoTestCoverageVerification now pass, with
the line ratio recovered to about 0.807 (5,044 covered / 6,249 total lines) - the
same 5,044 covered lines as before, just a smaller, honestly-JVM-testable denominator.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 14:19:22 -05:00
JMR-devandClaude Opus 4.8 b0ca5421b6 chore(preflight): run jacocoTestCoverageVerification in the fast gate
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 14:18:40 -05:00
JMR-devandClaude Opus 4.8 bfe5d46654 fix(security): fail closed on cache-encryption load failure with an error gate (#359)
Reworks #367. When SQLCipher native library loading fails while the opt-in
encrypted cache is enabled, the app previously degraded to a plaintext cache
(a silent fail-open that defeats the feature). Now it FAILS CLOSED.

DatabaseProvisioner raises a distinct CacheEncryptionUnavailableException
instead of degrading: it does NOT open plaintext, NOT wipe the on-disk
ciphertext, and NOT write the encryptCache setting. The throw is not
memoized, so a later launch re-attempts and recovers automatically if the
library loads.

A new CacheEncryptionGate wraps the app inside AppLockGateHost (so the
passphrase is already unlocked), probes prepareCache() before any DB-backed
screen composes, and on failure shows CacheEncryptionErrorScreen with the
exact message "Error - decryption could not proceed. Native decryption
library load failure." plus a "Report a problem" action. That action
generates an EPHEMERAL PII-free report via the existing DiagnosticsCollector
(never written to ReportStore, since encryption is unavailable in that
moment) for on-screen review and explicit Copy/Save; the copy says so.

The plaintext AccountDatabase tolerates the exception so accounts stay
readable for the error gate and the report. The encryptCache setting is now
written by exactly one caller: the user Settings toggle.

Tests: fail-closed raises the signal with no plaintext open / no wipe / no
setting write / not memoized; the gate VM resolves Ready vs Unavailable and
builds the ephemeral report; an instrumented error-screen UI test and an
AccountDatabase-resilience instrumented test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 13:48:35 -05:00
JMR-devandClaude Opus 4.8 551a2df66f feat(security): back cache-key Keystore keys with StrongBox, fall back to TEE (#359)
Bind the non-exportable AES-256-GCM keys that seal the SQLCipher cache
passphrase to the hardware StrongBox secure element when the device has
one. Applied in the single shared place, AesGcmKeystoreCipher, so it
covers both the master (KeystoreCrypto) and auth-bound (DatabaseKeyCipher)
keys.

Devices without StrongBox throw StrongBoxUnavailableException at
KeyGenerator.generateKey(); a new generate-with-fallback path catches it
and regenerates a TEE-backed key so key creation still succeeds
everywhere. Guarded on API 28+ (minSdk is 29). Framing, seal/unseal, and
the missing-key policies are unchanged; the passphrase is still never
plaintext at rest and never logged.

Adds a JVM regression test for the StrongBox->TEE fallback via the
existing test seams (existingKey / a new generateKey seam).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 13:46:03 -05:00
JMR-devandClaude Opus 4.8 e436503eaf feat(scripts): device-testing perf harness
Add a cross-platform, standard-library-only Python package under
scripts/device-testing/ that replicates LibreMail's on-device performance-test
scenarios and logging capture, codifying the methodology run by hand on
2026-07-05 (Pixel 10 Pro XL).

Modules:
- breadcrumbs.py: a pure, unit-tested parser for the ImapPerf / MailReader /
  Reader / MailBackfiller breadcrumbs, plus open-correlation that reproduces the
  manual timing-tables.md figures exactly.
- adb.py: a safety-guarded adb wrapper -- an allow-list of adb subcommands and a
  deny-list + assertions on shell commands. The only sanctioned app-state
  mutation is clearing LibreMail's own cache/ (exact-match); no pm clear /
  uninstall / data wipe, and no touching databases/ files/ shared_prefs/
  datastore/ can be constructed.
- uidump.py: uiautomator XML parser + screen recognition (mailbox rows with the
  cached "Available offline" flag, reader, and the keyguard / foreign-app guards).
- scenarios.py: cold-open, message-open (uncached), back-nav, prefetch A/B
  (fetch-policy toggle) and cross-provider, each keyguard-guarded and driven
  through the guarded wrapper.
- report.py + perf_harness.py: aggregates, a timing-tables.md renderer mirroring
  the manual write-up, and the CLI (timestamped run dir with the raw logcat, a
  filtered breadcrumb extract, and the tables). --dry-run prints the exact
  command plan without touching device state.

Tests (stdlib unittest, 68 cases) validate the parser, guardrails, UI
recognition and report against the manual run's real captures (fixtures include
a verbatim perf-extract slice and the reader/lockscreen/alarm dumps). Track the
*.log fixture past the gitignore *.log rule via a scoped negation.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 13:21:34 -05:00
Jason Ross 5c715e1676 Merge branch 'main' into fix-359-sqlcipher-16kb 2026-07-06 12:50:29 -05:00
Jason Ross b8d67557a0 Merge pull request #368 from JMR-dev/feat-125-imap-connection-reuse
perf(mail): enable IMAP connection reuse by default with a hardened cache
2026-07-05 22:18:11 -05:00
JMR-dev e36afc8ade changed conditional style to easier to read/maintain when (like switch) statement 2026-07-05 21:18:49 -05:00
JMR-devandClaude Opus 4.8 81a3b7ea34 refactor(mail): early-return guard in ImapConnectionCache (#357 review)
Restructure isConnectionDrop as leading guard clauses (definite-drop
types, then a not-MessagingException early return) instead of a when
expression, per maintainer review feedback on PR #368. Behavior is
unchanged; verified by the existing ImapConnectionCacheTest suite
(all 8 cases still pass), including the FolderClosedException /
StoreClosedException cases that depend on the check running before
the MessagingException .cause guard.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 20:36:38 -05:00
JMR-devandClaude Opus 4.8 060b7b1a71 Merge origin/main into feat-125-imap-connection-reuse
Resolve the IdleService.kt conflict as a union of both intents:
- #354 (already on main): foreground-service lifecycle rework —
  onStartCommand delegates to the IdleForegroundStarter seam
  (START_NOT_STICKY), cap-window skip/degrade.
- #357 Part 2 / #368: reused-connection idle-eviction sweep and
  low-battery teardown of reused connections.

In startWatchingIfNeeded(), reconcileWatchers() stays inside the
cache-lock-guarded launch and evictIdleReuseConnectionsLoop() launches as
a sibling coroutine that runs while the service lives (its original #368
placement, independent of the cache-lock guard). No behavior change to
either side.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 20:15:53 -05:00
JMR-devandClaude Opus 4.8 cc067408b8 perf(mail): enable IMAP connection reuse by default with a hardened cache
An on-device drilldown proved Gmail server-side throttles LibreMail's
connect-per-operation IMAP: every op was a fresh CONNECT+TLS+LOGIN, and
full-history backfill's body+attachment prefetch generated ~601 connections in
~22 min, tripping (and sustaining) Gmail's per-account rate/bandwidth clamp
(body download collapsed to ~4 KB/s). The `live` gauge peaked at only 5 (Gmail
allows ~15), so it is connection *volume*, not count. Outlook IMAP on the same
device opened in 2-3 s. Reusing one warm socket per account (~601 -> ~1) removes
the throttle's trigger. This wires the reuse path the #125 spike built and left
OFF (issue #357 Part 2 — connection reuse only; prefetch is a separate PR).

How it is enabled (with a safety switch):
- New `BuildConfig.IMAP_CONNECTION_REUSE` (default true) drives the production
  `ImapClient` no-arg `@Inject` constructor. To disable if a server misbehaves,
  flip it to "false" in app/build.gradle.kts — a build-config change, no Kotlin
  edit. The internal `ImapClient(reuseConnections, reuseIdleTimeoutMillis)`
  constructor stays the test/harness seam.
- Universal: applies to all providers (incl. Outlook). No per-provider caps or
  throttling here — that is a separate effort (#356/#360-#364).

Hardening `ImapConnectionCache` for production (was a spike):
- Transparent stale recovery: broadened drop detection to Angus's own
  `iap.ConnectionException` (and a MessagingException caused by one) — the real
  signal `folder.open()` throws on a server-dropped idle socket, which the
  IOException-only check missed, so the reconnect now actually fires. A dropped
  reused socket is rebuilt once and the op retried, so callers see no spurious
  error; a genuine app error (e.g. message-not-found) is never retried.
- Idle eviction: `evictIdle()` closes a connection unused past the reuse idle
  timeout (default 5 min), swept every 2 min by `IdleService`; skips any
  in-use connection.
- Teardown: `IdleService` also tears down reused connections on the low-battery
  push-teardown path (#88/#89/#90), mirroring the IDLE connection teardown.
- Concurrency: one connection per account behind a per-account mutex; the
  eviction sweep takes the lock non-blockingly so it never stalls or interrupts
  an in-flight op. Coexists with IMAP IDLE (its own separate connection).
- PII-free AppLog on the lifecycle (open / reuse-hit / reconnect-stale / evict /
  teardown) keyed by an opaque per-cache ordinal, plus the #358 ImapPerf
  breadcrumb (connect~=0ms on a reuse hit).

Tests (all via the fast gate, no emulator):
- ImapConnectionCacheTest: reuse, retry-once stale recovery, narrow drop
  detection, deterministic idle eviction (injected clock), teardown.
- ImapFolderOpenLatencyTest (GreenMail + counting proxy): N ops share one
  connection/LOGIN; a force-dropped socket is transparently reconnected; an app
  error does not reconnect; idle eviction LOGS-OUT and the next op reconnects.
- Correctness suites (ImapClientTest/ImapClientBackfillTest/MailBackfillerTest)
  pinned to reuse-off to keep their connect-per-op assertions unchanged.

Fast gate green: assembleDebug, testDebugUnitTest, compileDebugAndroidTestKotlin,
lintDebug, ktlintCheck, detekt.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 19:47:01 -05:00
Jason Ross a069188b65 Merge pull request #366 from JMR-dev/fix-354-idleservice-fgs
fix(push): stop IdleService dataSync FGS crash-loop on exhausted 24h cap (#354)
2026-07-05 19:36:27 -05:00
JMR-devandClaude Opus 4.8 921681812c fix(data): degrade encrypted cache to plaintext on SQLCipher native-load failure
On Android 15+ / SDK 37 devices with 16 KB memory pages (e.g. Pixel 10 Pro XL,
and the API-37 `google_apis_ps16k` emulator image), a native `.so` not aligned
for 16 KB pages fails to load with `UnsatisfiedLinkError` at
`SQLiteConnection.nativeOpen`. With the opt-in SQLCipher encrypted cache on, this
crashed the app on every cold start (issue #359, x4 on-device) instead of
degrading, and encryption silently never applied.

Fix: DatabaseProvisioner's encryption gate now catches `LinkageError`
(UnsatisfiedLinkError and related native-link failures) when opening/converting
the encrypted cache and degrades cleanly instead of propagating the crash — it
turns `encryptCache` off (so the next start does not re-attempt and re-wipe),
clears any on-disk ciphertext the plaintext framework opener cannot parse
(resetting its now-useless seals), and opens the cache unencrypted. The cache is
a re-syncable copy of server mail, so clearing it loses nothing that cannot be
re-fetched. PII-free AppLog.w breadcrumb on the degrade path.

Dependency: no bump needed or available. The repo already pins the newest
SQLCipher it references, `net.zetetic:sqlcipher-android:4.16.0`, which
docs/play-compliance.md certifies (ELF p_align = 0x4000) as 16 KB-aligned on
every ABI; SQLCipher has shipped 16 KB-aligned binaries since well before it, and
the other two bundled `.so` files (Compose graphics-path, DataStore
shared-counter) are already 16 KB-aligned per that doc. The graceful-degrade
catch is therefore the actionable fix.

Tests:
- Unit (DatabaseProvisionerTest): a simulated native-load failure degrades to a
  plaintext open without crashing, turns encryptCache off, and wipes + reseals an
  already-encrypted cache.
- Instrumented (DatabaseProvisionerInstrumentedTest): a fresh encrypt-on start
  loads the real SQLCipher native library and opens the keyed cache — CI's API-37
  `google_apis_ps16k` 16 KB job exercises the actual `.so` load, catching any
  future 16 KB-alignment regression.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 19:26:05 -05:00
JMR-devandClaude Opus 4.8 52503c000c fix(push): stop IdleService dataSync FGS crash-loop on exhausted 24h cap
Root cause: after #302's runtime-cap fallBackToPeriodicSync() stops the
dataSync foreground service, IdleService was restarted (START_STICKY
null-intent redelivery + explicit startForegroundService) and onStartCommand
unconditionally called startForeground(DATA_SYNC) while the rolling-24h budget
was still exhausted. The platform rejected the start with
ForegroundServiceStartNotAllowedException; it was uncaught, the process
crashed, and START_STICKY restarted straight back into the same rejection -- a
crash loop until the 24h window freed budget (#354).

Fix (IdleService.kt):
- onStartCommand now returns START_NOT_STICKY. Push is app-managed
  (LibreMailApplication.ensurePushStarted deterministically restarts it), so the
  sticky null-intent auto-restart was redundant and fired exactly when a dataSync
  FGS start is illegal.
- Guard the foreground start via a new JVM-testable IdleForegroundStarter seam:
  a ForegroundServiceStartNotAllowedException (caught via its IllegalStateException
  supertype, so no minSdk-29 class load) degrades like the cap handler --
  schedulePeriodicSync(), keep the degraded POLLING notification, stopSelf()
  promptly (avoids the "did not call startForeground in time" ANR) -- instead of
  propagating.
- Record the cap event (elapsedRealtime); while still inside the cap window,
  onStartCommand skips the now-guaranteed-illegal foreground start entirely.
- onTimeout stop path kept fast so ForegroundServiceDidNotStopInTimeException
  stays mitigated.

PII-free AppLog.w/i on the degrade paths.

Tests:
- Unit (IdleForegroundStarterTest): onStartCommand returns START_NOT_STICKY; a
  rejected start is caught and routed to degrade without propagating; the cap
  window skips the attempt; a non-ISE propagates.
- Instrumented (IdleServiceForegroundStartInstrumentedTest): the degrade path on
  a real Context -- rejection caught, periodic-sync fallback scheduled, degraded
  "instant delivery paused" notification built, watching skipped.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 19:17:19 -05:00
Jason Ross 6cb179bd8b Merge pull request #365 from JMR-dev/feat-358-reader-perf-logging
feat(reporting): PII-free latency breadcrumbs on the message-open path (#358)
2026-07-05 18:54:32 -05:00
JMR-devandClaude Opus 4.8 35b869944e feat(reporting): PII-free latency breadcrumbs on the message-open path (#358)
Adds AppLog breadcrumbs to the message-open path so a debug report can show
where the reader's spinner time goes:

- ImapClient.withStore: per-op connect vs. work timing plus a live
  connect-per-op connection gauge (issue #125's provider-ceiling context).
- fetchBodyMarkingSeen: select/body/flag phase timings plus PII-free size
  counts (RFC822 size, body chars, attachment count).
- MailRepositoryImpl.openMessage: end-to-end open latency plus the
  cached-vs-fetched branch, keyed by accountLogRef and logSafeFolderLabel.
- ReaderViewModel: spinner-to-ready latency, split success vs. failure.

All breadcrumbs are PII-free: accounts are logged via the existing
accountLogRef hash, folders via the existing logSafeFolderLabel allowlist,
and everything else is sizes/durations/booleans only.

Fixes the 4 unit-test classes that exercise this code without mocking
android.util.Log (a throwing stub under plain JVM tests): mockkStatic(Log)
is now installed in MailRepositoryImplCoverageTest, ImapClientBackfillTest,
ImapFolderOpenLatencyTest, and ReaderViewModelActionsTest, following the
existing MailBackfillerTest/ImapClientTest conventions. detekt.yml gains two
more ForbiddenImport excludes for the newly Log-importing test files.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 17:46:48 -05:00
Jason Ross f93e2dc2f0 Merge pull request #353 from JMR-dev/ci-extract-traffic-control
ci: extract traffic-control into its own workflow file
2026-07-05 16:17:10 -05:00
JMR-devandClaude Opus 4.8 939906986b ci: extract traffic-control into its own workflow file
Move the traffic-control (runner-priority orchestration) job verbatim out of
.github/workflows/ci.yml into a new standalone workflow,
.github/workflows/traffic-control.yml, so the heavy CI jobs no longer depend
on it. The job's YAML (name, runs-on, timeout-minutes, permissions, env,
steps) and its documentation comment move unchanged; the decision core
.github/scripts/traffic_control.py is untouched and still unit-tested by the
traffic-control-tests job in ci.yml.

In ci.yml: removed the traffic-control job, dropped needs: traffic-control
from the five heavy jobs (static-analysis, debug-build, unit-tests, e2e,
e2e-preview) and from traffic-control-tests (its only needs, which would
otherwise dangle at a now-deleted job), and updated the now-stale header and
ci-passed comments to point at the extracted workflow.

The new workflow will be disabled pending a rebuild as a published GitHub
Action.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 14:48:03 -05:00
Jason Ross 6118b6ded0 Merge pull request #285 from JMR-dev/perf-187-covering-index
perf(db): covering index for unified-inbox summary scans
2026-07-05 14:31:22 -05:00
github-actions[bot] 1cd0d66fa4 Merge main into perf-187-covering-index 2026-07-05 18:48:15 +00:00
Jason Ross c3933a3612 Merge pull request #352 from JMR-dev/ci-351-pat-triggering
ci: trigger CI with the PAT so dispatches don't need manual approval (#351)
2026-07-05 13:47:39 -05:00
JMR-devandClaude Opus 4.8 2fcee291ce ci: trigger CI with the PAT so dispatches don't need manual approval (#351)
#350 made ci-trigger.yml dispatch ci.yml with the built-in GITHUB_TOKEN, on the
claim that a workflow_dispatch is anti-recursion-exempt so no PAT is needed. In
practice a GITHUB_TOKEN-triggered run is held in `action_required` awaiting manual
approval and never runs un-attended, so auto-updated PRs' CI never ran (stalled
#285). The original #349 design was right: dispatch with a PAT so the run executes
as the authorized owner with no approval gate.

- ci-trigger.yml: the trigger step's GH_TOKEN is now
  `${{ secrets.AUTOUPDATE_TOKEN || github.token }}` (was `${{ github.token }}`).
  AUTOUPDATE_TOKEN (the PAT) is REQUIRED for the scheduler; the `|| github.token`
  fallback stays fail-open but only starts CI if repo settings don't gate
  GITHUB_TOKEN-triggered runs.
- autoupdate.yml: branch update stays on GITHUB_TOKEN (must NOT retrigger CI --
  that would re-introduce the cascade). Clarified that AUTOUPDATE_TOKEN is still
  required by the repo (by ci-trigger.yml) so the secret isn't deleted.
- Corrected the now-wrong "no PAT needed / workflow_dispatch anti-recursion-exempt"
  comments in ci-trigger.yml and the traffic_control.py docstrings.

updates = GITHUB_TOKEN, triggering = PAT.

Validation: all three workflow YAMLs parse clean; traffic-control unit tests still
pass (59 tests) -- the change is workflow-env only, script logic unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 13:28:10 -05:00
github-actions[bot] 33217c6cb2 Merge main into perf-187-covering-index 2026-07-05 18:10:50 +00:00
Jason Ross d8f6a66856 Merge pull request #317 from JMR-dev/fix-306-outlook-redundant-token
fix(auth): drop redundant second Outlook token request on sign-in
2026-07-05 13:10:19 -05:00
JMR-devandClaude Opus 4.8 46bc0e2258 Merge main into perf-187-covering-index
Resolve DatabaseModule conflict from #320: main replaced the explicit .addMigrations(...) chain with .addMigrations(*ALL_MIGRATIONS) plus an introspectable ALL_MIGRATIONS list guarded by databaseModuleRegistersEveryDeclaredMigration (registered == declared). Add MIGRATION_19_20 to ALL_MIGRATIONS so the unified-inbox covering-index migration (cache schema v19->v20) is both registered on the Room builder and satisfies that safety-net test. Schema 20.json, the v20 @Database version, and DatabaseEncryptionTest's schema-version assertion (20) are unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 13:05:07 -05:00
github-actions[bot] 44d6d1edb6 Merge main into fix-306-outlook-redundant-token 2026-07-05 17:52:20 +00:00
Jason Ross 6f4275ee1e Merge pull request #350 from JMR-dev/ci-349-traffic-owns-triggering
ci: traffic-controller owns CI triggering (GITHUB_TOKEN updates, priority-ordered dispatch)
2026-07-05 12:51:48 -05:00
JMR-devandClaude Opus 4.8 05d06eb45b ci: traffic-controller owns CI triggering (GITHUB_TOKEN updates, priority-ordered dispatch)
End the merge cascade and give the traffic-controller ownership of CI *triggering*.

- autoupdate.yml updates PR branches with the built-in GITHUB_TOKEN instead of a PAT,
  so an update push no longer auto-retriggers CI (GitHub's anti-recursion rule) — the
  cascade (every merge re-runs every PR, cancel-in-progress thrashing them) is gone.
- New scheduler ci-trigger.yml -> traffic_control.py --mode trigger (re-)triggers CI
  for the highest-priority PR(s) whose head SHA has absent/stale checks, a few at a
  time (inflight cap), in the existing P0-P9 / broken-draft priority order — a
  poor-man's merge queue reusing the priority core. It runs after autoupdate finishes
  (workflow_run, race-free) plus a cron backstop plus manual dispatch.
- Triggering uses workflow_dispatch, which is EXEMPT from anti-recursion, so the
  built-in GITHUB_TOKEN (actions: write) starts the run — NO PAT / secret change needed.
- ci.yml gains a workflow_dispatch trigger (pr/head_sha/reason inputs) and a per-PR
  concurrency group unifying pull_request and dispatch runs; its on: pull_request path
  is kept so brand-new PRs, human pushes, and fork PRs always get CI (fail-open).

Pure select_triggers / classify_sha_runs decision core added to traffic_control.py with
24 new unit tests (priority order, oldest-first fairness, inflight cap, fork skip, P0
bypass+preempt, head-SHA needy classification, and a liveness/anti-starvation simulation).

Closes #349

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 01:27:40 -05:00
Jason Ross 4bf6fa5535 Merge main into fix-306-outlook-redundant-token 2026-07-05 00:55:24 -05:00
Jason Ross 833dfc030a Merge pull request #339 from JMR-dev/chore-claudemd-dod-logging
chore(docs): require app logging in the definition of done
2026-07-05 00:54:55 -05:00
Jason Ross 11e0a85475 Merge main into fix-306-outlook-redundant-token 2026-07-05 00:37:36 -05:00
Jason Ross d795639a32 Merge main into chore-claudemd-dod-logging 2026-07-05 00:37:35 -05:00
Jason Ross 94a7562256 Merge pull request #344 from JMR-dev/docs-343-workflows-readme
docs: plain-English README for the CI traffic-controller
2026-07-05 00:37:08 -05:00
Jason Ross 7d60c7ac60 Merge main into docs-343-workflows-readme 2026-07-05 00:05:37 -05:00
Jason Ross 88b8f6cb07 Merge main into fix-306-outlook-redundant-token 2026-07-05 00:05:35 -05:00
Jason Ross dd9f6bb8d9 Merge main into chore-claudemd-dod-logging 2026-07-05 00:05:34 -05:00
Jason Ross 348d7adc28 Merge pull request #348 from JMR-dev/feat-331-detekt-log-guard
feat(logging): detekt guard banning android.util.Log outside AppLog (#331)
2026-07-05 00:05:06 -05:00
Jason Ross 45df6c7709 Merge main into chore-claudemd-dod-logging 2026-07-04 23:16:54 -05:00
Jason Ross cefd92864c Merge main into fix-306-outlook-redundant-token 2026-07-04 23:16:53 -05:00
Jason Ross b3ec24d9f3 Merge main into docs-343-workflows-readme 2026-07-04 23:16:52 -05:00
Jason Ross 5d50b60f68 Merge main into feat-331-detekt-log-guard 2026-07-04 23:16:52 -05:00
Jason Ross 5047e3eb2a Merge pull request #340 from JMR-dev/feat-326-logging-authlock
feat(logging): auth/lock -> AppLog + breadcrumbs (#326)
2026-07-04 23:16:23 -05:00
Jason Ross 494649b7d8 Merge main into feat-331-detekt-log-guard 2026-07-04 23:05:20 -05:00
JMR-devandClaude Opus 4.8 faab0e3260 feat(logging): detekt guard banning android.util.Log outside AppLog (#331)
Add a detekt style>ForbiddenImport rule that forbids `import android.util.Log`
so all logging flows through org.libremail.reporting.AppLog, which mirrors each
line into the debug-report RingLogBuffer. A raw android.util.Log import writes to
Logcat only and never reaches a user-reviewed DebugReport (epic #324, strangler
final step).

Excludes the AppLog facade itself (the one sanctioned wrapper) and the unit tests
that mockkStatic(Log) to verify forwarding — AppLog forwards to Log, a throwing
stub under plain JVM unit tests, so those tests must mock it; they do not bypass
the facade.

Closes #331
Part of #324

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 23:02:13 -05:00
Jason Ross a9f0c220da Merge main into docs-343-workflows-readme 2026-07-04 22:59:33 -05:00
Jason Ross a44f568f9c Merge main into feat-326-logging-authlock 2026-07-04 22:59:32 -05:00
Jason Ross c65fb26ad0 Merge main into chore-claudemd-dod-logging 2026-07-04 22:59:31 -05:00
Jason Ross 17d40119ca Merge main into fix-306-outlook-redundant-token 2026-07-04 22:59:31 -05:00
Jason Ross d0ec1949a2 Merge pull request #347 from JMR-dev/ci-346-traffic-control-tests
ci: run traffic-controller unit tests as a gate job
2026-07-04 22:59:01 -05:00
Jason Ross b2c017b52d Merge main into fix-306-outlook-redundant-token 2026-07-04 22:38:21 -05:00
Jason Ross 02153b5287 Merge main into chore-claudemd-dod-logging 2026-07-04 22:38:20 -05:00
Jason Ross ef16b1ef2c Merge main into feat-326-logging-authlock 2026-07-04 22:38:19 -05:00
Jason Ross 76caaddd33 Merge main into docs-343-workflows-readme 2026-07-04 22:38:18 -05:00
JMR-devandClaude Opus 4.8 8b89219734 ci: run traffic-controller unit tests as a gate job
Adds a fast traffic-control-tests job (ubuntu, actions/checkout +
actions/setup-python, no emulator/Gradle) that runs the 37 pure-stdlib
unit tests for .github/scripts/traffic_control.py on every PR, and
wires it into ci-passed's needs so a regression blocks merge instead
of only being caught locally.

Closes #346

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 22:35:39 -05:00
Jason Ross dc21411aed Merge pull request #345 from JMR-dev/ci-342-traffic-control-python
ci: extract traffic-controller into a testable Python module (#342)
2026-07-04 22:33:37 -05:00
Jason Ross a13bc9eb07 Merge main into docs-343-workflows-readme 2026-07-04 22:29:35 -05:00
Jason Ross 37bdaae304 Merge main into feat-326-logging-authlock 2026-07-04 22:29:34 -05:00
Jason Ross 527e31fb5c Merge main into chore-claudemd-dod-logging 2026-07-04 22:29:33 -05:00
Jason Ross 28745827c2 Merge main into fix-306-outlook-redundant-token 2026-07-04 22:29:32 -05:00
Jason Ross 647490c1a8 Merge main into ci-342-traffic-control-python 2026-07-04 22:29:32 -05:00
Jason Ross 785c6aa6c6 Merge pull request #338 from JMR-dev/feat-328-logging-connsend
feat(logging): connectivity/send → AppLog + breadcrumbs, scrub email PII (#328)
2026-07-04 22:29:04 -05:00
JMR-devandClaude Opus 4.8 a34db54b97 fix(ci): let any strictly-higher PR reclaim a broken/draft run (#342)
Restore (and extend to drafts) the old bash's broken-reclaim behaviour that the
initial Python refactor had dropped. runs_to_cancel now cancels an OTHER PR's
active/queued runs when EITHER:
  (a) THIS PR is P0 and that PR is strictly-lower (reclaim every lower runner); OR
  (b) that PR is broken/draft (effective priority 10) and THIS PR is strictly-higher
      (effective priority < 10) — a wasted run any ready PR may reclaim.

P1-P9 still never bump a *normal* (non-broken/draft) lower run; a broken/draft PR
(P10) preempts nothing (nothing is strictly-lower than the bottom, and the
equal-or-higher invariant means a P10 never cancels another P10). Self / main-push /
equal-or-higher invariants unchanged.

Updates the module docstring + ci.yml comments (the "only P0 preempts" wording
becomes: P0 preempts everything strictly-lower; additionally, any strictly-higher PR
preempts a broken/draft run) and the job step/permission/needs comments. Adds unit
tests: P3 reclaims a broken P10 run and a draft P10 run; P3 does not bump a normal P5
run; a P10 self preempts nothing; plus an end-to-end P5-reclaims-draft-then-waits
scenario. 37 unit tests pass; ci.yml parses clean; --dry-run shows a P3 cancelling a
draft (and broken) run while still yielding to a higher P1.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 22:19:01 -05:00
Jason Ross 14eebcb6c1 Merge main into fix-306-outlook-redundant-token 2026-07-04 22:09:43 -05:00
Jason Ross cbf159160f Merge main into feat-328-logging-connsend 2026-07-04 22:09:42 -05:00
Jason Ross d51778379a Merge main into chore-claudemd-dod-logging 2026-07-04 22:09:41 -05:00
Jason Ross 30026c1b29 Merge main into feat-326-logging-authlock 2026-07-04 22:09:40 -05:00
Jason Ross 6055fdc44d Merge main into docs-343-workflows-readme 2026-07-04 22:09:39 -05:00
Jason Ross ab41d8f24b Merge main into ci-342-traffic-control-python 2026-07-04 22:09:15 -05:00
Jason Ross c0dc5c5e75 Merge pull request #336 from JMR-dev/feat-330-logging-stragglers
feat(logging): stragglers -> AppLog (#330)
2026-07-04 22:09:07 -05:00
JMR-devandClaude Opus 4.8 8b7f895d6a ci: extract traffic-controller into a testable Python module (#342)
The priority-based runner orchestration ("traffic-control") lived as a large
inline-bash step in ci.yml — a two-pass preemption + hold-back script that was
effectively untestable in YAML. Move it into .github/scripts/traffic_control.py,
structured as a pure decision CORE + a thin gh-I/O SHELL:

* Pure functions (no network/clock/subprocess), unit-testable in isolation:
  - effective_priority(pr): lowest-numbered P0-P9, default P5; broken OR draft => 10.
  - runs_to_cancel(this_pr, all_prs, self_run_id): PASS 1 — run ids to cancel,
    empty unless THIS PR is P0; only strictly-lower running/queued runs; never self
    (by number or run id), never equal-or-higher.
  - wait_blockers(this_pr, all_prs): PASS 2 — yield to any strictly-higher PR with
    an active/queued run, and to same-level peers ordered ahead (running-first,
    then oldest createdAt). Empty => proceed.
* Shell (run_live): gathers the snapshot via gh, applies cancels, runs the bounded
  hold-back poll loop; always exits 0. --dry-run feeds the core a snapshot JSON and
  prints decisions with zero network.
* No jq/bash dependency (cross-platform, per the repo's Python-stdlib convention).

ci.yml's traffic-control job now checks out the repo and runs the module. Job
permissions gain `contents: read` (for checkout) alongside the existing
`actions: write` / `pull-requests: read`; env and downstream `needs:` wiring
unchanged; step stays `continue-on-error`.

33 stdlib unittest cases cover priority resolution, P0-only preemption, the
self/main/equal-or-higher invariants, and the same-level running-first/oldest
ordering.

Behaviour is preserved except the ticket's refinements: (1) drafts now count as
P10 (bottom); (2) an explicit same-level running-first-then-oldest tiebreaker; and
(3) per the ticket's order-of-operations, ONLY P0 preempts — the old bash also let
any higher-priority PR cancel a `broken` target's run, which no longer happens
(a broken/draft run is only cancelled by a P0, via the same strictly-lower rule).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 22:08:13 -05:00
JMR-devandClaude Opus 4.8 a24435a4fd docs(ci): add plain-English README for the CI traffic-controller
Explains the traffic-control job in ci.yml (priority labels, preemption
vs. bounded hold-back, safety invariants, and the known FIFO-runner
limitation) for developers new to the repo. Notes that #342 will refactor
this logic into a tested Python module, at which point this doc gets
updated.

Closes #343

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 22:06:09 -05:00
Jason Ross fc38255be5 Merge main into feat-326-logging-authlock 2026-07-04 21:52:01 -05:00
Jason Ross 6f20a23ae9 Merge main into chore-claudemd-dod-logging 2026-07-04 21:51:59 -05:00
Jason Ross 7e0b5c2b1c Merge main into feat-328-logging-connsend 2026-07-04 21:51:58 -05:00
Jason Ross 793346ab46 Merge main into feat-330-logging-stragglers 2026-07-04 21:51:57 -05:00
Jason Ross fff8afb453 Merge main into fix-306-outlook-redundant-token 2026-07-04 21:51:56 -05:00
Jason Ross b912af0ee2 Merge pull request #341 from JMR-dev/feat-329-logging-sync
feat(logging): sync-engine breadcrumbs via AppLog (#329)
2026-07-04 21:51:26 -05:00
Jason Ross 03179cda15 Merge main into fix-306-outlook-redundant-token 2026-07-04 21:34:12 -05:00
Jason Ross 8cc792c06e Merge main into feat-330-logging-stragglers 2026-07-04 21:34:11 -05:00
Jason Ross 59f86016cb Merge main into feat-328-logging-connsend 2026-07-04 21:34:10 -05:00
Jason Ross a891be4b14 Merge main into chore-claudemd-dod-logging 2026-07-04 21:34:09 -05:00
Jason Ross ca4c310a66 Merge main into feat-326-logging-authlock 2026-07-04 21:34:08 -05:00
Jason Ross 3d4dde6bc5 Merge main into feat-329-logging-sync 2026-07-04 21:34:07 -05:00
Jason Ross ff49c6c410 Merge pull request #337 from JMR-dev/feat-327-logging-dbkeystore
feat(logging): DB/keystore -> AppLog + breadcrumbs (#327)
2026-07-04 21:33:41 -05:00
Jason Ross 209371c9a5 Merge main into feat-329-logging-sync 2026-07-04 21:32:36 -05:00
JMR-devandClaude Opus 4.8 59b4252ce8 feat(logging): sync-engine breadcrumbs via AppLog (#329)
The sync engine (MailSyncer, MailBackfiller, MailPruner, and their WorkManager
workers) was completely silent, so a submitted debug report showed nothing
about whether sync ran, how much it fetched, or why it was skipped. Add
net-new AppLog breadcrumbs at each class's lifecycle points per the #324
strangler-migration plan: sync start/done/failed and per-folder fetch counts,
backfill slice start/done and per-folder page counts, prune's removed count,
and each worker's cache-locked deferral and success/retry outcome (the retry
path now also carries the scrubbed failure throwable via AppLog's #325
overloads).

Every breadcrumb is PII-safe by construction: accounts are identified only via
accountLogRef(account.id) (never the id or email directly), and a new
logSafeFolderLabel() helper logs a folder's name only when it matches a fixed
allowlist of known system folders (INBOX, Sent, Drafts, Trash, Spam/Junk,
Archive, and their common provider variants) — every other folder, however
nested or named, logs as a fixed placeholder.

Adding logging to these previously-silent classes meant every existing test
exercising them now hits android.util.Log (a throwing stub under plain JVM
unit tests), so each affected suite gains the same static Log mock already
established by AppLogTest/SendWorkerTest/ImapClientTest.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 21:31:46 -05:00
Jason Ross d6e886bde0 Merge main into feat-326-logging-authlock 2026-07-04 21:28:28 -05:00
JMR-devandClaude Opus 4.8 0236494ecb chore(docs): require app logging in the definition of done
Codify the maintainer rule that app source-code changes must add
appropriate, PII-free logging via the AppLog facade at key points
(lifecycle transitions, error/fallback paths, state changes) so
behaviour is diagnosable from a user's debug report.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 21:27:39 -05:00
JMR-devandClaude Opus 4.8 6e0fe14b06 feat(logging): auth/lock -> AppLog + breadcrumbs (#326)
Migrate AppLockViewModel (9 sites) and AccountSetupViewModel (1 site) off
raw android.util.Log onto the AppLog seam (#325), so their diagnostic
lines land in the RingLogBuffer and reach a submitted DebugReport instead
of only Logcat. Adds three new breadcrumbs that were previously silent:
auth-seal unlock success, the clear-cache-and-restart recovery trigger
(with the disableAppLock flag), and every onForeground LockAction
decision. AccountSetupViewModel's success path also now logs "Outlook
account added" (no email). None of these call sites carry PII; where a
throwable is attached, AppLog's StackTraceScrubber redacts it before it
reaches the buffer.

Both ViewModel test suites now install a real RingLogBuffer and assert
against it instead of `verify { Log... }`, including dedicated no-PII
assertions (a known test email never appears in a recorded line). A
minimal `mockkStatic(Log::class)` stub stays in both test files' shared
setUp — AppLog still forwards to the real android.util.Log internally,
which throws "not mocked" in JVM unit tests when uninvoked; the not-yet
-landed guard-rule ticket (#331) will need to reconcile that with a
repo-wide "no raw Log outside AppLog.kt" rule.

Closes #326
Part of #324

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 21:27:30 -05:00
Jason Ross c149411a9d Merge main into feat-328-logging-connsend 2026-07-04 21:23:29 -05:00
JMR-devandClaude Opus 4.8 21c74df794 feat(logging): connectivity/send -> AppLog + breadcrumbs, scrub email PII
Migrate ImapClient/SendWorker/IdleService off raw android.util.Log to AppLog,
per the debug-logging strangler epic (#324), so their diagnostics reach the
RingLogBuffer (and a submitted DebugReport) instead of logcat-only:

- ImapClient: IDLE connect + IDLE push (message count) breadcrumbs.
- SendWorker: outbox-drain count on entry, per-message sent/failed result,
  and the existing Graph->SMTP fallback warning.
- IdleService: IDLE watch start, cache-locked defer, and the existing
  IDLE-dropped/retrying warning.

Also closes #297: SendWorker and IdleService logged the raw account.email via
Log.w on the Graph->SMTP fallback and IDLE-drop paths. Both now log
accountLogRef(account.id) instead -- a short, stable, non-reversible
per-account reference -- so the account's email never reaches Logcat or a
report.

Rewrites ImapClientTest/SendWorkerTest to install a real RingLogBuffer via
AppLog.install(...) and assert on its contents (migrated calls + new
breadcrumbs), instead of verifying a mocked Log; every assertion also checks
no line carries the test account's email, regression-covering #297. Adds a
SendWorkerTest case that drives a real SmtpSender against an in-process
GreenMail SMTP server end to end. android.util.Log is still stubbed (by
fully-qualified name, without importing it) where AppLog's Logcat passthrough
would otherwise crash the unmocked Android stub in a JVM test.

Closes #328
Closes #297
Part of #324

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 21:22:37 -05:00
Jason Ross f642f2c2b5 Merge main into fix-306-outlook-redundant-token 2026-07-04 21:15:54 -05:00
Jason Ross 73c636704f Merge main into feat-330-logging-stragglers 2026-07-04 21:15:53 -05:00
Jason Ross eee1475a61 Merge main into feat-327-logging-dbkeystore 2026-07-04 21:15:53 -05:00
Jason Ross a2f96b930d Merge pull request #335 from JMR-dev/ci-334-api37-boot-diagnostics
ci(e2e): boot diagnostics + verbose/debug logging on the API 37 preview job
2026-07-04 21:15:24 -05:00
JMR-devandClaude Opus 4.8 0cb9bb2906 refactor(data): route DB/keystore logging through AppLog (#327)
Migrates the DB/keystore area's raw android.util.Log calls to AppLog so
key-invalidation and DB-conversion breadcrumbs land in the process
RingLogBuffer (and thus a user-reviewed debug report) even in release
builds, where Log.d is otherwise stripped from Logcat only.

- DatabaseKeyCipher: 4 auth-bound-key decision points (encrypt retry,
  isInvalidated's three branches) now log via AppLog.d(tag, msg, e).
- DatabaseEncryption.migrate: adds an AppLog.i "converting local cache
  database (targetEncrypted=...)" breadcrumb at the start, alongside the
  existing "converted" completion line now routed through AppLog.d.
- AccountDataMigrator: the "moved account tables into the account
  database: $present" breadcrumb (table names only) now routed through
  AppLog.d.

No PII or key material is logged; table-name sets and boolean flags only.

Adds instrumented tests (DatabaseKeyCipher is device-only and
behavior-preserving, so no new test there) asserting the breadcrumbs
land in a RingLogBuffer and never contain the seeded email, secret, or
passphrase.

Part of #324.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 21:12:59 -05:00
JMR-devandClaude Opus 4.8 9e771a9590 ci(e2e): boot diagnostics + verbose/debug logging on the API 37 preview job
The API 37 preview E2E boot has flaked twice (#285, #333) with only
"did not boot within 300s" and no root-cause signal. Add rich boot
diagnostics by default, kept in parity between CI and the local
hand-provisioning script (api37_e2e.py):

- Launch the emulator with `-verbose -debug init,avd_config,kernel`
  (diagnostics only; no boot-affecting flag changed), still redirecting
  to $EMU_LOG.
- Stream `adb logcat -v time` to a file from the moment the device
  registers (via `adb wait-for-device logcat`, backgrounded).
- On a boot timeout, dump accel-check, /dev/kvm presence, GPU mode,
  free mem/disk, the AVD config.ini and the emulator.log tail; CI writes
  these to a boot-diagnostics file, the local script prints them.
- CI uploads emulator.log + logcat.txt + boot-diagnostics.txt as an
  artifact with `if: always()` so they survive a timeout/cancel, and
  prints a concise summary (accel/KVM status + last 50 lines of
  emulator.log) to the step log.

The existing 2-attempt boot retry + boot-completed wait loop are
unchanged; the diagnostics are additive.

Closes #334

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 21:09:28 -05:00
JMR-devandClaude Opus 4.8 d92d4c1a14 feat(logging): stragglers -> AppLog (#330)
Migrate RestartActivity's one raw Log.w site to AppLog, clearing the
final raw android.util.Log site outside the auth/lock, DB/keystore,
connectivity/send, and sync-engine migration areas so the codebase is
ready for the detekt android.util.Log guard (#331).

RestartActivity runs in the separate :restart trampoline process,
where LibreMailApplication.onCreate returns early and never calls
AppLog.install, so this breadcrumb reaches Logcat only, never a
DebugReport. The migration is guard-compliance + Logcat-consistency
only; behavior is unchanged since AppLog forwards to Logcat.

RestartActivity is DEVICE-ONLY (multi-process kill/relaunch), so a
JVM buffer-capture test doesn't apply here. Added
RestartActivityLoggingTest, which instead pins the null-buffer shape
this call runs under in the trampoline process: it forwards to
Logcat and no-ops the buffer cleanly.

Closes #330
Part of #324

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 21:09:22 -05:00
Jason Ross 6893649f0d Merge main into fix-306-outlook-redundant-token 2026-07-04 21:02:12 -05:00
Jason Ross 41015a5c40 Merge pull request #333 from JMR-dev/feat-325-applog-seam
feat(logging): AppLog seam — record scrubbed throwables + accountLogRef (#325)
2026-07-04 21:01:39 -05:00
Jason Ross 9ea339436d Merge main into feat-325-applog-seam 2026-07-04 20:24:29 -05:00
JMR-devandClaude Opus 4.8 1a1fbf8d7f feat(logging): AppLog seam — record scrubbed throwables + accountLogRef (#325)
Add throwable-recording overloads to AppLog.d/w and make AppLog.e record the
throwable it is given: the throwable's stack trace is scrubbed via the existing
StackTraceScrubber (exception class names + frames kept; host/email-bearing
exception messages stripped) and appended to the buffered log line, so a
throwable can reach a user-reviewed DebugReport without leaking PII. The
existing no-throwable overloads are unchanged.

Add accountLogRef(accountId): a short, stable, non-reversible reference
(scheme prefix + truncated SHA-256 of the id) so downstream logging can
identify an account without logging the raw Account.id, which embeds the email.

Foundation for the #324 debug-logging strangler epic; consumed by #326–#330.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 20:23:38 -05:00
Jason Ross 0d9761dda2 Merge main into fix-306-outlook-redundant-token 2026-07-04 20:17:19 -05:00
Jason Ross b352e3838c Merge pull request #332 from JMR-dev/feat-281-python-dev-scripts
test-infra: port local_instrumented.sh to Python + rewire preflight off GMD (#281)
2026-07-04 20:16:51 -05:00
JMR-devandClaude Opus 4.8 f54e9c67fa test-infra: port local_instrumented.sh to Python + rewire preflight off GMD (#281)
Port the last bash dev-script (local_instrumented.sh) to a cross-platform,
stdlib-only Python 3 script (local_instrumented.py), matching api37_e2e.py's
style, and rewire the /preflight skill's local E2E off the GMD
apiXXDebugAndroidTest tasks (which fail locally under AEHD 2.2) onto it.

- local_instrumented.py preserves the .sh's behavior exactly: comma-separated
  test-class CLI arg, pre-boot orphan-kill, manual cold-boot of the dev36 AVD
  (no GMD, no snapshot), targeted connectedDebugAndroidTest, and the EXIT-trap
  teardown (now try/finally + atexit + SIGINT/SIGTERM handlers, idempotent).
  Exit codes 0/2/3/4 preserved.
- Cross-platform process kill abstracted per-OS: taskkill /F /IM on Windows,
  pkill -f qemu-system on *nix; process listing via tasklist / ps ax.
- Teardown hardened vs the .sh: it now also reaps the emulator *launcher*
  image, not just qemu -- the Windows -no-window emulator spawns a sibling
  emulator.exe that briefly outlives the qemu VM, which a qemu-only sweep left
  as an orphan on return (caught by the smoke run).
- SKILL.md + CLAUDE.md: replace the local api35/api36 GMD E2E steps with
  local_instrumented.py; CI's own multi-API matrix is untouched. CLAUDE.md
  documents the Python-first dev-script convention.

Validated: py_compile, argparse (--help / no-arg exit 2), and a guarded
emulator smoke run of org.libremail.data.local.DatabaseEncryptionTest -- boots,
passes, and tears down clean (no qemu/emulator orphan on return).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 19:58:14 -05:00
Jason Ross be2afd42c4 Merge main into fix-306-outlook-redundant-token 2026-07-04 11:46:19 -05:00
Jason Ross 47aaebccf5 Merge pull request #323 from JMR-dev/feat-251-292-coverage-noregression-gate
feat(ci): no-regression JVM coverage gate scoped to the testable surface
2026-07-04 11:45:47 -05:00
JMR-devandClaude Opus 4.8 2e0eaef4e9 feat(ci): no-regression JVM coverage gate scoped to the testable surface
Scope :app:jacocoTestReport's denominator to the JVM-testable surface and
add a :app:jacocoTestCoverageVerification no-regression gate that shares the
same classDirectories/executionData/sourceDirectories, wired into both the
`check` lifecycle task and CI's unit-test job (part of the `CI passed` gate).

Excluded from the denominator (structurally unreachable from a JVM unit
test): Compose screen/component render code, Android framework entry points
(*Activity/*Service/Application/*BackupAgent), Hilt DI (**/di/**), and the
src/debug cold-open probe. Kept in scope: ViewModels, repositories, mappers,
DAOs, utils, richtext, mail, reporting logic, and the six WorkManager Workers.

Corrects PR #292, which excluded **/*Worker*: SyncWorker, BackfillWorker,
PruneWorker, SendWorker, ReportPurgeWorker and ReportUploadWorker are all
directly unit-tested, so they stay counted in both numerator and denominator
(only their Hilt wiring, WorkManagerModule, is excluded, via **/di/**).

Baseline: 80.21% line (4838/6032). Floor: 0.79 (~1.2% headroom) so ordinary
noise doesn't red-flag it while a real drop fails. Manual ratchet for now:
bump the floor up in the same PR when coverage rises materially.

Closes #251
Closes #292

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 11:27:17 -05:00
Jason Ross cc1f4442b6 Merge main into fix-306-outlook-redundant-token 2026-07-04 04:00:45 -05:00
JMR-devandClaude Opus 4.8 e4457cfec1 test(db): expect schema v20 in encryption round-trip after v19->v20 bump
MIGRATION_19_20 (issue #187) bumped the Room cache schema to version 20,
but DatabaseEncryptionTest.schemaVersionIsCarriedOntoTheEncryptedFile still
asserted the plaintext -> encrypted conversion carried version 19, so it
failed across all E2E levels after the rebase onto main.

DatabaseEncryption.migrate() carries PRAGMA user_version dynamically
(userVersion = source.version -> target.version = userVersion), and a fresh
Room open now stamps 20, so v20 genuinely survives the conversion. Update the
expected constant to 20; the assertion's intent (the version survives the
round-trip) is unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 03:53:46 -05:00
Jason Ross 2e04934e19 Merge main into fix-306-outlook-redundant-token 2026-07-04 03:43:43 -05:00
Jason Ross 8cb57cba50 Merge main into fix-306-outlook-redundant-token 2026-07-04 03:21:29 -05:00
Jason Ross 2289f402e5 Merge main into perf-187-covering-index 2026-07-04 03:01:21 -05:00
Jason Ross 97b782ecb2 Merge main into fix-306-outlook-redundant-token 2026-07-04 03:01:19 -05:00
Jason Ross b33f73273d Merge main into fix-306-outlook-redundant-token 2026-07-04 02:45:03 -05:00
Jason Ross 1d4bd6346c Merge main into perf-187-covering-index 2026-07-04 02:44:59 -05:00
Jason Ross df57dcd18d Merge main into perf-187-covering-index 2026-07-04 02:26:52 -05:00
Jason Ross e80f8333f4 Merge main into fix-306-outlook-redundant-token 2026-07-04 02:26:49 -05:00
JMR-devandClaude Opus 4.8 da2963aa86 fix(auth): drop redundant second Outlook token request on sign-in
exchangeToken's authorization-code exchange already requests
`openid email offline_access $OUTLOOK_SCOPE`, so the returned access token
is an outlook.office.com token usable for IMAP verification and the AuthState
already carries the refresh token and expiry. The immediate follow-up
refreshForScope(authState, OUTLOOK_SCOPE) was a second round-trip for the
same resource that only rotated the just-issued refresh token and added a
needless onboarding failure point (a transient network error there failed
sign-in after consent + code-exchange had already succeeded).

Build OAuthResult directly from the code-exchange tokenResponse
(accessToken + authState.jsonSerializeString()), dropping the extra refresh.
The durable AuthState is still serialized and persisted for later token
refresh; the Graph token remains a distinct resource minted on demand via
freshGraphToken.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 02:22:00 -05:00
Jason Ross 8230dd0341 Merge main into perf-187-covering-index 2026-07-04 01:56:44 -05:00
Jason Ross 5c34b01bf0 Merge main into perf-187-covering-index 2026-07-04 01:22:27 -05:00
Jason Ross f4a95a6051 Merge main into perf-187-covering-index 2026-07-04 00:59:28 -05:00
Jason Ross 991f9b77e4 Merge main into perf-187-covering-index 2026-07-04 00:38:29 -05:00
Jason Ross bdb49f996f Merge main into perf-187-covering-index 2026-07-04 00:17:05 -05:00
JMR-devandClaude Opus 4.8 5a3669f017 perf(db): covering index for unified-inbox summary scans
The paged "All inboxes" query (MessageDao.pagingUnifiedFolderSummaries:
WHERE folder = ? AND inInbox = 1 ORDER BY timestampMillis DESC) had no
folder-leading index, so it SCANned the whole messages table via
index_messages_timestampMillis and filtered folder/inInbox per row.

Add a (folder, inInbox, timestampMillis) index so the two equality
predicates become an index seek and the ORDER BY is supplied by the
index. EXPLAIN QUERY PLAN for the query goes from
  SCAN messages USING INDEX index_messages_timestampMillis
to
  SEARCH messages USING INDEX index_messages_folder_inInbox_timestampMillis (folder=? AND inInbox=?)
with no temp B-tree sort.

Pure additive index (no column/table change): bump the Room DB to v20
with MIGRATION_19_20 (CREATE INDEX IF NOT EXISTS), register it in
DatabaseModule, export 20.json, and add a MigrationTest that runs the
migration and asserts the index shape plus the SEARCH plan on real
Android SQLite.

Closes #187

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-04 00:14:39 -05:00
163 changed files with 19401 additions and 1058 deletions
+49 -33
View File
@@ -1,6 +1,6 @@
---
name: preflight
description: Run LibreMail's fast CI gate locally (assembleDebug + testDebugUnitTest + compileDebugAndroidTestKotlin + lintDebug + ktlintCheck + detekt) plus the top-of-matrix emulator E2E (currently API 35 + 36 via Gradle Managed Devices, then API 37 preview via the hand-provisioning script) before pushing or opening a PR. Mirrors the merge gate; the rest of the multi-API matrix stays CI-only. Use before treating a change as done.
description: Run LibreMail's fast CI gate locally (assembleDebug + testDebugUnitTest + jacocoTestCoverageVerification + compileDebugAndroidTestKotlin + lintDebug + ktlintCheck + detekt) plus the local emulator E2E — the instrumented test class(es) you changed via local_instrumented.py (cold-boot, no Gradle Managed Devices), then the API 37 preview via api37_e2e.py — before pushing or opening a PR. Mirrors the merge gate; CI runs the full multi-API matrix. Use before treating a change as done.
---
# /preflight
@@ -13,20 +13,24 @@ Run the same fast checks CI enforces on every PR, in order, and report the outco
with a JDK/AGP version mismatch, check `java -version` / `JAVA_HOME` and point it at a 17–21
JDK (e.g. Android Studio's bundled JBR) before retrying.
- PowerShell: invoke the wrapper as `.\gradlew`. Git Bash / the Bash tool: `./gradlew`.
- The final three E2E steps each boot an emulator, so the host needs a **free hardware
- The final two E2E steps each boot an emulator, so the host needs a **free hardware
hypervisor** (Intel VT-x / AMD-V, exposed as WHPX on Windows, KVM on Linux, HVF on macOS).
Shut down VirtualBox, Hyper-V-based VMs, WSL2, Docker Desktop, or any other emulator first — a
VM holding the hypervisor starves the AVD, and it hangs at 0% CPU and never reaches
`sys.boot_completed`. The api35/api36 steps run through Gradle Managed Devices (Gradle
downloads the image and boots/tears down each AVD itself); the api37 step is hand-provisioned
by `api37_e2e.py` (see Steps). The first run per API level is slow while its system image
`sys.boot_completed`. Both steps are hand-provisioned by cross-platform Python scripts (**not**
Gradle Managed Devices, which fail locally on this box — see Steps): `local_instrumented.py`
cold-boots one existing AVD and runs the instrumented class(es) you changed; `api37_e2e.py`
installs + boots the API 37 preview image. The first API 37 run is slow while its system image
downloads. If the host has no accelerated emulator and a device cannot boot, report the E2E
step as not run rather than treating the gate as green.
- The api37 step is a stdlib-only, cross-platform **Python 3** script and needs `python3` plus
the Android SDK command-line tools (`sdkmanager`/`avdmanager`) and `emulator` on the machine,
located via `ANDROID_SDK_ROOT`/`ANDROID_HOME` (or the per-OS default:
`%LOCALAPPDATA%\Android\Sdk` on Windows, `~/Library/Android/sdk` on macOS, `~/Android/Sdk` on
Linux). The script installs the preview system image itself on first run.
- Both E2E steps are stdlib-only, cross-platform **Python 3** scripts. `local_instrumented.py`
needs `python3` plus the Android SDK `emulator` + `adb` on `PATH` and an existing AVD (any local
`apiXXDebugAndroidTest` run creates one; override with `LOCAL_INSTRUMENTED_AVD` /
`ANDROID_AVD_HOME`), and it pins `JAVA_HOME` to a JDK 17–21 itself (override with
`LOCAL_INSTRUMENTED_JDK`). `api37_e2e.py` additionally needs the Android SDK command-line tools
(`sdkmanager`/`avdmanager`), located via `ANDROID_SDK_ROOT`/`ANDROID_HOME` (or the per-OS
default: `%LOCALAPPDATA%\Android\Sdk` on Windows, `~/Library/Android/sdk` on macOS,
`~/Android/Sdk` on Linux); it installs the preview system image itself on first run.
## Steps
@@ -34,15 +38,19 @@ Run these, stopping at the first failure:
```bash
./gradlew :app:assembleDebug
./gradlew :app:testDebugUnitTest
./gradlew :app:testDebugUnitTest :app:jacocoTestCoverageVerification
./gradlew :app:compileDebugAndroidTestKotlin
./gradlew :app:lintDebug
./gradlew :app:ktlintCheck :app:detekt
./gradlew :app:api35DebugAndroidTest # top-of-matrix emulator E2E (2nd-highest stable level)
./gradlew :app:api36DebugAndroidTest # top-of-matrix emulator E2E (highest stable level)
python3 .claude/skills/preflight/local_instrumented.py <your.Changed.TestClass>[,<Class2>,...] # local instrumented/E2E, cold-boot (no GMD)
python3 .claude/skills/preflight/api37_e2e.py # api37 preview E2E (hand-provisioned; on Windows: py or python)
```
`jacocoTestCoverageVerification` runs right after `testDebugUnitTest` because it reads that
task's JVM exec data — it enforces the whole-app **no-regression line-coverage floor (currently
0.84)**, so a coverage regression is caught locally instead of only in CI (the exact class of
failure that reached CI on #367).
`compileDebugAndroidTestKotlin` compiles the `androidTest` source set — the E2E/instrumented
tests — without needing an emulator. `assembleDebug`, `testDebugUnitTest`, and `lintDebug` never
compile that source set, so a change that breaks it (e.g. an instrumented test calling a UI API
@@ -54,12 +62,19 @@ merge gate even when the build and lint are green. Add `--continue` to any comma
`:app:ktlintCheck :app:detekt --continue`) to collect every failure in one pass instead of
stopping at the first.
The three E2E steps run last because they are the slowest. `api35DebugAndroidTest` and
`api36DebugAndroidTest` run the full instrumented/E2E suite on `api35`, then `api36` — the top
two stable levels in the E2E matrix — each via its own Gradle Managed Device, which Gradle
provisions, boots, and tears down automatically.
The two E2E steps run last because they are the slowest, and they run through **cross-platform
Python scripts, not Gradle Managed Devices (GMD)**. GMD's `apiXXDebugAndroidTest` tasks fail
locally on this box — GMD's AVD-snapshot step times out under the AEHD 2.2 hypervisor
(`AvdSnapshotHandler$EmulatorSnapshotCannotCreatedException`), cycling for hours — so preflight
does **not** call them (issue #269/#281). `local_instrumented.py` instead cold-boots one existing
AVD by hand with `-no-snapshot` (the exact `connectedDebugAndroidTest` technique CI and
`api37_e2e.py` use), runs `:app:connectedDebugAndroidTest` filtered to the instrumented class(es)
you pass, then tears the emulator down and verifies no orphaned `qemu` process is left behind
(exit 3 if one survives). Pass the instrumented/E2E class(es) you actually changed
(comma-separated, no spaces) — the full ~114-test suite tends to wedge mid-run locally, so
targeted runs are deliberate; the whole suite across every API level is CI's job.
`api37_e2e.py` then runs the same suite on the **API 37 preview** emulator. API 37 has no Gradle
`api37_e2e.py` then runs the instrumented/E2E suite on the **API 37 preview** emulator. API 37 has no Gradle
Managed Device — its only published system image is the nonstandard `android-37.0` /
`google_apis_ps16k` pairing, which neither `ManagedVirtualDevice`'s `apiLevel` (Int) nor
`apiPreview` (codename) DSL resolves (see the comment above `testOptions.managedDevices` in
@@ -72,26 +87,27 @@ match `e2e-preview` with one deliberate local exception: the **GPU mode**. CI us
local run uses `-gpu auto-no-window`, which renders on the host GPU — faster, and the mode that
boots cleanly on a dev machine.
All three levels are the E2E that preflight runs locally and all three must pass; CI fans the
same suite out across the whole matrix (API 29–36 in `e2e`, plus API 37 in `e2e-preview`). Keep
`api35DebugAndroidTest` / `api36DebugAndroidTest` in lockstep with the top of the managed-device
list in `app/build.gradle.kts`, and keep `api37_e2e.py` in lockstep with the `e2e-preview` job in
`.github/workflows/ci.yml` (same image string + emulator flags, apart from the intentional GPU-mode
difference noted above) — when a newer API level is added there, run the new top levels instead.
Both E2E steps are the E2E that preflight runs locally and both must pass; CI then fans the full
instrumented/E2E suite out across the whole matrix (API 29–36 in `e2e`, plus API 37 in
`e2e-preview`). Keep `local_instrumented.py` pointed at the instrumented class(es) you changed,
and keep `api37_e2e.py` in lockstep with the `e2e-preview` job in `.github/workflows/ci.yml`
(same image string + emulator flags, apart from the intentional GPU-mode difference noted above).
The local gate no longer runs the GMD `apiXXDebugAndroidTest` tasks (they are unusable locally —
see above); full multi-API coverage stays CI's job.
## Reporting
- If everything passes, say so plainly (e.g. "preflight green: build, unit tests, lint, ktlint, detekt, api35+api36+api37 E2E").
- If everything passes, say so plainly (e.g. "preflight green: build, unit tests, lint, ktlint, detekt, local + api37 E2E").
- On failure, surface the actual Gradle error and point at the relevant report:
- unit tests → `app/build/reports/tests/testDebugUnitTest/`
- lint → `app/build/reports/lint-results-debug.html`
- ktlint → `app/build/reports/ktlint/` (per source set, e.g. `ktlintTestSourceSetCheck/`)
- detekt → `app/build/reports/detekt/`
- api35/api36 E2E → `app/build/reports/androidTests/managedDevice/` (per-device HTML, e.g.
`.../api35/`, `.../api36/`)
- api37 E2E → `app/build/reports/androidTests/connected/` (the `connectedDebugAndroidTest`
report the script drives); the emulator's own boot log is at the temp path the script prints.
- Run the **top two stable-API** levels (`api35DebugAndroidTest` + `api36DebugAndroidTest`) plus
the **API 37 preview** via `api37_e2e.py`; the rest of the multi-API matrix (API 29–34) stays
CI's job. If the host has no accelerated emulator and a device cannot boot (see the hypervisor
note in Preconditions), report that E2E could not run rather than treating the gate as green.
- local + api37 E2E → `app/build/reports/androidTests/connected/` (the `connectedDebugAndroidTest`
report both scripts drive — the later run overwrites the earlier); each emulator's own boot log
is at the temp path the script prints (`local_instrumented.py` also exits **3** if it leaves an
orphaned `qemu`, **4** if the emulator never booted).
- Run the instrumented class(es) you changed via `local_instrumented.py`, plus the **API 37
preview** via `api37_e2e.py`; the full multi-API matrix (API 29–37) stays CI's job. If the host
has no accelerated emulator and a device cannot boot (see the hypervisor note in Preconditions),
report that E2E could not run rather than treating the gate as green.
+126 -8
View File
@@ -38,6 +38,7 @@ from __future__ import annotations
import argparse
import os
import platform
import shutil
import subprocess
import sys
import tempfile
@@ -52,6 +53,11 @@ BUILD_TOOLS = "build-tools;37.0.0"
AVD_NAME = "api37"
DEVICE_PROFILE = "pixel_2"
BOOT_TIMEOUT = 300
# GPU mode: the ONE deliberate divergence from CI's e2e-preview (which uses `swiftshader_indirect`
# for headless determinism). Locally we render on the host GPU -- faster, and the mode that boots
# cleanly on a dev machine. See start_emulator. Kept as a constant so start_emulator and the
# boot-failure diagnostics dump report the same value.
GPU_MODE = "auto-no-window"
IS_WINDOWS = os.name == "nt"
BAT = ".bat" if IS_WINDOWS else ""
@@ -120,15 +126,17 @@ def create_avd(avdmanager: str, emulator: str) -> None:
def start_emulator(emulator: str, emu_log: Path, attempt: int) -> subprocess.Popen:
print(f"Starting API 37 emulator (attempt {attempt})...")
# Flags mirror .github/workflows/ci.yml e2e-preview (cold headless boot, hardware accel
# required, no cameras), with ONE deliberate LOCAL exception -- the GPU mode. CI uses
# `-gpu swiftshader_indirect` (software rendering, deterministic on a headless CI runner);
# locally we use `-gpu auto-no-window`, which renders on the host GPU: faster, and the mode
# that boots cleanly on a dev machine. Keep everything except the GPU mode in lockstep with
# that job.
# required, no cameras), with ONE deliberate LOCAL exception -- the GPU mode (GPU_MODE above:
# CI uses `-gpu swiftshader_indirect`, deterministic on a headless CI runner; locally we render
# on the host GPU -- faster, and the mode that boots cleanly on a dev machine). `-verbose -debug
# init,avd_config,kernel` turns emulator boot logging on by default (mirrors CI) so a boot flake
# is diagnosable from $EMU_LOG; it is DIAGNOSTICS ONLY and does not change any boot-affecting
# flag. Keep everything except the GPU mode in lockstep with that job.
flags = [
"-avd", AVD_NAME,
"-no-window", "-no-audio", "-no-boot-anim", "-no-snapshot", "-accel", "on",
"-gpu", "auto-no-window", "-camera-back", "none", "-camera-front", "none",
"-gpu", GPU_MODE, "-camera-back", "none", "-camera-front", "none",
"-verbose", "-debug", "init,avd_config,kernel",
]
log = open(emu_log, "wb") # noqa: SIM115 - handed to the child; closed in the parent below
try:
@@ -189,6 +197,102 @@ def tail(path: Path, lines: int = 80) -> None:
pass
def _accel_check(emulator: str) -> str:
"""`emulator -accel-check` output -- the accelerator status (WHPX / KVM / HVF availability)."""
try:
out = subprocess.run(cmd(emulator, "-accel-check"), capture_output=True, text=True,
check=False)
return (out.stdout + out.stderr).strip() or f"(no output; exit {out.returncode})"
except OSError as exc:
return f"(accel-check failed: {exc})"
def _kvm_status() -> str:
"""/dev/kvm presence (Linux). Off-Linux the accelerator is WHPX/HVF -- see -accel-check."""
if os.path.exists("/dev/kvm"):
return "/dev/kvm present"
return f"/dev/kvm absent (expected off-Linux; platform={platform.system()})"
def _mem_info() -> str:
"""Free/total memory. Reads /proc/meminfo on Linux (where CI runs); best-effort elsewhere."""
try:
meminfo = Path("/proc/meminfo")
if meminfo.exists():
wanted = {"MemTotal", "MemFree", "MemAvailable"}
lines = [line.strip() for line in meminfo.read_text().splitlines()
if line.split(":", 1)[0] in wanted]
if lines:
return "; ".join(lines)
except OSError:
pass
return f"(memory stats unavailable on {platform.system()})"
def _disk_info(path: Path) -> str:
"""Free/total disk for the filesystem holding `path` (cross-platform via shutil.disk_usage)."""
try:
usage = shutil.disk_usage(path)
gib = 1024 ** 3
return f"total={usage.total / gib:.1f}GiB free={usage.free / gib:.1f}GiB ({path})"
except OSError as exc:
return f"(disk stats unavailable: {exc})"
def start_logcat(adb: str, logcat_log: Path, attempt: int) -> subprocess.Popen | None:
"""Background `adb wait-for-device logcat -v time` to a file. wait-for-device blocks until the
device registers, so streaming starts the moment the emulator appears and captures the whole
boot. Mirrors CI's e2e-preview logcat capture; appended (with a header) per boot attempt."""
try:
with open(logcat_log, "a") as marker:
marker.write(f"===== logcat (attempt {attempt}) =====\n")
log = open(logcat_log, "ab") # noqa: SIM115 - child inherits fd; parent copy closed below
try:
return subprocess.Popen(cmd(adb, "wait-for-device", "logcat", "-v", "time"),
stdout=log, stderr=subprocess.STDOUT)
finally:
log.close() # the child has inherited its own fd; the parent's copy is no longer needed
except OSError as exc:
print(f"WARNING: could not start logcat capture: {exc}", file=sys.stderr)
return None
def stop_logcat(proc: subprocess.Popen | None) -> None:
if proc and proc.poll() is None:
proc.terminate()
try:
proc.wait(timeout=5)
except subprocess.TimeoutExpired:
proc.kill()
def dump_diagnostics(adb: str, emulator: str, emu_log: Path, avd_home: Path, attempt: int) -> None:
"""Print boot diagnostics + a concise failure summary to the console -- the local mirror of CI's
e2e-preview boot-timeout dump (accel/KVM/GPU/mem/disk/AVD config + emulator.log tail). Local
runs PRINT these; CI uploads the same set as an artifact and prints only the concise summary."""
accel = _accel_check(emulator)
kvm = _kvm_status()
config_ini = avd_home / f"{AVD_NAME}.avd" / "config.ini"
print(f"===== API 37 boot diagnostics (attempt {attempt}) =====")
print("--- adb devices ---")
subprocess.run(cmd(adb, "devices"), check=False)
print(f"--- emulator -accel-check ---\n{accel}")
print(f"--- KVM/hypervisor ---\n{kvm}")
print(f"--- GPU mode ---\n{GPU_MODE}")
print(f"--- free memory ---\n{_mem_info()}")
print(f"--- free disk ---\n{_disk_info(Path(tempfile.gettempdir()))}")
print("--- AVD config.ini ---")
try:
print(config_ini.read_text(errors="replace"))
except OSError as exc:
print(f"(could not read {config_ini}: {exc})")
# Concise failure summary (mirrors CI): accel/KVM status + the last 50 lines of emulator.log.
print(f"----- BOOT FAILURE SUMMARY (attempt {attempt}) -----")
print(f"accel-check: {accel}")
print(f"kvm: {kvm}")
tail(emu_log, 50)
def main() -> int:
parser = argparse.ArgumentParser(
description="Hand-provision + run the API 37 preview E2E suite.")
@@ -216,8 +320,14 @@ def main() -> int:
os.environ["ANDROID_AVD_HOME"] = str(avd_home)
emu_log = Path(tempfile.gettempdir()) / "libremail-api37-emulator.log"
logcat_log = Path(tempfile.gettempdir()) / "libremail-api37-logcat.txt"
try:
logcat_log.unlink() # start fresh; start_logcat appends (with a header) per attempt
except OSError:
pass
adb: str | None = None
proc: subprocess.Popen | None = None
logcat_proc: subprocess.Popen | None = None
test_exit = 1
try:
@@ -234,16 +344,21 @@ def main() -> int:
# 2. Create the AVD, mirroring CI.
create_avd(avdmanager, emulator)
# 3. Cold-boot headless, retrying once (mirrors CI's two-attempt boot loop).
# 3. Cold-boot headless, retrying once (mirrors CI's two-attempt boot loop). Diagnostics
# (logcat capture + a boot-timeout dump) are ADDITIVE -- the retry/boot-wait is unchanged.
booted = False
for attempt in (1, 2):
proc = start_emulator(emulator, emu_log, attempt)
# Capture logcat from device registration onward (mirrors CI); killed on failure.
logcat_proc = start_logcat(adb, logcat_log, attempt)
if wait_for_boot(adb, proc, args.boot_timeout):
booted = True
break
print(f"API 37 emulator did not boot within {args.boot_timeout}s (attempt {attempt}).",
file=sys.stderr)
tail(emu_log)
dump_diagnostics(adb, emulator, emu_log, avd_home, attempt)
stop_logcat(logcat_proc)
logcat_proc = None
stop_emulator(adb, proc)
proc = None
time.sleep(5)
@@ -266,9 +381,12 @@ def main() -> int:
finally:
# 5. Always tear the emulator down and delete the AVD, even on failure.
print("Tearing down API 37 emulator and AVD...")
stop_logcat(logcat_proc)
stop_emulator(adb, proc)
subprocess.run(cmd(avdmanager, "delete", "avd", "-n", AVD_NAME), check=False,
stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
print(f"(emulator boot log: {emu_log})")
print(f"(logcat: {logcat_log})")
if test_exit != 0:
print(f"api37 connectedDebugAndroidTest failed (exit {test_exit}).", file=sys.stderr)
@@ -1,25 +1,28 @@
<!-- SPDX-License-Identifier: GPL-3.0-or-later -->
# `local_instrumented.sh` — reliable local instrumented/E2E runs
# `local_instrumented.py` — reliable local instrumented/E2E runs
A helper for running LibreMail's instrumented / E2E tests **locally** without Gradle
Managed Devices (GMD). Companion to `api37_e2e.py`; born from issue #269.
Managed Devices (GMD). Cross-platform, pure Python 3 standard library (Windows / Linux /
macOS). Companion to `api37_e2e.py`; born from issue #269, ported from bash to Python in
issue #281 so it runs the same on the Windows primary dev box and on \*nix — no Git Bash,
no `jq`, no `taskkill`-vs-`kill` gaps.
## Usage
```bash
# in Git Bash, from anywhere — invoke the script by path:
.claude/skills/preflight/local_instrumented.sh org.libremail.ui.compose.ComposeScreenE2ETest
# from anywhere — invoke the script by path (Windows: use `py` or `python`):
python .claude/skills/preflight/local_instrumented.py org.libremail.ui.compose.ComposeScreenE2ETest
# multiple classes (comma-separated, no spaces):
.claude/skills/preflight/local_instrumented.sh org.libremail.a.FooTest,org.libremail.b.BarTest
python .claude/skills/preflight/local_instrumented.py org.libremail.a.FooTest,org.libremail.b.BarTest
```
The script is CWD-independent: it resolves its own repo/worktree root from its script
location (three directories up from `.claude/skills/preflight`) and `cd`s there before
invoking gradlew, so it always builds *that* tree's `:app` — never whatever tree your
shell happens to be sitting in. This matters most when you have several worktrees
checked out side by side; run the copy of this script that lives inside the worktree you
want to test, regardless of your current directory (issue #284).
location (three directories up from `.claude/skills/preflight`) and runs gradlew there, so
it always builds *that* tree's `:app` — never whatever tree your shell happens to be
sitting in. This matters most when you have several worktrees checked out side by side; run
the copy of this script that lives inside the worktree you want to test, regardless of your
current directory (issue #284).
It cold-boots **one** emulator (`-no-snapshot`, no GMD), waits for `sys.boot_completed`,
runs `:app:connectedDebugAndroidTest` filtered to the class(es) you pass, then tears the
@@ -35,16 +38,25 @@ emulator down and verifies no orphaned `qemu` process is left behind (exit **3**
- **Keep runs targeted.** The full ~114-test suite tends to wedge mid-run on this machine;
small, targeted class sets do not. That's why the script requires an explicit class list —
run only what you changed. The full matrix is CI's job.
- **Emulator hygiene is mandatory.** A hung `adb emu kill` leaves a detached
`qemu-system-x86_64-headless.exe`; accumulated orphans have frozen this machine. The
script force-kills stragglers before booting and after tearing down, and fails loudly if
a zombie survives.
- **Emulator hygiene is mandatory.** A hung `adb emu kill` leaves a detached qemu VM
(`qemu-system-x86_64-headless.exe` on Windows, a `qemu-system-*` process on \*nix);
accumulated orphans have frozen this machine. The script force-kills stragglers before
booting and after tearing down, and fails loudly (exit 3) if a zombie survives. The
orphan-kill is abstracted per-OS (`taskkill /F /IM …` on Windows, `pkill -f qemu-system`
on \*nix), and teardown always runs — even on Ctrl-C / error / SIGTERM (try/finally +
atexit + SIGINT/SIGTERM handlers).
See the header comment of `local_instrumented.sh` for the full rationale, requirements, and
the `LOCAL_INSTRUMENTED_*` environment overrides (AVD name, JDK home, boot timeout, …).
Exit codes: **0** pass · **2** usage/precondition failure · **3** a qemu zombie survived
teardown · **4** emulator never booted · any other non-zero = `connectedDebugAndroidTest`'s
own test-failure exit code.
See the module docstring at the top of `local_instrumented.py` for the full rationale,
requirements, and the `LOCAL_INSTRUMENTED_*` environment overrides (AVD name, JDK home,
boot timeout, …).
## Requirements
Git Bash; Android SDK `emulator` + `adb` on `PATH`; a JDK **17–21** (AGP 9.2 fails on 25+ —
the script pins `JAVA_HOME` to a known JDK 21, overridable via `LOCAL_INSTRUMENTED_JDK`); and
a free hardware hypervisor (shut down VirtualBox / other VMs first).
`python3` (Windows: `py`/`python`); Android SDK `emulator` + `adb` on `PATH`; a JDK
**17–21** (AGP 9.2 fails on 25+ — the script pins `JAVA_HOME` to a known JDK 21, overridable
via `LOCAL_INSTRUMENTED_JDK`); and a free hardware hypervisor (shut down VirtualBox / other
VMs first).
@@ -0,0 +1,451 @@
#!/usr/bin/env python3
# SPDX-License-Identifier: GPL-3.0-or-later
"""local_instrumented.py -- reliable LOCAL instrumented / E2E test runner for LibreMail.
Usage: local_instrumented.py <fully.qualified.TestClass>[,<Class2>,...]
Example:
python .claude/skills/preflight/local_instrumented.py \
org.libremail.ui.compose.ComposeScreenE2ETest
python .claude/skills/preflight/local_instrumented.py \
org.libremail.ui.compose.ComposeScreenE2ETest,org.libremail.ui.compose.RecipientChipTest
Cross-platform (Windows / Linux / macOS), pure standard library. Companion to
``api37_e2e.py``; ported from the original ``local_instrumented.sh`` (issue #281) so the
helper runs the same on the Windows primary dev box and on *nix -- no Git Bash, no ``jq``,
no ``taskkill`` vs ``kill`` portability gaps.
WHY THIS SCRIPT EXISTS (issue #269)
------------------------------------
On this machine (Windows + the AEHD 2.2 hypervisor) the Gradle Managed Device (GMD)
instrumented tasks -- ``apiXXDebugAndroidTest`` -- FAIL during setup. GMD tries to
save/load an AVD *snapshot* and AEHD 2.2 cannot complete it:
AvdSnapshotHandler$EmulatorSnapshotCannotCreatedException: Snapshot creation timed out
GMD retries the snapshot ~5x, rebooting the AVD each time -- that endless reboot is the
"cycling" that eats hours. The emulator ITSELF is healthy (8 GB RAM, sys.boot_completed=1,
shell-responsive); only GMD's snapshot step is broken. So every LOCAL GMD task is affected:
the coverage lanes and the /preflight api35/api36 steps. CI is unaffected -- it uses
reactivecircus/android-emulator-runner + ``connectedDebugAndroidTest``, never GMD.
THE RELIABLE LOCAL PATH (this script):
Cold-boot ONE emulator by hand with ``-no-snapshot`` (no GMD, no snapshot machinery),
then run ``:app:connectedDebugAndroidTest`` -- the exact technique CI and ``api37_e2e.py``
already use. We reuse a GMD-provisioned AVD by name so we don't re-download a system
image; GMD re-provisions its own copy on its next run, so the ``-wipe-data`` cold boot
here does not disturb it.
KEEP RUNS TARGETED -- THE ~114-TEST MID-SUITE WEDGE
---------------------------------------------------
Running the WHOLE instrumented suite (~114 tests) via ``connectedDebugAndroidTest`` on this
box tends to wedge partway through -- the emulator stops making progress mid-run. Small,
targeted class sets do NOT hit that wedge. That is why this helper takes an explicit
``<fully.qualified.TestClass>[,...]`` argument and filters the run with
``-Pandroid.testInstrumentationRunnerArguments.class=...`` instead of running everything.
Run the class(es) you actually changed; do not use this to run the full suite (that is
CI's / preflight's job across the API matrix).
FREEZE / HYGIENE RATIONALE -- WHY THE ORPHAN-KILL + TEARDOWN VERIFY ARE MANDATORY
--------------------------------------------------------------------------------
A hung ``adb emu kill`` (or an interrupted run) leaves a detached qemu VM process behind
(``qemu-system-x86_64-headless.exe`` on Windows; a ``qemu-system-*`` process on *nix).
These orphans do not show up in ``adb devices``, they keep holding the hypervisor + RAM,
and accumulated orphans have FROZEN this machine outright. So this script:
* PREAMBLE -- force-kills any pre-existing qemu/emulator processes and resets the adb
server BEFORE booting, so we always start from a clean slate.
* TEARDOWN -- ``adb emu kill``, kill the launcher we spawned, then re-check for ANY
surviving emulator/qemu process and force-kill it (the ``-no-window`` emulator
can leave a sibling ``emulator.exe`` that briefly outlives the qemu VM). Teardown
runs even on Ctrl-C / error / SIGTERM (try/finally + atexit + SIGINT/SIGTERM
handlers) and is idempotent.
* VERIFY -- if a qemu process is STILL alive after the force-kill, the script exits
non-zero (code 3) so the leak is never silently ignored.
Never leave an emulator running after this script; if it exits 3, hunt the zombie down by
hand (Windows: ``tasklist | findstr qemu`` then ``taskkill /F /IM
qemu-system-x86_64-headless.exe``; *nix: ``pgrep -fa qemu-system`` then ``pkill -f
qemu-system``).
CROSS-PLATFORM PROCESS KILL
---------------------------
Listing and force-killing the emulator/qemu processes is abstracted per-OS (see
``list_procs`` / ``force_kill``): Windows uses ``tasklist`` + ``taskkill /F /IM <image>``;
*nix uses ``ps ax`` + ``pkill -f qemu-system`` (alongside the graceful ``adb emu kill``).
The qemu VM is the freeze-causing orphan on every platform.
EXIT CODES (preserved from local_instrumented.sh)
0 tests passed
2 usage / precondition failure
3 a qemu zombie survived teardown -- clean it up by hand before the next run
4 emulator never reached sys.boot_completed
<n> connectedDebugAndroidTest's own non-zero exit code (test failures)
REQUIREMENTS
* Android SDK ``emulator`` + ``adb`` on PATH.
* A JDK 17-21 for the Gradle daemon -- AGP 9.2 fails on JDK 25+. This script pins
JAVA_HOME to a known JDK 21 (override with LOCAL_INSTRUMENTED_JDK) because the ambient
JAVA_HOME on the primary box points at JDK 25.
* A free hardware hypervisor (VT-x/WHPX/AEHD/KVM/HVF). Shut down VirtualBox / other VMs
first or the AVD hangs at 0% CPU and never reaches sys.boot_completed.
Overridable via environment (defaults target the primary Windows dev box):
LOCAL_INSTRUMENTED_AVD AVD name to boot (dev36_google_apis_x86_64_Pixel_2)
ANDROID_AVD_HOME AVD home dir (C:/Users/jasonross/.android/avd/gradle-managed)
LOCAL_INSTRUMENTED_JDK JDK 17-21 home (Eclipse Adoptium jdk-21.0.11.10-hotspot)
LOCAL_INSTRUMENTED_SERIAL adb serial (emulator-5554)
LOCAL_INSTRUMENTED_BOOT_TIMEOUT boot wait seconds (300)
"""
from __future__ import annotations
import argparse
import atexit
import os
import shutil
import signal
import subprocess
import sys
import tempfile
import time
from pathlib import Path
IS_WINDOWS = os.name == "nt"
# ---- configuration (env-overridable; defaults are correct for the primary dev box) ------
AVD_NAME = os.environ.get("LOCAL_INSTRUMENTED_AVD", "dev36_google_apis_x86_64_Pixel_2")
AVD_HOME = os.environ.get("ANDROID_AVD_HOME", "C:/Users/jasonross/.android/avd/gradle-managed")
JDK_HOME = os.environ.get(
"LOCAL_INSTRUMENTED_JDK", "C:/Program Files/Eclipse Adoptium/jdk-21.0.11.10-hotspot"
)
SERIAL = os.environ.get("LOCAL_INSTRUMENTED_SERIAL", "emulator-5554")
BOOT_TIMEOUT = int(os.environ.get("LOCAL_INSTRUMENTED_BOOT_TIMEOUT", "300"))
# Windows qemu/emulator image names (see FREEZE / HYGIENE above). The ``-headless`` variant is
# what a ``-no-window`` emulator launches; the plain qemu name is swept too, belt-and-suspenders.
QEMU_IMAGE = "qemu-system-x86_64-headless.exe"
QEMU_IMAGE_ALT = "qemu-system-x86_64.exe"
EMULATOR_IMAGE = "emulator.exe"
REPO_ROOT = Path(__file__).resolve().parents[3] # .claude/skills/preflight -> repo root
GRADLEW = REPO_ROOT / ("gradlew.bat" if IS_WINDOWS else "gradlew")
EMU_LOG = Path(tempfile.gettempdir()) / "libremail-local-instrumented-emulator.log"
class _RunState:
"""Mutable run state shared by main(), teardown(), the atexit hook and the signal
handlers -- mirrors the bash globals EMU_PID / TEST_EXIT / ZOMBIE / TEARDOWN_DONE."""
def __init__(self) -> None:
self.proc: subprocess.Popen | None = None
self.test_exit = 1
self.zombie = False
self.teardown_done = False
_STATE = _RunState()
def log(msg: str) -> None:
print(f"\n=== {msg} ===")
def warn(msg: str) -> None:
print(f"WARNING: {msg}", file=sys.stderr)
def die(msg: str) -> None:
"""Print an error and exit 2 (usage / precondition failure). Called before the teardown
backstops are armed, so nothing has booted and there is nothing to tear down."""
print(f"ERROR: {msg}", file=sys.stderr)
sys.exit(2)
def cmd(tool: str, *args: str) -> list[str]:
"""Build an argv list, wrapping Windows ``.bat``/``.cmd`` launchers (e.g. gradlew.bat)
through ``cmd /c`` -- matching api37_e2e.py. ``.exe`` tools pass through unchanged."""
if IS_WINDOWS and tool.lower().endswith((".bat", ".cmd")):
return ["cmd", "/c", tool, *args]
return [tool, *args]
def _run_quiet(argv: list[str]) -> None:
"""Run a command, discarding output and swallowing any error -- teardown/kill helpers
must always make progress."""
try:
subprocess.run(argv, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL, check=False)
except (OSError, subprocess.SubprocessError):
pass
def _run_capture(argv: list[str]) -> str:
try:
return subprocess.run(argv, capture_output=True, text=True, check=False).stdout or ""
except (OSError, subprocess.SubprocessError):
return ""
def list_procs(*needles: str) -> str:
"""Return the lines of currently-running processes whose name/command line contains any
of ``needles`` (case-insensitive); empty string if none. Cross-platform stand-in for the
.sh's ``tasklist | grep``: ``tasklist`` on Windows, ``ps ax`` on *nix."""
out = _run_capture(["tasklist"] if IS_WINDOWS else ["ps", "ax"])
lowered = [n.lower() for n in needles]
return "\n".join(ln for ln in out.splitlines() if any(n in ln.lower() for n in lowered))
def list_qemu() -> str:
return list_procs("qemu")
def list_emu_procs() -> str:
return list_procs("qemu", "emulator")
def force_kill(win_images: list[str], nix_patterns: list[str]) -> None:
"""Best-effort force-kill. Windows: ``taskkill /F /IM <image> ...``. *nix: ``pkill -f
<pattern>`` per pattern. Never raises -- teardown must always make progress."""
if IS_WINDOWS:
argv = ["taskkill", "/F"]
for image in win_images:
argv += ["/IM", image]
_run_quiet(argv)
else:
for pattern in nix_patterns:
_run_quiet(["pkill", "-f", pattern])
def tail(path: Path, lines: int = 40) -> None:
try:
with open(path, "r", errors="replace") as handle:
content = handle.readlines()[-lines:]
print("".join(content), file=sys.stderr, end="")
except OSError:
pass
def teardown() -> None:
"""Kill the emulator and verify no orphaned emulator/qemu process remains. Idempotent --
safe to call from the finally block, the atexit hook and the signal handlers (mirrors the
.sh TEARDOWN_DONE guard). Sets _STATE.zombie if a *qemu* process survives the force-kill --
the machine-freezing case (exit 3)."""
if _STATE.teardown_done:
return
_STATE.teardown_done = True
log("Teardown: killing emulator and verifying no orphaned emulator/qemu remains")
adb = shutil.which("adb")
if adb:
_run_quiet(cmd(adb, "-s", SERIAL, "emu", "kill"))
time.sleep(2)
# Belt-and-suspenders: kill the emulator launcher process we started, if still alive.
proc = _STATE.proc
if proc is not None and proc.poll() is None:
proc.terminate()
try:
proc.wait(timeout=1)
except subprocess.TimeoutExpired:
proc.kill()
# Reap any lingering emulator/qemu process, then verify. Unlike the original .sh -- which
# swept qemu ONLY -- we also force-kill the emulator *launcher* image: on Windows the
# ``-no-window`` emulator spawns a sibling ``emulator.exe`` that is NOT the Popen child we
# tracked and outlives both it and the qemu VM by a few seconds, so a qemu-only sweep
# returns while it is still shutting down -- an orphan the freeze-safety rule forbids. So we
# trigger on any emulator-or-qemu survivor and taskkill the launcher too.
if list_emu_procs():
warn("emulator/qemu still present after 'adb emu kill'; force-killing:")
print(list_emu_procs(), file=sys.stderr)
force_kill([QEMU_IMAGE, QEMU_IMAGE_ALT, EMULATOR_IMAGE], ["qemu-system"])
time.sleep(2)
# A surviving QEMU is the machine-freezing zombie (exit 3); a stray launcher is not.
remaining = list_qemu()
if remaining:
warn("qemu ZOMBIE survived teardown -- kill it by hand or the machine may freeze:")
print(remaining, file=sys.stderr)
_STATE.zombie = True
if adb:
_run_quiet(cmd(adb, "kill-server"))
def orphan_kill_preamble(adb: str) -> None:
"""Force-kill any pre-existing qemu/emulator processes and reset the adb server, so we
always cold-boot from a clean slate."""
log("Orphan-kill preamble: ensuring a clean slate before boot")
existing = list_emu_procs()
if existing:
warn("Pre-existing emulator/qemu processes found -- force-killing them first:")
print(existing, file=sys.stderr)
force_kill([QEMU_IMAGE, EMULATOR_IMAGE], ["qemu-system"])
time.sleep(2)
else:
print("No pre-existing qemu/emulator processes.")
_run_quiet(cmd(adb, "kill-server"))
_run_quiet(cmd(adb, "start-server"))
def start_emulator(emulator: str) -> subprocess.Popen:
"""Cold-boot ONE emulator by hand (no GMD, no snapshot), logging to EMU_LOG. Records the
launcher process in _STATE so teardown can reap it even if we are interrupted next."""
log(f"Cold-booting @{AVD_NAME} (no GMD, no snapshot); log -> {EMU_LOG}")
flags = [
f"@{AVD_NAME}",
"-no-window", "-no-snapshot", "-no-boot-anim", "-no-audio",
"-gpu", "auto-no-window", "-cores", "8", "-wipe-data",
]
logf = open(EMU_LOG, "wb") # noqa: SIM115 - handed to the child; parent copy closed below
try:
proc = subprocess.Popen(cmd(emulator, *flags), stdout=logf, stderr=subprocess.STDOUT)
finally:
logf.close() # the child inherited its own fd; the parent's copy is no longer needed
_STATE.proc = proc
print(f"emulator launcher pid={proc.pid}")
return proc
def wait_for_boot(adb: str, proc: subprocess.Popen, timeout: int) -> bool:
"""Poll ``adb get-state`` + ``getprop sys.boot_completed`` until the emulator is up, or the
launcher dies, or ``timeout`` seconds elapse. Mirrors the .sh boot loop."""
print(f"Waiting up to {timeout}s for sys.boot_completed on {SERIAL}...")
deadline = time.monotonic() + timeout
while time.monotonic() < deadline:
if proc.poll() is not None:
warn("emulator process exited during boot; last log lines:")
tail(EMU_LOG, 40)
return False
state = _run_capture(cmd(adb, "-s", SERIAL, "get-state")).strip()
if state == "device":
booted = _run_capture(
cmd(adb, "-s", SERIAL, "shell", "getprop", "sys.boot_completed")
).strip()
if booted == "1":
return True
time.sleep(3)
return False
def dismiss_keyguard(adb: str) -> None:
# Dismiss the keyguard (mirrors CI + api37_e2e.py). Best-effort: a cold -wipe-data boot
# rarely needs it, and the input service can lose a race right after boot.
_run_quiet(cmd(adb, "-s", SERIAL, "shell", "input", "keyevent", "82"))
def run_tests(test_classes: str) -> int:
"""Run :app:connectedDebugAndroidTest filtered to ``test_classes`` from the repo root
(JAVA_HOME / ANDROID_AVD_HOME are already in the environment)."""
log(f"Running :app:connectedDebugAndroidTest for: {test_classes}")
print(f"JAVA_HOME={os.environ.get('JAVA_HOME', '')}")
return subprocess.run(
cmd(
str(GRADLEW),
":app:connectedDebugAndroidTest",
f"-Pandroid.testInstrumentationRunnerArguments.class={test_classes}",
"--stacktrace",
),
cwd=str(REPO_ROOT),
check=False,
).returncode
def main() -> int:
parser = argparse.ArgumentParser(
prog="local_instrumented.py",
formatter_class=argparse.RawDescriptionHelpFormatter,
description=(
"Cold-boot ONE emulator (no GMD, no snapshot) and run "
":app:connectedDebugAndroidTest filtered to the given instrumented test class(es)."
),
epilog=(
"Keep the class set small and targeted -- the full ~114-test suite tends to wedge\n"
"mid-run on this box (see the module docstring). The full matrix is CI's job.\n"
"Example:\n"
" python .claude/skills/preflight/local_instrumented.py \\\n"
" org.libremail.ui.compose.ComposeScreenE2ETest,org.libremail.ui.compose.RecipientChipTest"
),
)
parser.add_argument(
"test_classes",
metavar="TEST_CLASSES",
help=(
"Comma-separated fully-qualified instrumented test class(es), no spaces "
"(e.g. org.libremail.a.FooTest,org.libremail.b.BarTest)."
),
)
args = parser.parse_args()
# ---- preconditions (before arming teardown; nothing has booted yet) ------------------
emulator = shutil.which("emulator")
adb = shutil.which("adb")
if not emulator:
die("emulator not on PATH (install Android SDK emulator).")
if not adb:
die("adb not on PATH (install Android SDK platform-tools).")
kill_tool = "taskkill" if IS_WINDOWS else "pkill"
if not shutil.which(kill_tool):
die(f"{kill_tool} not found -- required to force-kill orphaned emulator/qemu processes.")
if not GRADLEW.is_file():
die(f"gradlew not found at {GRADLEW}.")
if not os.path.isdir(JDK_HOME):
die(f"JDK 17-21 not found at '{JDK_HOME}'. Set LOCAL_INSTRUMENTED_JDK.")
if not os.path.isfile(os.path.join(AVD_HOME, AVD_NAME + ".ini")):
die(
f"AVD '{AVD_NAME}' not found under '{AVD_HOME}'. "
"Set LOCAL_INSTRUMENTED_AVD / ANDROID_AVD_HOME. "
"(GMD AVDs are created by any local apiXXDebugAndroidTest run.)"
)
# Move gradlew's working dir to this tree's repo root (below) and pin JAVA_HOME/AVD home,
# exactly like the .sh -- the ambient JAVA_HOME on this box points at JDK 25 (AGP-incompatible).
os.environ["JAVA_HOME"] = JDK_HOME
os.environ["ANDROID_AVD_HOME"] = AVD_HOME
# ---- arm teardown backstops BEFORE touching the emulator -----------------------------
# try/finally is the primary path; atexit covers sys.exit()/unhandled-exception exits; the
# signal handlers make SIGINT/SIGTERM tear down too (Python does not raise on SIGTERM by
# default). teardown() is idempotent, so firing from several paths is safe (mirrors the
# .sh's ``trap teardown EXIT INT TERM`` + TEARDOWN_DONE guard).
atexit.register(teardown)
def _signal_teardown(signum: int, _frame: object) -> None:
teardown()
sys.exit(128 + signum)
signal.signal(signal.SIGINT, _signal_teardown)
if hasattr(signal, "SIGTERM"):
signal.signal(signal.SIGTERM, _signal_teardown)
boot_failed = False
try:
orphan_kill_preamble(adb)
proc = start_emulator(emulator)
if wait_for_boot(adb, proc, BOOT_TIMEOUT):
print("Emulator booted.")
dismiss_keyguard(adb)
_STATE.test_exit = run_tests(args.test_classes)
else:
warn(f"Emulator did not reach sys.boot_completed within {BOOT_TIMEOUT}s.")
tail(EMU_LOG, 40)
boot_failed = True
finally:
teardown()
if boot_failed:
return 4
if _STATE.zombie:
warn(
"Exiting 3: a qemu zombie was left behind (see above) -- "
"clean it up before the next run."
)
return 3
if _STATE.test_exit != 0:
warn(
f"connectedDebugAndroidTest failed (exit {_STATE.test_exit}). "
"Report: app/build/reports/androidTests/connected/"
)
return _STATE.test_exit
log(f"PASS -- instrumented tests green for: {args.test_classes}")
return 0
if __name__ == "__main__":
sys.exit(main())
@@ -1,253 +0,0 @@
#!/usr/bin/env bash
# SPDX-License-Identifier: GPL-3.0-or-later
#
# local_instrumented.sh — reliable LOCAL instrumented / E2E test runner for LibreMail.
#
# Usage: local_instrumented.sh <fully.qualified.TestClass>[,<Class2>,...]
# Example:
# .claude/skills/preflight/local_instrumented.sh \
# org.libremail.ui.compose.ComposeScreenE2ETest
# .claude/skills/preflight/local_instrumented.sh \
# org.libremail.ui.compose.ComposeScreenE2ETest,org.libremail.ui.compose.RecipientChipTest
#
# =============================================================================
# WHY THIS SCRIPT EXISTS (issue #269)
# -----------------------------------------------------------------------------
# On this machine (Windows + the AEHD 2.2 hypervisor) the Gradle Managed Device
# (GMD) instrumented tasks — `apiXXDebugAndroidTest` — FAIL during setup. GMD tries
# to save/load an AVD *snapshot* and AEHD 2.2 cannot complete it:
#
# AvdSnapshotHandler$EmulatorSnapshotCannotCreatedException: Snapshot creation timed out
#
# GMD retries the snapshot ~5x, rebooting the AVD each time — that endless reboot is
# the "cycling" that eats hours. The emulator ITSELF is healthy (8 GB RAM,
# `sys.boot_completed=1`, shell-responsive); only GMD's snapshot step is broken. So
# every LOCAL GMD task is affected: coverage lanes 3/5 (#248/#250) and the /preflight
# api35/api36 steps (#266). CI is unaffected — it uses reactivecircus/android-emulator-runner
# + `connectedDebugAndroidTest`, never GMD.
#
# THE RELIABLE LOCAL PATH (this script):
# Cold-boot ONE emulator by hand with `-no-snapshot` (no GMD, no snapshot machinery),
# then run `:app:connectedDebugAndroidTest` — the exact technique CI and
# `api37_e2e.py` already use. We reuse a GMD-provisioned AVD by name so we don't have
# to re-download a system image; GMD re-provisions its own copy on its next run, so
# the `-wipe-data` cold boot here does not disturb it.
#
# =============================================================================
# KEEP RUNS TARGETED — THE ~114-TEST MID-SUITE WEDGE
# -----------------------------------------------------------------------------
# Running the WHOLE instrumented suite (~114 tests) via `connectedDebugAndroidTest`
# on this box tends to wedge partway through — the emulator stops making progress
# mid-run. Small, targeted class sets do NOT hit that wedge. That is why this helper
# takes an explicit `<fully.qualified.TestClass>[,...]` argument and filters the run
# with `-Pandroid.testInstrumentationRunnerArguments.class=...` instead of running
# everything. Run the class(es) you actually changed; do not use this to run the full
# suite (that is CI's / preflight's job across the API matrix).
#
# =============================================================================
# FREEZE / HYGIENE RATIONALE — WHY THE ORPHAN-KILL + TEARDOWN VERIFY ARE MANDATORY
# -----------------------------------------------------------------------------
# A hung `adb emu kill` (or an interrupted run) leaves a detached
# `qemu-system-x86_64-headless.exe` behind. These orphans do not show up in
# `adb devices`, they keep holding the hypervisor + RAM, and accumulated orphans have
# FROZEN this machine outright. So this script:
# * PREAMBLE — force-kills any pre-existing qemu/emulator processes and resets the
# adb server BEFORE booting, so we always start from a clean slate.
# * TEARDOWN — `adb emu kill`, then re-checks `tasklist` for qemu and force-kills any
# survivor. The teardown runs even on Ctrl-C / error (EXIT/INT/TERM trap).
# * VERIFY — if a qemu process is STILL alive after the force-kill, the script exits
# non-zero (code 3) so the leak is never silently ignored.
# Never leave an emulator running after this script; if it exits 3, hunt the zombie
# down by hand (`tasklist | grep -i qemu`; `taskkill //F //IM qemu-system-x86_64-headless.exe`).
#
# =============================================================================
# REQUIREMENTS
# * Git Bash (this is a bash script; it shells out to Windows `tasklist`/`taskkill`).
# * Android SDK `emulator` + `adb` on PATH (SDK at C:\Android here).
# * A JDK 17–21 for the Gradle daemon — AGP 9.2 fails on JDK 25+. This script pins
# JAVA_HOME to a known JDK 21 (override with LOCAL_INSTRUMENTED_JDK) because the
# ambient JAVA_HOME on this box points at JDK 25.
# * A free hardware hypervisor (VT-x/WHPX/AEHD). Shut down VirtualBox / other VMs first
# or the AVD hangs at 0% CPU and never reaches sys.boot_completed.
#
# Overridable via environment (defaults target THIS machine):
# LOCAL_INSTRUMENTED_AVD AVD name to boot (dev36_google_apis_x86_64_Pixel_2)
# ANDROID_AVD_HOME AVD home dir (C:/Users/jasonross/.android/avd/gradle-managed)
# LOCAL_INSTRUMENTED_JDK JDK 17–21 home (Eclipse Adoptium jdk-21.0.11.10-hotspot)
# LOCAL_INSTRUMENTED_SERIAL adb serial (emulator-5554)
# LOCAL_INSTRUMENTED_BOOT_TIMEOUT boot wait seconds (300)
# =============================================================================
set -uo pipefail
# ---- configuration (env-overridable; defaults are correct for this machine) -----------
AVD_NAME="${LOCAL_INSTRUMENTED_AVD:-dev36_google_apis_x86_64_Pixel_2}"
AVD_HOME="${ANDROID_AVD_HOME:-C:/Users/jasonross/.android/avd/gradle-managed}"
JDK_HOME="${LOCAL_INSTRUMENTED_JDK:-C:/Program Files/Eclipse Adoptium/jdk-21.0.11.10-hotspot}"
SERIAL="${LOCAL_INSTRUMENTED_SERIAL:-emulator-5554}"
BOOT_TIMEOUT="${LOCAL_INSTRUMENTED_BOOT_TIMEOUT:-300}"
QEMU_IMAGE="qemu-system-x86_64-headless.exe"
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
REPO_ROOT="$(cd "${SCRIPT_DIR}/../../.." && pwd)" # .claude/skills/preflight -> repo root
GRADLEW="${REPO_ROOT}/gradlew"
EMU_LOG="${TMPDIR:-/tmp}/libremail-local-instrumented-emulator.log"
EMU_PID=""
TEST_EXIT=1
ZOMBIE=0
TEARDOWN_DONE=0
log() { printf '\n=== %s ===\n' "$*"; }
warn() { printf 'WARNING: %s\n' "$*" >&2; }
die() { printf 'ERROR: %s\n' "$*" >&2; exit 2; }
# All emulator/qemu processes Windows currently sees (empty string if none).
list_emu_procs() { tasklist 2>/dev/null | grep -iE 'qemu|emulator' || true; }
list_qemu() { tasklist 2>/dev/null | grep -i 'qemu' || true; }
# ---- argument parsing -----------------------------------------------------------------
TEST_CLASSES="${1:-}"
if [[ -z "${TEST_CLASSES}" ]]; then
cat >&2 <<'USAGE'
usage: local_instrumented.sh <fully.qualified.TestClass>[,<Class2>,...]
Cold-boots ONE emulator (no GMD, no snapshot) and runs :app:connectedDebugAndroidTest
filtered to the given instrumented test class(es). Keep the set small and targeted —
see the header for the ~114-test mid-suite wedge.
USAGE
exit 2
fi
# ---- preconditions --------------------------------------------------------------------
command -v emulator >/dev/null 2>&1 || die "emulator not on PATH (install Android SDK emulator)."
command -v adb >/dev/null 2>&1 || die "adb not on PATH (install Android SDK platform-tools)."
command -v tasklist >/dev/null 2>&1 || die "tasklist not found — this helper targets Windows/Git Bash."
[[ -f "${GRADLEW}" ]] || die "gradlew not found at ${GRADLEW}."
[[ -d "${JDK_HOME}" ]] || die "JDK 17-21 not found at '${JDK_HOME}'. Set LOCAL_INSTRUMENTED_JDK."
[[ -f "${AVD_HOME}/${AVD_NAME}.ini" ]] || \
die "AVD '${AVD_NAME}' not found under '${AVD_HOME}'. Set LOCAL_INSTRUMENTED_AVD / ANDROID_AVD_HOME.
(GMD AVDs are created by any local apiXXDebugAndroidTest run.)"
# Move into the resolved repo/worktree root before invoking gradlew. GRADLEW above is an
# absolute path, but the gradlew wrapper script picks the *project* to build from the
# process's current directory, not from its own script location — so without this `cd`,
# running this helper from a different tree (e.g. another worktree, or the main repo
# while iterating on a worktree's copy of this script) silently builds the CALLER's CWD
# tree instead of this one (issue #284).
cd "${REPO_ROOT}" || die "Could not cd to repo root '${REPO_ROOT}'."
export JAVA_HOME="${JDK_HOME}"
export ANDROID_AVD_HOME="${AVD_HOME}"
# ---- teardown: always runs (normal exit, error, or Ctrl-C) ----------------------------
teardown() {
[[ "${TEARDOWN_DONE}" == "1" ]] && return 0
TEARDOWN_DONE=1
log "Teardown: killing emulator and verifying no orphaned qemu remains"
adb -s "${SERIAL}" emu kill >/dev/null 2>&1 || true
sleep 2
# Belt-and-suspenders: kill the emulator launcher process we started, if still alive.
if [[ -n "${EMU_PID}" ]] && kill -0 "${EMU_PID}" 2>/dev/null; then
kill "${EMU_PID}" 2>/dev/null || true
sleep 1
kill -9 "${EMU_PID}" 2>/dev/null || true
fi
# Verify: any surviving qemu is a machine-freezing zombie — force-kill and re-check.
local remaining
remaining="$(list_qemu)"
if [[ -n "${remaining}" ]]; then
warn "qemu still present after 'adb emu kill'; force-killing:"
printf '%s\n' "${remaining}" >&2
taskkill //F //IM "${QEMU_IMAGE}" >/dev/null 2>&1 || true
# Sweep any other stray qemu-system image name, too.
taskkill //F //IM "qemu-system-x86_64.exe" >/dev/null 2>&1 || true
sleep 2
remaining="$(list_qemu)"
if [[ -n "${remaining}" ]]; then
warn "qemu ZOMBIE survived teardown — kill it by hand or the machine may freeze:"
printf '%s\n' "${remaining}" >&2
ZOMBIE=1
fi
fi
adb kill-server >/dev/null 2>&1 || true
}
trap teardown EXIT INT TERM
# ---- 1. orphan-kill preamble ----------------------------------------------------------
log "Orphan-kill preamble: ensuring a clean slate before boot"
existing="$(list_emu_procs)"
if [[ -n "${existing}" ]]; then
warn "Pre-existing emulator/qemu processes found — force-killing them first:"
printf '%s\n' "${existing}" >&2
taskkill //F //IM "${QEMU_IMAGE}" //IM "emulator.exe" >/dev/null 2>&1 || true
sleep 2
else
echo "No pre-existing qemu/emulator processes."
fi
adb kill-server >/dev/null 2>&1 || true
adb start-server >/dev/null 2>&1 || true
# ---- 2. cold-boot ONE emulator (no snapshot) ------------------------------------------
log "Cold-booting @${AVD_NAME} (no GMD, no snapshot); log -> ${EMU_LOG}"
emulator "@${AVD_NAME}" \
-no-window -no-snapshot -no-boot-anim -no-audio \
-gpu auto-no-window -cores 8 -wipe-data \
>"${EMU_LOG}" 2>&1 &
EMU_PID=$!
echo "emulator launcher pid=${EMU_PID}"
echo "Waiting up to ${BOOT_TIMEOUT}s for sys.boot_completed on ${SERIAL}..."
deadline=$(( $(date +%s) + BOOT_TIMEOUT ))
booted=0
while (( $(date +%s) < deadline )); do
if ! kill -0 "${EMU_PID}" 2>/dev/null; then
warn "emulator process exited during boot; last log lines:"
tail -n 40 "${EMU_LOG}" >&2 || true
break
fi
state="$(adb -s "${SERIAL}" get-state 2>/dev/null | tr -d '\r')"
if [[ "${state}" == "device" ]]; then
bc="$(adb -s "${SERIAL}" shell getprop sys.boot_completed 2>/dev/null | tr -d '\r\n ')"
if [[ "${bc}" == "1" ]]; then booted=1; break; fi
fi
sleep 3
done
if [[ "${booted}" != "1" ]]; then
warn "Emulator did not reach sys.boot_completed within ${BOOT_TIMEOUT}s."
tail -n 40 "${EMU_LOG}" >&2 || true
# teardown runs via the EXIT trap; surface a boot failure distinctly.
exit 4
fi
echo "Emulator booted."
# Dismiss the keyguard (mirrors CI + api37_e2e.py). Best-effort: a cold -wipe-data boot
# rarely needs it, and the input service can lose a race right after boot.
adb -s "${SERIAL}" shell input keyevent 82 >/dev/null 2>&1 || true
# ---- 3. run the targeted instrumented tests -------------------------------------------
log "Running :app:connectedDebugAndroidTest for: ${TEST_CLASSES}"
echo "JAVA_HOME=${JAVA_HOME}"
"${GRADLEW}" :app:connectedDebugAndroidTest \
"-Pandroid.testInstrumentationRunnerArguments.class=${TEST_CLASSES}" \
--stacktrace
TEST_EXIT=$?
# ---- 4. teardown + verify, then exit --------------------------------------------------
teardown
if (( ZOMBIE != 0 )); then
warn "Exiting 3: a qemu zombie was left behind (see above) — clean it up before the next run."
exit 3
fi
if (( TEST_EXIT != 0 )); then
warn "connectedDebugAndroidTest failed (exit ${TEST_EXIT}). Report: app/build/reports/androidTests/connected/"
exit "${TEST_EXIT}"
fi
log "PASS — instrumented tests green for: ${TEST_CLASSES}"
exit 0
+306
View File
@@ -0,0 +1,306 @@
#!/usr/bin/env python3
# SPDX-License-Identifier: GPL-3.0-or-later
"""Hardened Android SDK setup for CI (issue #389).
The dominant merge-blocking flake was the **Set up Android SDK** step
(`android-actions/setup-android`) dying *before* the emulator ever starts:
Wrong version in preinstalled sdkmanager
Warning: ... preparing SDK package Android Emulator: Error reading Zip
content from a SeekableByteChannel.
Error: The process '.../sdkmanager' failed with exit code 1
Two root causes, both a corrupt/truncated download that a bare `sdkmanager`
turns into an un-retried exit 1:
* the action's own **unverified** cmdline-tools re-download (its default
cmdline-tools version rarely matches the runner image's preinstalled one, so
it logs "Wrong version in preinstalled sdkmanager" and re-fetches with *no*
checksum), and
* the action's default ``packages: tools platform-tools`` install (the "SDK
Tools" corrupt zip seen on a #388 preview shard) plus the emulator/platform
package installs.
This module hardens both with **verify -> reject -> retry**, never trusting
sdkmanager's exit code alone:
``bootstrap`` Download the *pinned* Android command-line tools zip, verify it
against a pinned size + SHA-256, and install it to
``$ANDROID_SDK_ROOT/cmdline-tools/<rev>`` -- the exact path
setup-android probes first, so the action reuses our verified
tree and never does its own unverified "Wrong version"
re-download. A size/hash mismatch (corrupt OR wrong version)
=> delete the bad zip + any half-extracted dir => re-download
clean. Only a verified tree is ever left in place, so the
success-gated cache can never bake in a corrupt SDK.
``install`` Run ``sdkmanager --install <packages>`` with retry + backoff.
"Error reading Zip content from a SeekableByteChannel" is a
corrupt package zip, so on failure each requested package's dir
(and sdkmanager's temp/intermediate dirs) is PURGED before the
retry -- forcing a fresh re-download instead of a re-read of the
corrupt file.
stdlib only (urllib/hashlib/zipfile/...), cross-platform, per the repo's "prefer
Python for dev/CI-helper scripts" rule. The pure helpers are unit-tested in
``test_setup_android_sdk.py`` (run by the ``traffic-control-tests`` job); the
full download/install path is validated by CI itself.
"""
from __future__ import annotations
import argparse
import hashlib
import os
import shutil
import subprocess
import sys
import tempfile
import time
import urllib.request
import zipfile
# --- Pinned Android command-line tools (revision 20.0) --------------------
# android-actions/setup-android v4.0.1 defaults to this same build (its
# getVersionShort() maps "14742923" -> "20.0"). We provision it OURSELVES,
# integrity-checked, into the path the action looks for first
# ($ANDROID_SDK_ROOT/cmdline-tools/20.0), so the action finds it, skips its own
# unverified download, and never prints "Wrong version in preinstalled
# sdkmanager".
#
# CLT_SIZE + the SHA-1 are Google's published values for this immutable,
# build-numbered zip (repository2-3.xml). CLT_SHA256 was computed locally from
# bytes that matched BOTH of Google's published values, so it is an authoritative
# integrity pin. A build-numbered URL is immutable, so these never drift; bumping
# the tools means bumping all four constants together.
CLT_VERSION_LONG = "14742923"
CLT_VERSION_SHORT = "20.0"
CLT_URL = (
"https://dl.google.com/android/repository/"
f"commandlinetools-linux-{CLT_VERSION_LONG}_latest.zip"
)
CLT_SIZE = 172789259
CLT_SHA256 = "04453066b540409d975c676d781da1477479dde3761310f1a7eb92a1dfb15af7"
# Total tries (1 initial + retries). Backoff is linear: 10s, 20s, 30s ...
MAX_ATTEMPTS = 4
def log(msg: str) -> None:
print(msg, flush=True)
def warn(msg: str) -> None:
print(f"::warning::{msg}", flush=True)
def error(msg: str) -> None:
print(f"::error::{msg}", flush=True)
def backoff_seconds(attempt: int) -> int:
"""Linear backoff before the next attempt: 10s after attempt 1, 20s after 2..."""
return 10 * attempt
def sdk_root() -> str:
"""The Android SDK root. GitHub-hosted runners preset ANDROID_SDK_ROOT /
ANDROID_HOME to /usr/local/lib/android/sdk; fall back to the SDK's default."""
root = os.environ.get("ANDROID_SDK_ROOT") or os.environ.get("ANDROID_HOME")
if not root:
root = os.path.join(os.path.expanduser("~"), ".android", "sdk")
return root
def sha256_of(path: str) -> str:
h = hashlib.sha256()
with open(path, "rb") as fh:
for chunk in iter(lambda: fh.read(1024 * 1024), b""):
h.update(chunk)
return h.hexdigest()
def verify_download(path, expected_size, expected_sha256):
"""(ok, detail) for a downloaded file: size first (cheap), then SHA-256.
A mismatch means a corrupt/truncated download OR the wrong version -- both
must be rejected and re-fetched."""
if not os.path.exists(path):
return False, "download missing"
actual_size = os.path.getsize(path)
if actual_size != expected_size:
return False, f"size {actual_size} != expected {expected_size}"
actual_sha = sha256_of(path)
if actual_sha != expected_sha256:
return False, f"sha256 {actual_sha} != expected {expected_sha256}"
return True, "ok"
def package_dir(root: str, package: str) -> str:
"""On-disk dir for an sdkmanager package id. sdkmanager lays packages out by
turning the ';' separators into path separators, e.g.
'platforms;android-37.0' -> <root>/platforms/android-37.0, so this is exactly
the tree to purge to force a corrupt package to re-download."""
return os.path.join(root, *package.split(";"))
def _rm(path: str) -> None:
"""Best-effort recursive delete of a file or dir (reject a bad download)."""
if os.path.islink(path) or os.path.isfile(path):
try:
os.remove(path)
except FileNotFoundError:
pass
elif os.path.isdir(path):
shutil.rmtree(path, ignore_errors=True)
def _extract_preserving_perms(zip_path: str, target_dir: str) -> None:
"""Extract a zip, restoring the unix permission bits stored in each entry's
external attributes. ZipFile.extractall drops the executable bit, which would
leave bin/sdkmanager non-executable and break the action's `sdkmanager
--licenses`; Google's zip is unix-built, so external_attr carries the +x."""
with zipfile.ZipFile(zip_path) as zf:
for info in zf.infolist():
extracted = zf.extract(info, target_dir)
mode = (info.external_attr >> 16) & 0o7777
if mode:
os.chmod(extracted, mode)
class _RejectAndRetry(Exception):
"""Internal signal: discard this attempt's download and retry from scratch."""
def bootstrap() -> int:
"""Ensure $ANDROID_SDK_ROOT/cmdline-tools/<rev> is a verified install."""
root = sdk_root()
dest = os.path.join(root, "cmdline-tools", CLT_VERSION_SHORT)
sdkmanager = os.path.join(dest, "bin", "sdkmanager")
if os.path.exists(sdkmanager):
# Cache hit (or already provisioned): the cache is populated only after a
# passing integrity check, so a present tree is trusted -> no re-download.
log(f"cmdline-tools {CLT_VERSION_SHORT} already present at {dest} "
"(cache hit) -- skipping verified download")
return 0
tools_parent = os.path.join(root, "cmdline-tools")
os.makedirs(tools_parent, exist_ok=True)
for attempt in range(1, MAX_ATTEMPTS + 1):
log(f"::group::Download + verify cmdline-tools {CLT_VERSION_SHORT} "
f"(attempt {attempt}/{MAX_ATTEMPTS})")
tmp_zip = os.path.join(tempfile.gettempdir(), f"clt-{CLT_VERSION_LONG}.zip")
# Extract on the SAME filesystem as `dest` so the final move is an atomic
# rename that preserves the restored +x bit on bin/sdkmanager.
tmp_extract = tempfile.mkdtemp(prefix=".clt-extract-", dir=tools_parent)
_rm(tmp_zip)
try:
log(f"Downloading {CLT_URL}")
urllib.request.urlretrieve(CLT_URL, tmp_zip) # noqa: S310 (pinned https)
ok, detail = verify_download(tmp_zip, CLT_SIZE, CLT_SHA256)
if not ok:
warn(f"cmdline-tools integrity check failed: {detail} -- "
"rejecting the bad download and retrying clean")
raise _RejectAndRetry()
log(f"Integrity OK (size {CLT_SIZE}, sha256 {CLT_SHA256})")
_extract_preserving_perms(tmp_zip, tmp_extract)
unpacked = os.path.join(tmp_extract, "cmdline-tools")
if not os.path.isdir(unpacked):
warn("extracted zip has no top-level cmdline-tools/ dir -- retrying")
raise _RejectAndRetry()
_rm(dest) # drop any half-extracted leftover before moving the good tree
shutil.move(unpacked, dest)
# Mirror the action: touch repositories.cfg so sdkmanager is happy.
open(os.path.join(root, "repositories.cfg"), "a", encoding="utf-8").close()
if os.path.exists(sdkmanager):
log(f"Installed verified cmdline-tools to {dest}")
return 0
warn("sdkmanager missing after extract -- retrying")
except _RejectAndRetry:
pass
except Exception as exc: # noqa: BLE001 - any transient error is retryable
warn(f"cmdline-tools bootstrap attempt {attempt} failed: {exc}")
finally:
_rm(tmp_zip)
_rm(tmp_extract)
log("::endgroup::")
if attempt < MAX_ATTEMPTS:
time.sleep(backoff_seconds(attempt))
error(f"Failed to provision verified cmdline-tools after {MAX_ATTEMPTS} attempts")
return 1
def find_sdkmanager(root: str):
"""Locate sdkmanager: our pinned rev first, then the action's `latest`, then PATH."""
candidates = [
os.path.join(root, "cmdline-tools", CLT_VERSION_SHORT, "bin", "sdkmanager"),
os.path.join(root, "cmdline-tools", "latest", "bin", "sdkmanager"),
]
for candidate in candidates:
if os.path.exists(candidate):
return candidate
return shutil.which("sdkmanager")
def install(packages) -> int:
"""`sdkmanager --install <packages>` with retry + purge-on-corrupt-zip."""
root = sdk_root()
sdkmanager = find_sdkmanager(root)
if not sdkmanager:
error("sdkmanager not found -- run the cmdline-tools bootstrap step first")
return 1
# Feed 'y' repeatedly in case any license needs accepting (setup-android's
# --licenses runs first, but this keeps the step self-contained).
accept = ("y\n" * 32).encode()
for attempt in range(1, MAX_ATTEMPTS + 1):
log(f"::group::sdkmanager --install {' '.join(packages)} "
f"(attempt {attempt}/{MAX_ATTEMPTS})")
result = subprocess.run([sdkmanager, "--install", *packages], input=accept)
log("::endgroup::")
if result.returncode == 0:
log(f"Installed SDK packages: {' '.join(packages)}")
return 0
warn(f"sdkmanager attempt {attempt} failed (exit {result.returncode}) -- "
"purging partial/corrupt packages before retry")
# REJECT: a corrupt package zip must be re-downloaded, not re-read. Purge
# each requested package's dir + sdkmanager's temp/intermediate dirs so
# the retry starts clean.
for pkg in packages:
_rm(package_dir(root, pkg))
_rm(os.path.join(root, ".temp"))
_rm(os.path.join(root, ".downloadIntermediates"))
if attempt < MAX_ATTEMPTS:
time.sleep(backoff_seconds(attempt))
error(f"sdkmanager failed to install {list(packages)} after {MAX_ATTEMPTS} attempts")
log("--- sdkmanager --list_installed ---")
subprocess.run([sdkmanager, "--list_installed"])
return 1
def build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
description="Hardened Android SDK setup for CI (issue #389)."
)
sub = parser.add_subparsers(dest="command", required=True)
sub.add_parser(
"bootstrap",
help="Download + SHA-256-verify the pinned Android command-line tools.",
)
installer = sub.add_parser(
"install",
help="sdkmanager --install with retry + purge-on-corrupt-zip.",
)
installer.add_argument("packages", nargs="+", help="sdkmanager package ids")
return parser
def main(argv) -> int:
args = build_parser().parse_args(argv)
if args.command == "bootstrap":
return bootstrap()
if args.command == "install":
return install(args.packages)
return 2 # pragma: no cover - argparse requires a subcommand
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
+178
View File
@@ -0,0 +1,178 @@
#!/usr/bin/env python3
# SPDX-License-Identifier: GPL-3.0-or-later
"""Unit tests for the pure helpers of setup_android_sdk.py (no network, no SDK).
Covers the bits whose correctness is load-bearing for the hardening in #389:
the package-id -> purge-path mapping (a wrong mapping would purge the wrong dir),
the size/SHA-256 integrity gate (verify -> reject), the pinned-constant
self-consistency, the backoff schedule, sdkmanager discovery, and that extraction
restores the executable bit that sdkmanager needs. The full download/install path
is exercised by CI itself."""
from __future__ import annotations
import hashlib
import os
import stat
import tempfile
import unittest
import zipfile
import setup_android_sdk as sdk
class PackageDirTests(unittest.TestCase):
def test_semicolon_ids_map_to_nested_dirs(self):
root = os.path.join("opt", "sdk")
self.assertEqual(
sdk.package_dir(root, "platforms;android-37.0"),
os.path.join(root, "platforms", "android-37.0"),
)
self.assertEqual(
sdk.package_dir(root, "build-tools;37.0.0"),
os.path.join(root, "build-tools", "37.0.0"),
)
self.assertEqual(
sdk.package_dir(root, "system-images;android-37.0;google_apis_ps16k;x86_64"),
os.path.join(root, "system-images", "android-37.0", "google_apis_ps16k", "x86_64"),
)
def test_flat_ids_map_to_single_dir(self):
root = os.path.join("opt", "sdk")
self.assertEqual(sdk.package_dir(root, "emulator"), os.path.join(root, "emulator"))
self.assertEqual(
sdk.package_dir(root, "platform-tools"), os.path.join(root, "platform-tools")
)
def test_purge_target_stays_under_root(self):
# The purge path must never escape the SDK root (no absolute/`..` package ids).
root = os.path.abspath(os.path.join("opt", "sdk"))
target = os.path.abspath(sdk.package_dir(root, "platforms;android-37.0"))
self.assertTrue(target.startswith(root + os.sep))
class VerifyDownloadTests(unittest.TestCase):
def _write(self, data: bytes) -> str:
fd, path = tempfile.mkstemp()
with os.fdopen(fd, "wb") as fh:
fh.write(data)
self.addCleanup(lambda: os.path.exists(path) and os.remove(path))
return path
def test_accepts_matching_size_and_hash(self):
data = b"correct-cmdline-tools-bytes"
path = self._write(data)
ok, detail = sdk.verify_download(path, len(data), hashlib.sha256(data).hexdigest())
self.assertTrue(ok, detail)
self.assertEqual(detail, "ok")
def test_rejects_wrong_size_before_hashing(self):
data = b"truncated"
path = self._write(data)
ok, detail = sdk.verify_download(path, len(data) + 1, hashlib.sha256(data).hexdigest())
self.assertFalse(ok)
self.assertIn("size", detail)
def test_rejects_corrupt_bytes_with_right_size(self):
good = b"aaaaaaaa"
corrupt = b"aaaaaaab" # same length, different content (silent corruption)
path = self._write(corrupt)
ok, detail = sdk.verify_download(path, len(good), hashlib.sha256(good).hexdigest())
self.assertFalse(ok)
self.assertIn("sha256", detail)
def test_rejects_missing_file(self):
ok, detail = sdk.verify_download(
os.path.join(tempfile.gettempdir(), "does-not-exist-clt.zip"), 1, "0" * 64
)
self.assertFalse(ok)
class PinnedConstantsTests(unittest.TestCase):
def test_url_embeds_the_pinned_build_number(self):
self.assertIn(sdk.CLT_VERSION_LONG, sdk.CLT_URL)
self.assertTrue(sdk.CLT_URL.startswith("https://"))
self.assertTrue(sdk.CLT_URL.endswith("_latest.zip"))
def test_sha256_is_a_full_hex_digest(self):
self.assertEqual(len(sdk.CLT_SHA256), 64)
int(sdk.CLT_SHA256, 16) # raises if not hex
self.assertEqual(sdk.CLT_SHA256, sdk.CLT_SHA256.lower())
def test_size_is_positive(self):
self.assertGreater(sdk.CLT_SIZE, 0)
class BackoffTests(unittest.TestCase):
def test_backoff_is_linear_and_increasing(self):
seq = [sdk.backoff_seconds(a) for a in range(1, sdk.MAX_ATTEMPTS + 1)]
self.assertEqual(seq, [10, 20, 30, 40][: sdk.MAX_ATTEMPTS])
self.assertEqual(seq, sorted(seq))
class SdkRootTests(unittest.TestCase):
def test_prefers_android_sdk_root_over_home(self):
with _env(ANDROID_SDK_ROOT="/a/sdk-root", ANDROID_HOME="/b/home"):
self.assertEqual(sdk.sdk_root(), "/a/sdk-root")
def test_falls_back_to_android_home(self):
with _env(ANDROID_SDK_ROOT=None, ANDROID_HOME="/b/home"):
self.assertEqual(sdk.sdk_root(), "/b/home")
class FindSdkManagerTests(unittest.TestCase):
def test_prefers_pinned_revision_dir(self):
with tempfile.TemporaryDirectory() as root:
pinned = os.path.join(root, "cmdline-tools", sdk.CLT_VERSION_SHORT, "bin")
latest = os.path.join(root, "cmdline-tools", "latest", "bin")
for d in (pinned, latest):
os.makedirs(d)
open(os.path.join(d, "sdkmanager"), "w").close()
self.assertEqual(
sdk.find_sdkmanager(root),
os.path.join(pinned, "sdkmanager"),
)
class ExtractPermsTests(unittest.TestCase):
@unittest.skipUnless(os.name == "posix", "unix exec bit only meaningful on POSIX")
def test_executable_bit_is_restored(self):
with tempfile.TemporaryDirectory() as work:
zip_path = os.path.join(work, "clt.zip")
with zipfile.ZipFile(zip_path, "w") as zf:
info = zipfile.ZipInfo("cmdline-tools/bin/sdkmanager")
info.external_attr = 0o755 << 16 # -rwxr-xr-x, as Google's zip stores it
zf.writestr(info, "#!/bin/sh\n")
out = os.path.join(work, "out")
sdk._extract_preserving_perms(zip_path, out)
mode = os.stat(os.path.join(out, "cmdline-tools", "bin", "sdkmanager")).st_mode
self.assertTrue(mode & stat.S_IXUSR, "sdkmanager must be executable after extract")
class _env:
"""Context manager to set/clear env vars for a test, restoring them after."""
def __init__(self, **values):
self._values = values
self._saved = {}
def __enter__(self):
for key, value in self._values.items():
self._saved[key] = os.environ.get(key)
if value is None:
os.environ.pop(key, None)
else:
os.environ[key] = value
return self
def __exit__(self, *exc):
for key, previous in self._saved.items():
if previous is None:
os.environ.pop(key, None)
else:
os.environ[key] = previous
return False
if __name__ == "__main__":
unittest.main()
+509
View File
@@ -0,0 +1,509 @@
#!/usr/bin/env python3
# SPDX-License-Identifier: GPL-3.0-or-later
"""Unit tests for the pure decision core of traffic_control.py (no network).
Covers: priority resolution (P-label / broken / draft / default P5), PASS 1
preemption (P0 reclaims all strictly-lower; ANY higher PR reclaims a broken/draft
lower run; P1-P9 never bump a *normal* lower run; self / main / equal-or-higher
never cancelled), PASS 2 hold-back (yield to strictly-higher with an active run;
same-level running-first then oldest-first), and a few end-to-end decision
scenarios.
Also covers the --mode trigger scheduler core (issue #349): head-SHA run
classification (absent/cancelled => needy; success/failure => not needy),
select_triggers (priority order, oldest-first fairness, inflight cap, fork skip,
P0 bypasses-cap-and-preempts), and a liveness/anti-starvation simulation proving
every eligible PR is triggered within a bounded number of passes."""
from __future__ import annotations
import json
import unittest
import traffic_control as tc
from traffic_control import PullRequest
def pr(number, labels=(), *, draft=False, created_at="", status=tc.NONE, run_ids=()):
"""Terse PullRequest builder for tests."""
return PullRequest(
number=number,
labels=tuple(labels),
is_draft=draft,
created_at=created_at,
run_status=status,
run_ids=tuple(run_ids),
)
class EffectivePriorityTests(unittest.TestCase):
def test_no_labels_defaults_to_p5(self):
self.assertEqual(tc.effective_priority(pr(1)), 5)
def test_non_priority_labels_ignored_default_p5(self):
self.assertEqual(tc.effective_priority(pr(1, ["bug", "enhancement"])), 5)
def test_single_p_label(self):
self.assertEqual(tc.effective_priority(pr(1, ["P3"])), 3)
self.assertEqual(tc.effective_priority(pr(1, ["P0"])), 0)
def test_lowest_numbered_p_label_wins(self):
self.assertEqual(tc.effective_priority(pr(1, ["P4", "P1", "P7"])), 1)
def test_broken_is_bottom_p10_overriding_p0(self):
self.assertEqual(tc.effective_priority(pr(1, ["broken", "P0"])), 10)
def test_draft_is_bottom_p10_overriding_p0(self):
self.assertEqual(tc.effective_priority(pr(1, ["P0"], draft=True)), 10)
def test_draft_and_broken_still_p10(self):
self.assertEqual(tc.effective_priority(pr(1, ["broken"], draft=True)), 10)
def test_double_digit_pseudo_label_is_not_a_priority(self):
# Only P0-P9 count (regex ^P[0-9]$); "P10" is not a valid priority label.
self.assertEqual(tc.effective_priority(pr(1, ["P10"])), 5)
def test_priority_label_text(self):
self.assertEqual(tc.priority_label(0), "P0")
self.assertEqual(tc.priority_label(5), "P5")
self.assertIn("bottom", tc.priority_label(10))
class RunsToCancelTests(unittest.TestCase):
def test_non_p0_self_does_not_bump_normal_lower_run(self):
# P1-P9 never preempt a *normal* strictly-lower run — they yield instead.
me = pr(1, ["P1"])
others = [pr(2, ["P5"], status=tc.RUNNING, run_ids=[200])]
self.assertEqual(tc.runs_to_cancel(me, [me, *others]), [])
def test_non_p0_self_reclaims_broken_lower_run(self):
# Any higher-priority PR (not just P0) may reclaim a broken target's runner.
me = pr(1, ["P3"])
broken = pr(2, ["broken"], status=tc.RUNNING, run_ids=[200])
self.assertEqual(tc.runs_to_cancel(me, [me, broken]), [200])
def test_non_p0_self_reclaims_draft_lower_run(self):
# A draft is not merge-ready — its run is likewise reclaimable by any higher PR.
me = pr(1, ["P3"])
draft = pr(2, [], draft=True, status=tc.QUEUED, run_ids=[200])
self.assertEqual(tc.runs_to_cancel(me, [me, draft]), [200])
def test_broken_self_does_not_cancel_equal_broken(self):
# Both effective P10 — the equal-or-higher invariant still forbids cancelling.
me = pr(1, ["broken"])
peer = pr(2, ["broken"], status=tc.RUNNING, run_ids=[200])
self.assertEqual(tc.runs_to_cancel(me, [me, peer]), [])
def test_bottom_self_preempts_nothing(self):
# A broken/draft PR (P10) is the bottom: nothing is strictly-lower, so it
# cancels neither a higher (P5) nor an equal (P10) run.
me = pr(1, [], draft=True) # P10
prs = [
me,
pr(2, ["P5"], status=tc.RUNNING, run_ids=[200]), # higher
pr(3, ["broken"], status=tc.RUNNING, run_ids=[300]), # equal P10
]
self.assertEqual(tc.runs_to_cancel(me, prs), [])
def test_p0_cancels_strictly_lower_active_runs(self):
me = pr(1, ["P0"])
low = pr(2, ["P5"], status=tc.RUNNING, run_ids=[200])
queued = pr(3, ["P9"], status=tc.QUEUED, run_ids=[300])
self.assertEqual(
sorted(tc.runs_to_cancel(me, [me, low, queued])), [200, 300])
def test_p0_cancels_broken_and_draft_lower_runs(self):
me = pr(1, ["P0"])
broken = pr(2, ["broken"], status=tc.RUNNING, run_ids=[200])
draft = pr(3, ["P2"], draft=True, status=tc.RUNNING, run_ids=[300])
self.assertEqual(
sorted(tc.runs_to_cancel(me, [me, broken, draft])), [200, 300])
def test_p0_never_cancels_equal_priority_p0(self):
me = pr(1, ["P0"])
peer = pr(2, ["P0"], status=tc.RUNNING, run_ids=[200])
self.assertEqual(tc.runs_to_cancel(me, [me, peer]), [])
def test_p0_never_cancels_self(self):
me = pr(1, ["P0"], status=tc.RUNNING, run_ids=[100])
self.assertEqual(tc.runs_to_cancel(me, [me]), [])
def test_p0_excludes_own_run_id_defensively(self):
me = pr(1, ["P0"], status=tc.RUNNING, run_ids=[100])
# A lower PR that somehow reports our own run id must not be cancelled.
low = pr(2, ["P5"], status=tc.RUNNING, run_ids=[100, 200])
self.assertEqual(
tc.runs_to_cancel(me, [me, low], self_run_id=100), [200])
def test_p0_skips_lower_with_no_active_run(self):
me = pr(1, ["P0"])
idle = pr(2, ["P5"], status=tc.NONE, run_ids=[])
self.assertEqual(tc.runs_to_cancel(me, [me, idle]), [])
class WaitBlockersTests(unittest.TestCase):
def test_p0_never_waits(self):
me = pr(1, ["P0"])
higher = pr(2, ["P0"], status=tc.RUNNING) # nothing outranks P0 anyway
self.assertEqual(tc.wait_blockers(me, [me, higher]), [])
def test_yields_to_strictly_higher_with_active_run(self):
me = pr(2, ["P5"], status=tc.RUNNING)
higher = pr(1, ["P2"], status=tc.RUNNING)
blockers = tc.wait_blockers(me, [me, higher])
self.assertEqual([b.number for b in blockers], [1])
self.assertEqual(blockers[0].kind, "higher-priority")
def test_does_not_yield_to_higher_without_active_run(self):
me = pr(2, ["P5"], status=tc.RUNNING)
higher_idle = pr(1, ["P2"], status=tc.NONE)
self.assertEqual(tc.wait_blockers(me, [me, higher_idle]), [])
def test_does_not_yield_to_lower_priority(self):
me = pr(1, ["P2"], status=tc.RUNNING)
lower = pr(2, ["P5"], status=tc.RUNNING)
self.assertEqual(tc.wait_blockers(me, [me, lower]), [])
def test_same_level_oldest_running_proceeds(self):
me = pr(1, ["P5"], created_at="2026-07-01T00:00:00Z", status=tc.RUNNING)
newer = pr(2, ["P5"], created_at="2026-07-02T00:00:00Z", status=tc.RUNNING)
self.assertEqual(tc.wait_blockers(me, [me, newer]), [])
def test_same_level_newer_running_yields_to_older(self):
# Coordinator clarification: within a level, older createdAt goes first.
older = pr(1, ["P5"], created_at="2026-07-01T00:00:00Z", status=tc.RUNNING)
me = pr(2, ["P5"], created_at="2026-07-02T00:00:00Z", status=tc.RUNNING)
blockers = tc.wait_blockers(me, [me, older])
self.assertEqual([b.number for b in blockers], [1])
self.assertEqual(blockers[0].kind, "same-level-ahead")
def test_same_level_running_first_beats_older_waiting(self):
# An in-flight peer keeps its place; a not-yet-running OLDER peer does not
# jump ahead of us while we are the one already running.
me = pr(2, ["P5"], created_at="2026-07-02T00:00:00Z", status=tc.RUNNING)
older_waiting = pr(1, ["P5"], created_at="2026-07-01T00:00:00Z", status=tc.NONE)
self.assertEqual(tc.wait_blockers(me, [me, older_waiting]), [])
def test_same_level_waiting_orders_oldest_before_newer(self):
# Neither running: strictly oldest-first among the waiting bucket.
oldest = pr(1, ["P5"], created_at="2026-07-01T00:00:00Z", status=tc.NONE)
middle = pr(2, ["P5"], created_at="2026-07-02T00:00:00Z", status=tc.NONE)
me = pr(3, ["P5"], created_at="2026-07-03T00:00:00Z", status=tc.NONE)
blockers = tc.wait_blockers(me, [oldest, middle, me])
self.assertEqual([b.number for b in blockers], [1, 2])
def test_same_level_queued_counts_as_waiting_ordered_by_age(self):
# A queued peer is "waiting to start", not in-flight: ordered purely by age.
me = pr(1, ["P5"], created_at="2026-07-01T00:00:00Z", status=tc.NONE)
newer_queued = pr(2, ["P5"], created_at="2026-07-02T00:00:00Z", status=tc.QUEUED)
self.assertEqual(tc.wait_blockers(me, [me, newer_queued]), [])
def test_broken_self_yields_to_everyone_active(self):
me = pr(1, ["broken"], status=tc.RUNNING) # effective P10
normal = pr(2, ["P5"], status=tc.RUNNING)
blockers = tc.wait_blockers(me, [me, normal])
self.assertEqual([b.number for b in blockers], [2])
self.assertEqual(blockers[0].kind, "higher-priority")
class EndToEndDecisionTests(unittest.TestCase):
def test_p0_emergency_cancels_lower_and_proceeds(self):
me = pr(10, ["P0"], status=tc.RUNNING, run_ids=[1000])
prs = [
me,
pr(11, ["P2"], status=tc.RUNNING, run_ids=[1100]),
pr(12, ["P5"], status=tc.QUEUED, run_ids=[1200]),
pr(13, ["P0"], status=tc.RUNNING, run_ids=[1300]), # equal — spared
]
dec = tc.decide(me, prs, self_run_id=1000)
self.assertEqual(sorted(dec.cancel_run_ids), [1100, 1200])
self.assertTrue(dec.proceed)
def test_p5_waits_behind_running_higher(self):
me = pr(20, ["P5"], status=tc.RUNNING, run_ids=[2000])
higher = pr(21, ["P2"], status=tc.RUNNING, run_ids=[2100])
dec = tc.decide(me, [me, higher])
self.assertEqual(dec.cancel_run_ids, ()) # not P0 — cancels nothing
self.assertFalse(dec.proceed)
self.assertEqual([b.number for b in dec.blockers], [21])
def test_lone_p5_proceeds(self):
me = pr(30, ["P5"], status=tc.RUNNING, run_ids=[3000])
dec = tc.decide(me, [me])
self.assertEqual(dec.cancel_run_ids, ())
self.assertTrue(dec.proceed)
def test_p5_reclaims_draft_then_waits_behind_higher(self):
# A non-P0 PR can BOTH reclaim a broken/draft lower run (PASS 1) AND still
# yield to a strictly-higher PR (PASS 2) in the same evaluation.
me = pr(40, ["P5"], status=tc.RUNNING, run_ids=[4000])
prs = [
me,
pr(41, ["P2"], status=tc.RUNNING, run_ids=[4100]), # higher — blocks
pr(42, [], draft=True, status=tc.RUNNING, run_ids=[4200]), # draft — reclaimed
]
dec = tc.decide(me, prs)
self.assertEqual(list(dec.cancel_run_ids), [4200])
self.assertFalse(dec.proceed)
self.assertEqual([b.number for b in dec.blockers], [41])
class SnapshotParsingTests(unittest.TestCase):
def test_from_json_label_objects_and_fields(self):
obj = {
"number": 7,
"labels": [{"name": "P3"}, {"name": "bug"}],
"isDraft": True,
"createdAt": "2026-07-01T00:00:00Z",
"runStatus": "running",
"runIds": [42, 43],
}
p = PullRequest.from_json(obj)
self.assertEqual(p.number, 7)
self.assertEqual(p.labels, ("P3", "bug"))
self.assertTrue(p.is_draft)
self.assertEqual(p.run_status, tc.RUNNING)
self.assertEqual(p.run_ids, (42, 43))
self.assertEqual(tc.effective_priority(p), 10) # draft => bottom
def test_from_json_plain_string_labels_and_unknown_status(self):
p = PullRequest.from_json(
{"number": 8, "labels": ["P1"], "runStatus": "bogus"})
self.assertEqual(p.labels, ("P1",))
self.assertEqual(p.run_status, tc.NONE) # unknown -> none
def test_load_snapshot_roundtrip(self):
text = json.dumps({
"self": 2,
"self_run_id": 222,
"prs": [
{"number": 1, "labels": ["P2"], "runStatus": "running",
"runIds": [111]},
{"number": 2, "labels": ["P5"], "runStatus": "running",
"runIds": [222]},
],
})
this_pr, all_prs, self_run_id = tc._load_snapshot(text)
self.assertEqual(this_pr.number, 2)
self.assertEqual(len(all_prs), 2)
self.assertEqual(self_run_id, 222)
# ── --mode trigger scheduler core (issue #349) ───────────────────────────────
def npr(number, labels=("P5",), *, draft=False, created_at="", status=tc.NONE, run_ids=()):
"""Terse builder defaulting to a P5 PR (for the trigger tests)."""
return pr(number, labels, draft=draft, created_at=created_at, status=status,
run_ids=run_ids)
class ClassifyShaRunsTests(unittest.TestCase):
def test_no_runs_is_needy(self):
status, ids, needy = tc.classify_sha_runs([])
self.assertEqual(status, tc.NONE)
self.assertEqual(ids, ())
self.assertTrue(needy) # absent checks => must be triggered
def test_success_verdict_not_needy(self):
_, ids, needy = tc.classify_sha_runs(
[{"status": "completed", "conclusion": "success", "databaseId": 1}])
self.assertEqual(ids, ())
self.assertFalse(needy)
def test_failure_verdict_not_needy(self):
# A real failure is the author's to fix — never auto-retriggered (no fail loop).
_, _, needy = tc.classify_sha_runs(
[{"status": "completed", "conclusion": "failure", "databaseId": 1}])
self.assertFalse(needy)
def test_only_cancelled_is_needy(self):
# Cancelled leaves no verdict → re-trigger so the PR can reach a mergeable state.
_, ids, needy = tc.classify_sha_runs(
[{"status": "completed", "conclusion": "cancelled", "databaseId": 1}])
self.assertEqual(ids, ())
self.assertTrue(needy)
def test_in_progress_is_active_not_needy(self):
status, ids, needy = tc.classify_sha_runs(
[{"status": "in_progress", "conclusion": None, "databaseId": 9}])
self.assertEqual(status, tc.RUNNING)
self.assertEqual(ids, (9,))
self.assertFalse(needy)
def test_queued_is_active_not_needy(self):
status, ids, needy = tc.classify_sha_runs(
[{"status": "queued", "conclusion": None, "databaseId": 8}])
self.assertEqual(status, tc.QUEUED)
self.assertEqual(ids, (8,))
self.assertFalse(needy)
def test_running_beats_queued_in_status(self):
status, ids, _ = tc.classify_sha_runs([
{"status": "queued", "databaseId": 1},
{"status": "in_progress", "databaseId": 2},
])
self.assertEqual(status, tc.RUNNING)
self.assertEqual(sorted(ids), [1, 2])
def test_cancelled_plus_active_not_needy(self):
# An active run already covers the SHA — cancelled siblings don't make it needy.
_, ids, needy = tc.classify_sha_runs([
{"status": "completed", "conclusion": "cancelled", "databaseId": 1},
{"status": "in_progress", "databaseId": 2},
])
self.assertEqual(ids, (2,))
self.assertFalse(needy)
class SelectTriggersTests(unittest.TestCase):
def test_empty_needy_triggers_nothing(self):
dec = tc.select_triggers([npr(1), npr(2)], set(), max_inflight=3)
self.assertEqual(dec.trigger_numbers, ())
def test_single_needy_triggered(self):
dec = tc.select_triggers([npr(1)], {1}, max_inflight=3)
self.assertEqual(dec.trigger_numbers, (1,))
self.assertEqual(dec.cancel_run_ids, ())
def test_priority_order(self):
prs = [npr(1, ["P5"]), npr(2, ["P2"]), npr(3, ["P8"])]
dec = tc.select_triggers(prs, {1, 2, 3}, max_inflight=3)
self.assertEqual(dec.trigger_numbers, (2, 1, 3)) # P2, P5, P8
def test_same_level_oldest_first(self):
prs = [
npr(1, ["P5"], created_at="2026-07-03T00:00:00Z"),
npr(2, ["P5"], created_at="2026-07-01T00:00:00Z"),
npr(3, ["P5"], created_at="2026-07-02T00:00:00Z"),
]
dec = tc.select_triggers(prs, {1, 2, 3}, max_inflight=3)
self.assertEqual(dec.trigger_numbers, (2, 3, 1)) # oldest createdAt first
def test_cap_limits_triggers(self):
prs = [npr(1, ["P2"]), npr(2, ["P3"]), npr(3, ["P4"])]
dec = tc.select_triggers(prs, {1, 2, 3}, max_inflight=2)
self.assertEqual(dec.trigger_numbers, (1, 2)) # only 2 free slots
self.assertEqual(dec.slots, 2)
def test_inflight_consumes_slots(self):
prs = [
npr(1, ["P5"], status=tc.RUNNING, run_ids=[100]), # inflight — occupies a slot
npr(2, ["P2"]),
npr(3, ["P3"]),
]
dec = tc.select_triggers(prs, {2, 3}, max_inflight=2)
self.assertEqual(dec.slots, 1) # 2 cap - 1 inflight
self.assertEqual(dec.trigger_numbers, (2,)) # highest-priority needy only
self.assertEqual(dec.inflight_numbers, (1,))
def test_running_needy_is_not_retriggered(self):
# Defensive: a PR flagged needy but already running is never a candidate.
prs = [npr(1, ["P5"], status=tc.RUNNING, run_ids=[100])]
dec = tc.select_triggers(prs, {1}, max_inflight=3)
self.assertEqual(dec.trigger_numbers, ())
def test_fork_pr_skipped(self):
dec = tc.select_triggers([npr(1, ["P2"]), npr(2, ["P1"])],
{1, 2}, max_inflight=3, forks={2})
self.assertEqual(dec.trigger_numbers, (1,)) # fork #2 not token-triggerable
self.assertEqual(dec.skipped_fork_numbers, (2,))
def test_p0_bypasses_cap_and_preempts_lower(self):
# Cap full (a P5 running), but a needy P0 still triggers AND preempts the strictly-
# lower running run to free a runner immediately.
prs = [
npr(1, ["P5"], status=tc.RUNNING, run_ids=[500]), # inflight, strictly-lower
npr(2, ["P0"]), # needy emergency
]
dec = tc.select_triggers(prs, {2}, max_inflight=1)
self.assertEqual(dec.slots, 0) # cap is full
self.assertEqual(dec.trigger_numbers, (2,)) # P0 bypasses the cap
self.assertEqual(dec.cancel_run_ids, (500,)) # preempts the lower run
def test_p0_does_not_preempt_equal_priority(self):
prs = [
npr(1, ["P0"], status=tc.RUNNING, run_ids=[500]), # equal P0 — spared
npr(2, ["P0"]), # needy emergency
]
dec = tc.select_triggers(prs, {2}, max_inflight=1)
self.assertEqual(dec.trigger_numbers, (2,))
self.assertEqual(dec.cancel_run_ids, ()) # never preempts an equal P0
def test_needy_numbers_reports_all_candidates_in_order(self):
dec = tc.select_triggers([npr(1, ["P5"]), npr(2, ["P2"])], {1, 2}, max_inflight=1)
self.assertEqual(dec.needy_numbers, (2, 1)) # priority order, cap-independent
self.assertEqual(dec.trigger_numbers, (2,)) # but only 1 slot triggered
def test_zero_cap_is_clamped_to_one(self):
# A misconfigured cap must never stall everything: clamp to >= 1 so at least the
# top-priority needy PR still gets a slot (fail-safe forward progress).
dec = tc.select_triggers([npr(1, ["P5"])], {1}, max_inflight=0)
self.assertEqual(dec.max_inflight, 1)
self.assertEqual(dec.trigger_numbers, (1,))
class TriggerStarvationTests(unittest.TestCase):
def test_every_needy_pr_is_triggered_within_bounded_passes(self):
# Liveness / anti-starvation: with a fixed needy set and cap=2, simulate scheduler
# passes where a triggered PR gains a run (leaves the needy set) and its run finishes
# one pass later (freeing its slot). Mixed priorities prove lower-priority PRs are
# served too — never starved — while higher-priority PRs still go first.
labels = {1: ["P1"], 2: ["P1"], 3: ["P5"], 4: ["P5"],
5: ["P8"], 6: ["P8"], 7: ["P5"]}
created = {n: f"2026-07-{n:02d}T00:00:00Z" for n in labels}
needy = set(labels)
running: dict[int, int] = {} # number -> passes left running
triggered_ever: set[int] = set()
first_pass: dict[int, int] = {}
cap, max_passes = 2, 12
for pass_no in range(1, max_passes + 1):
snap = [
pr(n, labels[n], created_at=created[n],
status=tc.RUNNING if n in running else tc.NONE,
run_ids=[1000 + n] if n in running else [])
for n in labels
]
dec = tc.select_triggers(snap, needy, max_inflight=cap)
for n in dec.trigger_numbers:
triggered_ever.add(n)
first_pass.setdefault(n, pass_no)
needy.discard(n) # gained a run -> no longer needy
running[n] = 1 # occupies a slot for one pass
for n in list(running): # running PRs finish after one pass
running[n] -= 1
if running[n] <= 0:
del running[n]
if not needy and not running:
break
self.assertEqual(triggered_ever, set(labels),
"a PR was starved (never triggered)")
# Priority respected: the two P1s go in the very first pass; everything else later.
self.assertTrue(all(first_pass[n] == 1 for n in (1, 2)))
self.assertTrue(all(first_pass[n] >= 2 for n in (3, 4, 5, 6, 7)))
class TriggerSnapshotParsingTests(unittest.TestCase):
def test_load_trigger_snapshot_roundtrip(self):
text = json.dumps({
"max_inflight": 2,
"needy": [1, 3],
"forks": [3],
"prs": [
{"number": 1, "labels": ["P2"], "createdAt": "2026-07-01T00:00:00Z"},
{"number": 2, "labels": ["P5"], "runStatus": "running", "runIds": [22]},
{"number": 3, "labels": ["P1"], "createdAt": "2026-07-02T00:00:00Z"},
],
})
all_prs, needy, forks, cap = tc._load_trigger_snapshot(text)
self.assertEqual(len(all_prs), 3)
self.assertEqual(needy, {1, 3})
self.assertEqual(forks, {3})
self.assertEqual(cap, 2)
dec = tc.select_triggers(all_prs, needy, max_inflight=cap, forks=forks)
# #2 is inflight (uses a slot); #3 is a fork (skipped); only #1 fits the 1 free slot.
self.assertEqual(dec.inflight_numbers, (2,))
self.assertEqual(dec.skipped_fork_numbers, (3,))
self.assertEqual(dec.trigger_numbers, (1,))
if __name__ == "__main__":
unittest.main()
+805
View File
@@ -0,0 +1,805 @@
#!/usr/bin/env python3
# SPDX-License-Identifier: GPL-3.0-or-later
"""CI traffic-controller: priority-based runner orchestration for LibreMail's
`ci.yml`. Extracted out of the old inline-bash `traffic-control` step into a
Python module so the decision logic is developer-legible and, above all, unit-
testable (see `test_traffic_control.py`).
DESIGN: pure decision CORE + thin gh-I/O SHELL
----------------------------------------------
The decisions ("who do we cancel?", "do we proceed or wait?") are pure functions
over plain `PullRequest` snapshots — no network, no clock, no subprocess — so they
can be exercised exhaustively in unit tests. The SHELL (`run_live`) is the only part
that touches `gh`: it gathers the snapshot, applies the cancellations, and runs the
bounded hold-back poll loop. Feed the core a snapshot JSON (`--dry-run`) to see its
decisions with zero network.
TWO MODES
---------
* ``--mode orchestrate`` (default; unchanged behaviour): the in-run `traffic-control`
job of `ci.yml`. Orders runner ACCESS for the PR whose run is already executing —
PASS 1 preemption + PASS 2 hold-back (below). This is `run_live`.
* ``--mode trigger`` (issue #349): the *scheduler* (companion `ci-trigger.yml`, run
after each auto-update and on a cron backstop). It OWNS CI *triggering*: it
(re-)triggers CI for the highest-priority PR(s) whose head SHA has absent/stale
required checks — a few at a time (an inflight cap), in the SAME priority order —
via a `workflow_dispatch`. This is `run_trigger` / the pure `select_triggers`.
WHY --mode trigger EXISTS (issue #349): `autoupdate.yml` now updates PR branches with the
built-in GITHUB_TOKEN instead of a PAT, so an update push no longer auto-retriggers CI
(GitHub's anti-recursion rule) — killing the merge-cascade that cancelled every open PR's
run on every merge. The cost is that a freshly-updated PR's required checks go stale/absent
on its NEW head SHA, so this scheduler deliberately (re-)triggers them in priority order (a
poor-man's merge queue). The dispatch uses the AUTOUPDATE_TOKEN PAT, NOT the built-in
GITHUB_TOKEN: a GITHUB_TOKEN-triggered run is held for MANUAL approval (`action_required`) and
never runs un-attended, whereas a PAT dispatch runs as the authorized owner with no approval gate
(#350's "no PAT needed" claim was wrong — see ci-trigger.yml + issue #351). FAIL-OPEN,
structurally: `ci.yml` KEEPS its `on: pull_request`
trigger, so any human push — and a brand-new PR — always gets CI regardless of this
scheduler; the scheduler only fills the gap left by GITHUB_TOKEN auto-updates and can never
leave a PR un-triggerable. Fork PRs (no token/secret access) are skipped by the scheduler and
left to `on: pull_request`, so they are never wedged either.
ORDER OF OPERATIONS (issue #342)
--------------------------------
1. Effective priority orders everything: the lowest-numbered `P0`-`P9` label present
(P0 = highest), default `P5` if none. A `broken` OR `draft` PR is effectively P10
(bottom, below P9), overriding any P0-P9 label.
2. PASS 1 - preemption: a strictly-lower OTHER PR's in-progress / queued run is
cancelled iff THIS PR is P0 (an emergency reclaims ALL lower runners) OR the target
is broken/draft (a wasted run any higher-priority PR may reclaim). P1-P9 never bump
a *normal* lower run mid-flight — only a P0 does that.
3. PASS 2 - bounded hold-back: a non-P0 PR yields (cancels nothing) to any strictly-
higher-priority OTHER PR that has an active/queued run, and — among its OWN
priority level — to any PR ordered ahead of it (running-first, then oldest by
`createdAt`). It proceeds the moment it is at the front, or when the wait budget
elapses (a PR never blocks itself).
SAFETY INVARIANTS (preserved from the original step)
----------------------------------------------------
* never cancel a run on `main` / a push event — the shell's `gh run list` query
filters `--event pull_request` and drops `headBranch == main`, so only PR-event
runs ever reach the core;
* never cancel THIS PR's own run — skipped by PR number AND by run id;
* never cancel an equal-or-higher-priority PR — only strictly-lower (prio > self).
This is deliberately NOT a merge gate: every gh call is guarded, the shell always
exits 0, and the ci.yml step stays `continue-on-error`, so a hiccup (API error,
missing permission, fork PR) can never fail CI.
Pure standard library, cross-platform (the primary dev box is Windows, where the old
bash + `jq` pipeline had no clean equivalent).
"""
from __future__ import annotations
import argparse
import json
import os
import re
import subprocess
import sys
import time
from dataclasses import dataclass, field
# ── Priority model ───────────────────────────────────────────────────────────
BROKEN_LABEL = "broken"
DEFAULT_PRIORITY = 5 # a PR with no P0-P9 label
BOTTOM_PRIORITY = 10 # broken OR draft — below P9
TOP_PRIORITY = 0 # P0, the only priority that preempts
_P_LABEL = re.compile(r"^P([0-9])$") # single digit only, matching the old jq `^P[0-9]$`
# ── Run-status model (normalised from gh's raw run statuses) ─────────────────
RUNNING = "running" # gh status in_progress
QUEUED = "queued" # gh status queued / waiting / requested / pending
NONE = "none" # no active run (completed or absent)
ACTIVE = frozenset({RUNNING, QUEUED})
# For aggregating a branch's overall status from its runs: running beats queued
# beats none (most-active wins). NB: the same-level ORDER (see _ordering_key) is a
# coarser two-bucket split — in-flight (running) vs everything-else-by-age.
_STATUS_RANK = {RUNNING: 0, QUEUED: 1, NONE: 2}
# createdAt sentinel so a PR with an unknown timestamp sorts LAST (never wrongly
# "oldest"/front, so it yields rather than preempts another PR's front slot).
_FAR_FUTURE = "9999-12-31T23:59:59Z"
# Run conclusions that count as a FINAL VERDICT on a head SHA (issue #349, --mode trigger).
# A SHA with one of these is NOT re-triggered: success = green, failure/timeout/etc. = the
# author's to fix — auto-retriggering a real failure would waste runners and could loop.
# Everything else a completed run can report (cancelled / skipped / stale / startup_failure /
# null) is treated as "no verdict", so a SHA whose only runs are those — or that has no run at
# all (absent checks after a GITHUB_TOKEN auto-update) — is NEEDY and gets (re-)triggered.
VERDICT_CONCLUSIONS = frozenset(
{"success", "failure", "timed_out", "action_required", "neutral"}
)
@dataclass(frozen=True)
class PullRequest:
"""A snapshot of one open PR. The pure decision core consumes only these — no
network. `run_ids` are the PR's active (non-completed) CI run ids, already
filtered to pull_request events on a non-main head by the shell that built them."""
number: int
labels: tuple[str, ...] = ()
is_draft: bool = False
created_at: str = ""
run_status: str = NONE
run_ids: tuple[int, ...] = ()
@classmethod
def from_json(cls, obj: dict) -> "PullRequest":
"""Build from a snapshot dict. `labels` may be a list of names or of gh's
label objects (`{"name": ...}`)."""
raw_labels = obj.get("labels") or []
names: list[str] = []
for lab in raw_labels:
if isinstance(lab, dict):
name = lab.get("name")
else:
name = lab
if name:
names.append(str(name))
status = (obj.get("runStatus") or obj.get("run_status") or NONE).lower()
if status not in (RUNNING, QUEUED, NONE):
status = NONE
raw_ids = obj.get("runIds") or obj.get("run_ids") or ()
return cls(
number=int(obj["number"]),
labels=tuple(names),
is_draft=bool(obj.get("isDraft") or obj.get("is_draft")
or obj.get("draft") or False),
created_at=str(obj.get("createdAt") or obj.get("created_at") or ""),
run_status=status,
run_ids=tuple(int(r) for r in raw_ids),
)
@dataclass(frozen=True)
class Blocker:
"""A PR that THIS PR must yield to in PASS 2 (purely informational for logging)."""
number: int
priority: int
kind: str # "higher-priority" | "same-level-ahead"
@dataclass(frozen=True)
class Decision:
"""The full point-in-time decision for THIS PR (used by --dry-run and tests)."""
self_number: int
self_priority: int
cancel_run_ids: tuple[int, ...] = ()
blockers: tuple[Blocker, ...] = ()
proceed: bool = True
# ── Pure decision core (no network / clock / subprocess) ─────────────────────
def effective_priority(pr: PullRequest) -> int:
"""Effective priority: `broken` OR `draft` => 10 (bottom, overriding any P0-P9);
else the lowest-numbered P0-P9 label present; else the default P5."""
if pr.is_draft or BROKEN_LABEL in pr.labels:
return BOTTOM_PRIORITY
nums = [int(m.group(1)) for name in pr.labels if (m := _P_LABEL.match(name))]
return min(nums) if nums else DEFAULT_PRIORITY
def priority_label(prio: int) -> str:
"""Human-readable priority for logs."""
if prio >= BOTTOM_PRIORITY:
return f"P{BOTTOM_PRIORITY} (broken/draft — bottom, below P9)"
return f"P{prio}"
def _ordering_key(pr: PullRequest) -> tuple[int, str, int]:
"""Same-level ordering (issue #342 rule 3): an in-flight (RUNNING) run keeps its
place at the front — a same-level peer never reorders it — then, among the PRs
still waiting to start (QUEUED or no run yet), OLDEST createdAt first (ascending),
then PR number as a stable final tiebreak so the order is fully deterministic."""
in_flight = 0 if pr.run_status == RUNNING else 1
return (in_flight, pr.created_at or _FAR_FUTURE, pr.number)
def runs_to_cancel(
this_pr: PullRequest,
all_prs: list[PullRequest],
*,
self_run_id: int | None = None,
) -> list[int]:
"""PASS 1. Run ids to cancel. A strictly-lower OTHER PR's active (running/queued)
run is cancelled iff keeping it running is wasteful, i.e. EITHER:
* THIS PR is P0 — an emergency reclaims every strictly-lower runner now; OR
* the target is broken/draft (effective priority 10) — its run can't merge /
isn't merge-ready, so ANY higher-priority PR may reclaim its runner.
P1-P9 never cancel a *normal* strictly-lower run — they yield in PASS 2 instead.
Invariants: never cancel self (by number or run id), never cancel an
equal-or-higher-priority PR (only strictly-lower, prio > self)."""
self_prio = effective_priority(this_pr)
to_cancel: list[int] = []
seen: set[int] = set()
for pr in all_prs:
if pr.number == this_pr.number:
continue # never cancel self
target_prio = effective_priority(pr)
if target_prio <= self_prio:
continue # only strictly-lower (skip equal-or-higher)
if pr.run_status not in ACTIVE:
continue # nothing running/queued to cancel
# Strictly lower: preemptible iff we're P0 OR the target is broken/draft
# (a bottom, priority-10, wasted run that any higher PR may reclaim).
if self_prio != TOP_PRIORITY and target_prio < BOTTOM_PRIORITY:
continue # P1-P9 don't bump a *normal* lower run
for rid in pr.run_ids:
if self_run_id is not None and rid == self_run_id:
continue # never cancel our own run
if rid in seen:
continue
seen.add(rid)
to_cancel.append(rid)
return to_cancel
def wait_blockers(this_pr: PullRequest, all_prs: list[PullRequest]) -> list[Blocker]:
"""PASS 2. The PRs THIS PR must yield to right now (empty => proceed). P0 never
yields. Otherwise yield to (a) any strictly-higher-priority OTHER PR with an
active/queued run, and (b) any SAME-priority PR ordered ahead of THIS PR
(running-first, then oldest createdAt)."""
self_prio = effective_priority(this_pr)
if self_prio == TOP_PRIORITY:
return [] # P0 outranks everything — never wait
others = [pr for pr in all_prs if pr.number != this_pr.number]
blockers: list[Blocker] = []
# (a) strictly-higher-priority PRs that actually have an active/queued run.
for pr in others:
p = effective_priority(pr)
if p < self_prio and pr.run_status in ACTIVE:
blockers.append(Blocker(pr.number, p, "higher-priority"))
# (b) same-level ordering: THIS PR proceeds only when it is at the front.
same_level = [pr for pr in others if effective_priority(pr) == self_prio]
same_level.append(this_pr) # this_pr appears exactly once
for pr in sorted(same_level, key=_ordering_key):
if pr.number == this_pr.number:
break # reached self => nobody ahead remains
blockers.append(Blocker(pr.number, self_prio, "same-level-ahead"))
return blockers
def decide(
this_pr: PullRequest,
all_prs: list[PullRequest],
*,
self_run_id: int | None = None,
) -> Decision:
"""Convenience: the full point-in-time decision (both passes) for THIS PR."""
cancels = runs_to_cancel(this_pr, all_prs, self_run_id=self_run_id)
blockers = wait_blockers(this_pr, all_prs)
return Decision(
self_number=this_pr.number,
self_priority=effective_priority(this_pr),
cancel_run_ids=tuple(cancels),
blockers=tuple(blockers),
proceed=not blockers,
)
# ── Pure TRIGGER-decision core (issue #349, --mode trigger) ──────────────────
def classify_sha_runs(runs: list[dict]) -> tuple[str, tuple[int, ...], bool]:
"""PURE. Summarise the CI runs on ONE head SHA. Returns (run_status, active_run_ids,
needy):
* run_status: RUNNING if any run is in progress, else QUEUED if any is queued/pending,
else NONE;
* active_run_ids: databaseIds of the non-completed (running/queued) runs;
* needy: True iff the SHA has NO active run AND NO run with a final VERDICT — i.e. its
required checks are absent/stale (a fresh SHA after a GITHUB_TOKEN auto-update) or
only cancelled/infra-aborted, so the PR cannot merge until CI is (re-)triggered on
that SHA. A success/failure/timeout verdict is NOT needy (green, or the author's to
fix — never auto-retried)."""
status = NONE
active_ids: list[int] = []
has_verdict = False
for r in runs:
raw = (r.get("status") or "").lower()
if raw == "completed":
if (r.get("conclusion") or "").lower() in VERDICT_CONCLUSIONS:
has_verdict = True
continue
norm = _normalise_status(raw) # in_progress -> running; else queued
if _STATUS_RANK[norm] < _STATUS_RANK[status]:
status = norm
rid = r.get("databaseId")
if rid is not None:
active_ids.append(int(rid))
needy = not active_ids and not has_verdict
return status, tuple(active_ids), needy
def _trigger_order_key(pr: PullRequest) -> tuple[int, str, int]:
"""Trigger ordering: highest priority first (lowest effective-priority number), then
OLDEST createdAt first (the longest-waiting PR at a level goes first — the same-level
fairness / anti-starvation rule), then PR number as a stable final tiebreak."""
return (effective_priority(pr), pr.created_at or _FAR_FUTURE, pr.number)
@dataclass(frozen=True)
class TriggerDecision:
"""PURE output of `select_triggers`: which PR(s) the scheduler should (re-)trigger CI
for right now, in order, plus any strictly-lower runs a P0 emergency preempts to free a
runner. Exercised by `--mode trigger --dry-run` and the unit tests."""
trigger_numbers: tuple[int, ...] = ()
cancel_run_ids: tuple[int, ...] = ()
inflight_numbers: tuple[int, ...] = ()
needy_numbers: tuple[int, ...] = ()
skipped_fork_numbers: tuple[int, ...] = ()
slots: int = 0
max_inflight: int = 0
def select_triggers(
all_prs: list[PullRequest],
needy: "set[int] | frozenset[int]",
*,
max_inflight: int,
forks: "set[int] | frozenset[int]" = frozenset(),
) -> TriggerDecision:
"""PURE. Choose the PR(s) to (re-)trigger CI for now — a poor-man's merge queue over the
existing priority model. No network / clock / subprocess, so it is exhaustively unit-
tested (see TestSelectTriggers / TestTriggerStarvation).
* inflight = PRs already running/queued on their head SHA — they occupy the cap.
* candidates = NEEDY PRs (absent/stale checks on their head SHA) that are not already
running and are not forks (forks have no token/secret access — see `run_trigger`).
* order = effective priority, then oldest createdAt, then number (`_trigger_order_key`).
* P0 = EMERGENCY: always triggered, BYPASSING the cap, and it PREEMPTS its strictly-lower
OTHER runs (reusing `runs_to_cancel`) so a runner frees for it immediately.
* P1–P10 fill only the remaining ``slots = max_inflight - len(inflight)``; the rest wait
for a later pass.
STARVATION is bounded, not by aging but structurally: triggering a PR gives its head SHA
a run, so it LEAVES the needy set; between merges the needy set only shrinks, and the
scheduler re-runs on every auto-update plus a cron backstop, so every eligible PR is
triggered within a bounded number of passes (proved by TestTriggerStarvation). Ordering
is still by priority, so higher-priority PRs are simply served first, never exclusively
forever (a served PR stops being needy until its next push/auto-update)."""
cap = max(1, max_inflight)
inflight = [p for p in all_prs if p.run_status in ACTIVE]
candidates = [
p for p in all_prs
if p.number in needy and p.run_status not in ACTIVE and p.number not in forks
]
ordered = sorted(candidates, key=_trigger_order_key)
emergencies = [p for p in ordered if effective_priority(p) == TOP_PRIORITY]
normal = [p for p in ordered if effective_priority(p) != TOP_PRIORITY]
slots = max(0, cap - len(inflight))
chosen = emergencies + normal[:slots] # P0 bypasses the cap; P1–P10 fill free slots
cancel_ids: list[int] = []
seen: set[int] = set()
for emergency in emergencies: # P0 preempts its strictly-lower active runs
for rid in runs_to_cancel(emergency, all_prs):
if rid not in seen:
seen.add(rid)
cancel_ids.append(rid)
return TriggerDecision(
trigger_numbers=tuple(p.number for p in chosen),
cancel_run_ids=tuple(cancel_ids),
inflight_numbers=tuple(sorted(p.number for p in inflight)),
needy_numbers=tuple(p.number for p in ordered),
skipped_fork_numbers=tuple(sorted(n for n in needy if n in forks)),
slots=slots,
max_inflight=cap,
)
# ── gh I/O shell (the only part that touches the network) ────────────────────
def _log(msg: str) -> None:
print(msg, flush=True)
def _gh_json(args: list[str]) -> list | dict | None:
"""Run `gh <args> --json ...` and parse stdout as JSON. Returns None (never
raises) on any failure — the caller fails open."""
try:
proc = subprocess.run(
["gh", *args],
capture_output=True,
text=True,
check=False,
)
except (OSError, ValueError) as exc:
_log(f"::warning::gh invocation failed ({' '.join(args[:2])}): {exc}")
return None
if proc.returncode != 0:
_log(f"::warning::gh exited {proc.returncode} ({' '.join(args[:2])}): "
f"{proc.stderr.strip()}")
return None
try:
return json.loads(proc.stdout or "null")
except json.JSONDecodeError as exc:
_log(f"::warning::could not parse gh JSON ({' '.join(args[:2])}): {exc}")
return None
def _normalise_status(raw: str) -> str:
"""Map a gh run status onto our RUNNING / QUEUED / NONE model."""
if raw == "in_progress":
return RUNNING
if raw == "completed":
return NONE
return QUEUED # queued / waiting / requested / pending
def _runs_by_head(limit: int = 300) -> dict[str, dict]:
"""One bulk `gh run list` -> {headBranch: {"status", "ids"}} for active PR-event
runs. Enforces the 'never cancel main/push' invariant at the source: only
`event == pull_request`, non-`main`, non-completed runs are kept. Active runs are
the most recent, so `limit` most-recent runs comfortably covers them."""
rows = _gh_json([
"run", "list", "--workflow", "ci.yml", "--event", "pull_request",
"--limit", str(limit),
"--json", "databaseId,status,headBranch,event",
])
by_head: dict[str, dict] = {}
for row in rows or []:
if row.get("event") != "pull_request":
continue
head = row.get("headBranch")
if not head or head == "main":
continue
if row.get("status") == "completed":
continue
entry = by_head.setdefault(head, {"status": NONE, "ids": []})
entry["ids"].append(int(row["databaseId"]))
status = _normalise_status(row.get("status", ""))
# running beats queued beats none for the branch's aggregate status.
if _STATUS_RANK[status] < _STATUS_RANK[entry["status"]]:
entry["status"] = status
return by_head
def gather_snapshot(self_pr_number: int) -> tuple[PullRequest | None, list[PullRequest]]:
"""Build (this_pr, all_prs) from live gh data. this_pr is forced to RUNNING —
by definition our own run is in progress while this job executes."""
prs = _gh_json([
"pr", "list", "--state", "open", "--limit", "300",
"--json", "number,headRefName,labels,isDraft,createdAt",
])
if prs is None:
return None, []
runs = _runs_by_head()
all_prs: list[PullRequest] = []
this_pr: PullRequest | None = None
for obj in prs:
head = obj.get("headRefName") or ""
run_info = runs.get(head, {"status": NONE, "ids": []})
number = int(obj["number"])
is_self = number == self_pr_number
pr = PullRequest.from_json({
**obj,
# self is definitionally running (this job is in progress).
"runStatus": RUNNING if is_self else run_info["status"],
"runIds": run_info["ids"],
})
all_prs.append(pr)
if is_self:
this_pr = pr
return this_pr, all_prs
def _cancel_run(run_id: int) -> bool:
try:
proc = subprocess.run(
["gh", "run", "cancel", str(run_id)],
capture_output=True, text=True, check=False,
)
except (OSError, ValueError) as exc:
_log(f"::warning::could not cancel run {run_id}: {exc}")
return False
if proc.returncode == 0:
return True
_log(f"::warning::could not cancel run {run_id} — likely already finished. "
f"{proc.stderr.strip()}")
return False
def _positive_int(env_name: str, default: int) -> int:
raw = os.environ.get(env_name, "")
return int(raw) if raw.isdigit() and int(raw) > 0 else default
def run_live() -> int:
"""The gh-driven shell: gather, PASS 1 (cancel), PASS 2 (bounded hold-back). Always
returns 0 — the traffic-controller must never fail CI."""
event = os.environ.get("GITHUB_EVENT_NAME", "")
self_raw = os.environ.get("SELF_PR", "")
if event != "pull_request" or not self_raw.isdigit():
_log("Not a pull_request event (or no PR number) — nothing to do.")
return 0
self_number = int(self_raw)
self_run_id = int(os.environ["GITHUB_RUN_ID"]) if os.environ.get(
"GITHUB_RUN_ID", "").isdigit() else None
this_pr, all_prs = gather_snapshot(self_number)
if this_pr is None:
_log("::warning::Could not resolve THIS PR from the open-PR list — skipping.")
return 0
self_prio = effective_priority(this_pr)
_log(f"This PR #{self_number} effective priority: {priority_label(self_prio)} "
"(P0 = highest/emergency, P9 = lowest, broken/draft = bottom).")
# ── PASS 1: PREEMPTION (P0 reclaims all lower; anyone reclaims broken/draft) ──
to_cancel = runs_to_cancel(this_pr, all_prs, self_run_id=self_run_id)
if not to_cancel:
if self_prio == TOP_PRIORITY:
_log("P0 emergency — no strictly-lower active runs to cancel.")
else:
_log("No preemptible runs (P1-P9 only reclaim broken/draft lower runs; "
"none active).")
else:
cancelled = 0
for rid in to_cancel:
if _cancel_run(rid):
_log(f" cancelled run {rid} (freed its runner).")
cancelled += 1
_log(f"P0 preemption complete — cancelled {cancelled}/{len(to_cancel)} run(s).")
# ── PASS 2: BOUNDED HOLD-BACK (yield to higher / same-level-ahead) ────────
if self_prio == TOP_PRIORITY:
_log("P0 emergency — not yielding; proceeding immediately.")
return 0
budget = _positive_int("HOLD_BACK_BUDGET_SECONDS", 180)
poll = _positive_int("HOLD_BACK_POLL_SECONDS", 15)
deadline = time.monotonic() + budget
_log(f"{priority_label(self_prio)} — holding back up to {budget}s for higher / "
"earlier same-level PRs (no cancellation).")
while True:
remaining = deadline - time.monotonic()
if remaining <= 0:
_log("Hold-back budget elapsed — proceeding; higher-priority PRs got their "
"head start.")
break
# Refresh OTHER PRs so newly-opened higher-priority PRs are seen mid-wait;
# THIS PR's own identity/priority stays fixed (matching the original).
_, fresh = gather_snapshot(self_number)
if not fresh:
_log("::warning::Could not refresh open PRs — proceeding.")
break
blockers = wait_blockers(this_pr, fresh)
if not blockers:
_log("No higher-priority or earlier same-level PR is ahead — proceeding.")
break
tags = " ".join(f"#{b.number}(P{b.priority},{b.kind})" for b in blockers)
sleep_s = min(poll, int(remaining)) if remaining >= 1 else 0
_log(f"Yielding to: {tags} — re-checking in {sleep_s}s "
f"({int(remaining)}s budget left).")
if sleep_s > 0:
time.sleep(sleep_s)
_log("Hold-back complete — this PR's heavy jobs may now start.")
return 0
# ── --dry-run: feed the pure core a snapshot JSON, print its decisions ───────
def _load_snapshot(text: str) -> tuple[PullRequest, list[PullRequest], int | None]:
data = json.loads(text)
all_prs = [PullRequest.from_json(o) for o in data.get("prs", [])]
self_number = int(data["self"])
self_run_id = data.get("self_run_id")
self_run_id = int(self_run_id) if self_run_id is not None else None
this_pr = next((p for p in all_prs if p.number == self_number), None)
if this_pr is None:
raise ValueError(f"self #{self_number} not present in prs[]")
return this_pr, all_prs, self_run_id
def run_dry(text: str) -> int:
this_pr, all_prs, self_run_id = _load_snapshot(text)
dec = decide(this_pr, all_prs, self_run_id=self_run_id)
_log(f"This PR #{dec.self_number} effective priority: "
f"{priority_label(dec.self_priority)}")
if dec.cancel_run_ids:
why = ("P0 emergency (reclaims all strictly-lower)"
if dec.self_priority == TOP_PRIORITY
else "reclaiming broken/draft lower runs")
_log(f"PASS 1 (preemption): {why} — cancel run ids: "
f"{list(dec.cancel_run_ids)}")
else:
_log("PASS 1 (preemption): nothing to cancel.")
if dec.proceed:
_log("PASS 2 (hold-back): PROCEED — no blockers.")
else:
tags = ", ".join(f"#{b.number}(P{b.priority}, {b.kind})" for b in dec.blockers)
_log(f"PASS 2 (hold-back): WAIT — yielding to: {tags}")
return 0
# ── --mode trigger: the PAT-free scheduler shell (issue #349) ────────────────
def _gh_ok(args: list[str]) -> bool:
"""Run `gh <args>` for its side effect (no JSON parse). Returns True on exit 0; never
raises — the scheduler fails open on any I/O error."""
try:
proc = subprocess.run(["gh", *args], capture_output=True, text=True, check=False)
except (OSError, ValueError) as exc:
_log(f"::warning::gh invocation failed ({' '.join(args[:2])}): {exc}")
return False
if proc.returncode != 0:
_log(f"::warning::gh exited {proc.returncode} ({' '.join(args[:3])}): "
f"{proc.stderr.strip()}")
return False
return True
def _ci_workflow_file() -> str:
return os.environ.get("CI_WORKFLOW_FILE", "ci.yml")
def gather_trigger_snapshot() -> tuple[list[PullRequest], set[int], set[int], dict[int, dict]]:
"""Build (all_prs, needy, forks, meta) from live gh data for the trigger scheduler.
* all_prs: PullRequest snapshots whose run_status / run_ids reflect the runs on each
PR's CURRENT head SHA (so 'inflight' means a live run on the mergeable SHA, never a
stale one on a superseded SHA);
* needy: PR numbers whose head SHA has absent/stale checks (must be (re-)triggered);
* forks: cross-repository PR numbers — no token/secret access, so NOT token-triggerable;
* meta: number -> {headRefName, headRefOid} for the dispatch I/O.
Returns empty structures (never raises) if gh can't be reached — the caller fails open."""
prs = _gh_json([
"pr", "list", "--state", "open", "--limit", "300",
"--json", "number,headRefName,headRefOid,isCrossRepository,labels,isDraft,createdAt",
])
if prs is None:
return [], set(), set(), {}
runs = _gh_json([
"run", "list", "--workflow", _ci_workflow_file(), "--limit", "300",
"--json", "databaseId,status,conclusion,headSha,headBranch,event",
]) or []
by_sha: dict[str, list[dict]] = {}
for row in runs:
sha = row.get("headSha")
if sha:
by_sha.setdefault(sha, []).append(row)
all_prs: list[PullRequest] = []
needy: set[int] = set()
forks: set[int] = set()
meta: dict[int, dict] = {}
for obj in prs:
number = int(obj["number"])
sha = obj.get("headRefOid") or ""
status, run_ids, is_needy = classify_sha_runs(by_sha.get(sha, []))
all_prs.append(PullRequest.from_json(
{**obj, "runStatus": status, "runIds": list(run_ids)}))
meta[number] = {
"headRefName": obj.get("headRefName") or "",
"headRefOid": sha,
}
if obj.get("isCrossRepository"):
forks.add(number)
if is_needy:
needy.add(number)
return all_prs, needy, forks, meta
def _dispatch_ci(pr_number: int, head_ref: str, head_sha: str) -> bool:
"""Trigger `ci.yml` for one PR via a `workflow_dispatch` on the PR's head branch. The
dispatch runs as GH_TOKEN, which ci-trigger.yml sets to the AUTOUPDATE_TOKEN PAT: a run
triggered by the built-in GITHUB_TOKEN is held for MANUAL approval (`action_required`) and
never runs un-attended, so the PAT (authorized owner) is what actually starts the run with no
approval gate (see issue #351). Running on the head branch puts the run's checks on the PR
head SHA, so they satisfy branch protection's required checks."""
if not head_ref:
_log(f"::warning::PR #{pr_number} has no head branch — cannot dispatch; skipping.")
return False
ok = _gh_ok([
"workflow", "run", _ci_workflow_file(), "--ref", head_ref,
"-f", f"pr={pr_number}",
"-f", f"head_sha={head_sha}",
"-f", "reason=traffic-controller",
])
if ok:
short = head_sha[:8] if head_sha else "?"
_log(f" triggered CI for #{pr_number} on {head_ref} (head {short}).")
return ok
def run_trigger() -> int:
"""The scheduler shell (companion `ci-trigger.yml`): pick the highest-priority needy
PR(s) within the inflight cap and (re-)trigger their CI via workflow_dispatch; a P0
emergency additionally preempts its strictly-lower runs. ALWAYS returns 0 — the scheduler
must never wedge CI, and structurally it cannot: `ci.yml` keeps `on: pull_request`, so any
human push (and a brand-new PR) still gets CI independently of this scheduler."""
max_inflight = _positive_int("MAX_INFLIGHT_RUNS", 2)
all_prs, needy, forks, meta = gather_trigger_snapshot()
if not all_prs:
_log("No open PRs (or could not list them) — nothing to trigger.")
return 0
dec = select_triggers(all_prs, needy, max_inflight=max_inflight, forks=forks)
_log(f"Open PRs: {len(all_prs)} | needy (absent/stale checks): {list(dec.needy_numbers)} "
f"| inflight: {list(dec.inflight_numbers)} | cap {dec.max_inflight}, "
f"free slots {dec.slots}.")
if dec.skipped_fork_numbers:
_log(f"Fork PR(s) needing CI left to `on: pull_request` (no token access — not "
f"wedged): {list(dec.skipped_fork_numbers)}.")
for rid in dec.cancel_run_ids: # P0 emergency preemption
if _cancel_run(rid):
_log(f" P0 preemption: cancelled lower run {rid} (freed its runner).")
if not dec.trigger_numbers:
_log("Nothing to trigger this pass (no needy PR fits a free slot).")
return 0
triggered = 0
for number in dec.trigger_numbers:
info = meta.get(number, {})
if _dispatch_ci(number, info.get("headRefName", ""), info.get("headRefOid", "")):
triggered += 1
_log(f"Trigger pass complete — dispatched {triggered}/{len(dec.trigger_numbers)} "
"run(s) in priority order.")
return 0
def _load_trigger_snapshot(
text: str,
) -> tuple[list[PullRequest], set[int], set[int], int]:
data = json.loads(text)
all_prs = [PullRequest.from_json(o) for o in data.get("prs", [])]
needy = {int(n) for n in data.get("needy", [])}
forks = {int(n) for n in data.get("forks", [])}
max_inflight = int(data.get("max_inflight", 2))
return all_prs, needy, forks, max_inflight
def run_trigger_dry(text: str) -> int:
all_prs, needy, forks, max_inflight = _load_trigger_snapshot(text)
dec = select_triggers(all_prs, needy, max_inflight=max_inflight, forks=forks)
_log(f"Trigger decision (cap {dec.max_inflight}, free slots {dec.slots}):")
_log(f" inflight (occupying the cap): {list(dec.inflight_numbers)}")
_log(f" needy candidates (priority order): {list(dec.needy_numbers)}")
if dec.skipped_fork_numbers:
_log(f" skipped forks (no token access): {list(dec.skipped_fork_numbers)}")
if dec.cancel_run_ids:
_log(f" P0 preemption — cancel run ids: {list(dec.cancel_run_ids)}")
_log(f" => TRIGGER (in priority order): {list(dec.trigger_numbers)}")
return 0
def main(argv: list[str] | None = None) -> int:
# Emit UTF-8 regardless of the host console so the log typography is stable on
# the UTF-8 CI runners (and never raises on a legacy Windows code page).
try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
except (AttributeError, ValueError):
pass
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument(
"--mode", choices=("orchestrate", "trigger"), default="orchestrate",
help="orchestrate (default): in-run runner-priority for the executing PR "
"(unchanged). trigger: the scheduler that (re-)triggers CI by priority "
"(issue #349).")
parser.add_argument(
"--dry-run", action="store_true",
help="read a snapshot JSON (from --input or stdin), print decisions, no network.")
parser.add_argument(
"--input", help="snapshot JSON file for --dry-run (default: stdin).")
args = parser.parse_args(argv)
if args.dry_run:
text = (open(args.input, encoding="utf-8").read() if args.input
else sys.stdin.read())
return run_trigger_dry(text) if args.mode == "trigger" else run_dry(text)
return run_trigger() if args.mode == "trigger" else run_live()
if __name__ == "__main__":
try:
sys.exit(main())
except Exception as exc: # never let the controller fail CI
_log(f"::warning::traffic_control crashed, proceeding fail-open: {exc}")
sys.exit(0)
+153
View File
@@ -0,0 +1,153 @@
<!-- SPDX-License-Identifier: GPL-3.0-or-later -->
# GitHub Actions workflows
This directory holds the repo's workflows:
- **`ci.yml`** — the pull-request gate: build, unit tests, static analysis (ktlint /
detekt), and the E2E/instrumented-test matrix, aggregated into one `CI passed` check
that branch protection requires. It also runs the `traffic-control` job described
below.
- **`autoupdate.yml`** — rebases every open PR onto `main` whenever `main` advances, so
the "branches up to date" branch rule never needs a manual update.
- **`release.yml`** — turns a pushed version tag into signed, published release
artifacts; see [`docs/release.md`](../../docs/release.md).
The rest of this README is about **`traffic-control`** — the job (in the Checks tab it
shows up as **"Traffic control (runner priority)"**) that decides whose CI gets to run
first when several PRs are queued at once.
## Why this job exists
GitHub Actions has no concept of "run this PR's checks before that one" — every PR's
workflow run joins the same pool of runners and is served roughly first-come,
first-served. That's fine most of the time, but with several PRs open at once it means
an urgent one-line hotfix queues up as an equal to a routine refactor, and can end up
stuck waiting behind CI runs for changes that aren't in any hurry.
`traffic-control` addresses that by reading a **priority label** on the current PR,
comparing it against every other open PR, and then either freeing up a runner by
cancelling a lower-priority PR's run (**preemption**), or briefly waiting before this
PR's own heavy jobs start so a higher-priority PR's jobs get a head start
(**hold-back**). It runs first in every PR's CI: every other job in `ci.yml`
(`debug-build`, `unit-tests`, `static-analysis`, `e2e`, `e2e-preview`) declares
`needs: traffic-control`, so it always goes first —
```
PR's CI run starts
│
▼
traffic-control
│ 1. compute this PR's effective priority (see table below)
│ 2. PASS 1 — preemption: cancel strictly-lower-priority OTHER PRs' active
│ runs, but only if we're P0, or the target PR is `broken`
│ 3. PASS 2 — hold-back: if we're not P0, wait (up to 180s) while any
│ strictly-higher-priority OTHER PR still has an active run, then
│ proceed regardless
▼
debug-build · unit-tests · static-analysis · e2e · e2e-preview
```
## Effective priority
Priority comes from a label on the PR:
| Label | Effective priority | Meaning |
| --- | --- | --- |
| `P0` | 0 (highest) | **Emergency only** — production is broken, or an emergency security fix. |
| `P1` – `P9` | 1 – 9 | Higher number = lower priority. |
| *(no `P` label)* | 5 (default) | Normal priority — most PRs. |
| `broken` | 10 (lowest) | A stuck/failing PR, deprioritised below even `P9`. Overrides any `P0`–`P9` label also present. |
Apply at most one `P0`–`P9` label; if more than one is somehow present, the numerically
lowest (most urgent) one wins. The `broken` label is meant to be applied by a maintainer
to a PR whose CI is stuck or failing, as a "let everyone else go first while this gets
fixed" signal — not something a PR author sets on their own work. Removing it restores
whatever `P0`–`P9` priority (or the `P5` default) the PR would otherwise have.
## Preemption vs. holding back
### P0 preempts everyone lower
If *this* PR is `P0`, it's treated as an emergency: the job immediately cancels the
in-progress or queued CI runs of **every other open PR at a strictly lower priority**
(that is, anything that isn't also `P0`), freeing up their runners right away. A PR
that gets cancelled this way isn't harmed long-term — it simply reruns on its next push,
or the next time `autoupdate.yml` rebases it onto `main`. Because nothing outranks an
emergency, a `P0` PR also never does the hold-back wait described below.
### A `broken` PR can be preempted by anyone
A PR labelled `broken` can't merge while it's broken, so its CI run occupying a runner
is wasted capacity. Any PR that isn't itself `broken` — in other words, any PR with a
real `P0`–`P9` priority — outranks it and may cancel its active run to reclaim the
runner, not just a `P0` PR. `broken` is also the only priority level that yields to
*everything*: since it sits below every other level, it always waits for other PRs'
runs rather than the other way around.
### P1–P9 yield, but never cancel
Every other level (`P1`–`P9`, including the `P5` default) is cooperative rather than
aggressive: it never cancels a run that's already going, no matter how much lower that
run's priority is. Instead, before letting its own heavy jobs start, it checks whether
any **strictly higher**-priority PR currently has an active or queued CI run. If so, it
waits — polling every 15 seconds and re-checking the full list of open PRs each time, so
a newly opened higher-priority PR is picked up mid-wait too — giving that PR's jobs a
chance to reach the runner queue first. The wait is capped at **180 seconds**
(comfortably inside the job's 6-minute hard timeout); once the budget runs out, this PR
proceeds regardless. A PR should never be able to block itself indefinitely.
## Safety invariants
Whatever the priority math says, a few things are hard-coded to never happen:
- **Never touches `main` / push-triggered runs.** The job only acts on `pull_request`
events, and every run it's even allowed to consider cancelling is filtered down to
`event == pull_request` with `headBranch != main`.
- **Never cancels this PR's own run.** The current PR is excluded from the "other PRs"
list up front by PR number, and the currently-executing run ID is skipped too, just in
case.
- **Never cancels an equal-or-higher-priority run.** Only strictly-lower-priority PRs
(a numerically larger, i.e. worse, priority) are ever candidates for cancellation.
## Honest limitation
This is a **best-effort head start, not a real priority queue.** GitHub Actions has no
API for "give this run's jobs priority over that run's jobs" — runners are handed out
roughly FIFO no matter what this job does. Hold-back approximates priority by making
lower-priority PRs wait a little before their jobs even enter that FIFO queue, but under
sustained contention (many PRs queuing at once) the bounded wait can run out before a
higher-priority PR's jobs have actually made it through the runner pool. The waiting job
itself is cheap and short-lived, which is exactly why the wait is capped rather than
open-ended — occasionally under-prioritizing is preferable to a job that ties up a
runner indefinitely just to wait.
## Not a merge gate
`traffic-control` is an optimizer, not a check your PR needs to pass. It's deliberately
left out of `ci-passed`'s `needs:` list, every GitHub API call it makes is guarded
against failure, the script always exits `0`, and the step itself runs with
`continue-on-error: true`. A hiccup here — a transient API error, a missing permission,
a fork PR without write access — can never fail or block your PR.
That said, the heavy jobs still order themselves after it via `needs: traffic-control`,
so if this job were ever skipped or failed outright, GitHub would mark those jobs
`skipped` — and `ci-passed` treats a required job coming back `skipped` as a gate
failure. So the worst case is fail-safe: it blocks the merge rather than letting an
untested PR through.
It also needs very little to run: no checkout step (it only calls the `gh` CLI), and
just two permissions (`actions: write` to cancel runs, `pull-requests: read` to read
labels). Values that come from outside the repo — labels, branch names — are only ever
read through `gh`'s JSON output into shell variables, never interpolated as shell code.
## Where this is heading
**#342** is rewriting this logic as a tested Python module
(`.github/scripts/traffic_control.py`), with a couple of small behavior refinements:
draft PRs will also sink to the bottom (like `broken`), and PRs at the exact same
priority level get an explicit order (whichever run is already in flight finishes
first; among the rest, whoever has been waiting longest goes next). This README
describes the shell-script version currently in `ci.yml` — see the comment block above
the `traffic-control` job there for the byte-for-byte spec — and will be updated once
#342 lands.
+17 -8
View File
@@ -6,13 +6,20 @@ name: Auto-update PR branches
# touched (PR_FILTER: all) — this is no longer limited to PRs with GitHub auto-merge
# enabled.
#
# IMPORTANT: for the branch update to RE-TRIGGER the PR's CI (so it can pass and merge),
# this must run with a PAT, not the default GITHUB_TOKEN — pushes made by GITHUB_TOKEN do
# not start new workflow runs (GitHub's anti-recursion rule), so the updated PR would sit
# with stale checks. Create a fine-grained PAT scoped to this repo with
# contents:read/write + pull-requests:read/write and add it as the AUTOUPDATE_TOKEN secret.
# Without it this falls back to GITHUB_TOKEN, which updates the branch but will NOT re-run
# the PR's checks.
# IMPORTANT (issue #349): the branch update runs with the default GITHUB_TOKEN — ON PURPOSE.
# A GITHUB_TOKEN push does NOT start new workflow runs (GitHub's anti-recursion rule), so
# updating every behind PR here NO LONGER re-triggers every PR's CI. That deliberately breaks
# the old merge-cascade (every merge -> autoupdate rebases all PRs with a PAT -> all re-run ->
# ci.yml's cancel-in-progress kills each in-flight run -> PRs thrash and can't converge).
# Branches still go up to date (satisfying "require branches up to date"); they just don't
# auto-run CI on the new head SHA. Re-triggering that SHA's CI is now OWNED by the traffic-
# controller scheduler (`.github/workflows/ci-trigger.yml` -> `traffic_control.py --mode
# trigger`), which triggers the updated PRs deliberately, in priority order, a few at a time.
# So this workflow must NOT use the PAT for the update push (that would re-introduce the
# cascade). This workflow itself doesn't need AUTOUPDATE_TOKEN — but the secret is still REQUIRED
# by the repo: the ci-trigger.yml scheduler dispatches CI with it (a GITHUB_TOKEN dispatch would
# be held for manual approval and never run un-attended). Don't delete the secret. See
# ci-trigger.yml + issue #351.
on:
push:
@@ -38,7 +45,9 @@ jobs:
- name: Update all behind PRs
uses: chinthakagodawita/autoupdate@0707656cd062a3b0cf8fa9b2cda1d1404d74437e # v1.7.0
env:
GITHUB_TOKEN: ${{ secrets.AUTOUPDATE_TOKEN || secrets.GITHUB_TOKEN }}
# Default GITHUB_TOKEN — NOT a PAT — so this update push does not auto-retrigger CI
# (anti-recursion). See the header: re-triggering is owned by ci-trigger.yml.
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
PR_FILTER: "all"
PR_READY_STATE: "all"
MERGE_CONFLICT_ACTION: "ignore"
+100
View File
@@ -0,0 +1,100 @@
# SPDX-License-Identifier: GPL-3.0-or-later
name: CI trigger (traffic-controller)
# The traffic-controller SCHEDULER (issue #349). It OWNS CI *triggering*. After main advances,
# autoupdate.yml updates every behind PR's branch with the built-in GITHUB_TOKEN which, by
# GitHub's anti-recursion rule, does NOT start CI — so those PRs sit with absent/stale required
# checks on their new head SHA and cannot merge. This workflow then (re-)triggers CI for the
# highest-priority such PR(s), a few at a time (an inflight cap), in the existing P0–P9 /
# broken-draft priority order — a poor-man's merge queue that replaces the old "every merge
# re-runs every PR" thundering herd (the cascade; see the ci-merge-cascade note + issue #349).
#
# HOW IT TRIGGERS: `traffic_control.py --mode trigger` runs `gh workflow run ci.yml --ref
# <pr-head-branch>`, dispatching with the AUTOUPDATE_TOKEN PAT — NOT the built-in GITHUB_TOKEN.
# A workflow run triggered by GITHUB_TOKEN is held in the `action_required` state waiting on
# MANUAL approval and never runs un-attended (confirmed empirically on #285 / #350: it sits
# `action_required`, while the same dispatch by an authorized user runs immediately) — which
# would defeat the whole scheduler. A PAT dispatch runs AS the authorized token owner, so the
# run starts immediately with no approval gate (this is the original #349 design; #350's "no
# PAT needed / workflow_dispatch is anti-recursion-exempt" claim was WRONG — see #351).
# AUTOUPDATE_TOKEN is therefore REQUIRED for this scheduler. The dispatched run executes on the
# PR's head branch, so its checks land on the PR head SHA and satisfy branch protection.
#
# WHEN IT RUNS:
# • workflow_run, after "Auto-update PR branches" completes — the race-free moment: autoupdate
# has finished moving branches to their new (checkless) head SHAs, so this pass sees exactly
# the PRs that now need a run. (A bare `push: main` trigger would race autoupdate and often
# read the pre-update SHAs, missing them until the next pass.)
# • schedule (cron) — a backstop so no PR is ever permanently un-triggered even if a
# workflow_run is missed/skipped (part of the fail-open guarantee), and so a brand-new PR
# whose first `on: pull_request` run got cancelled is still picked up.
# • workflow_dispatch — manual kick.
#
# FAIL-OPEN: the script guards every gh call and always exits 0; and structurally, ci.yml keeps
# its `on: pull_request` trigger, so a human push (and a brand-new PR) always triggers CI
# regardless of this scheduler — CI can never become permanently un-triggerable. Fork PRs (no
# token/secret access) are skipped here and left to `on: pull_request`, so they are never wedged.
on:
workflow_run:
workflows: ["Auto-update PR branches"]
types: [completed]
schedule:
# Backstop cadence (UTC). GitHub may delay scheduled runs under load; that is fine — this
# is only a safety net behind the immediate workflow_run trigger above.
- cron: "*/15 * * * *"
workflow_dispatch:
# Trigger-only; this workflow never gates a merge. The gh calls run as GH_TOKEN, which is
# normally the AUTOUPDATE_TOKEN PAT (see the step below). These permissions govern the built-in
# GITHUB_TOKEN, used only on the fail-open fallback path when AUTOUPDATE_TOKEN is absent:
# `actions: write` lets it dispatch ci.yml (workflow_dispatch) and cancel strictly-lower runs
# when a P0 emergency preempts; `pull-requests: read` + `contents: read` cover the PR/label
# enumeration. (A GITHUB_TOKEN dispatch needs manual approval, so that fallback only actually
# starts CI if repo settings don't gate GITHUB_TOKEN-triggered runs — the PAT is the real path.)
permissions:
contents: read
pull-requests: read
actions: write
# One trigger pass at a time. Do NOT cancel an in-flight pass (cancel-in-progress: false):
# a half-finished pass could leave some needy PRs un-triggered until the next pass.
concurrency:
group: ci-trigger
cancel-in-progress: false
jobs:
trigger:
name: Trigger CI by priority
runs-on: ubuntu-latest
timeout-minutes: 10 # generous backstop; the script only enumerates + dispatches, no waits
steps:
- name: Check out source
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
- name: Set up Python
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6.3.0
with:
python-version: "3.x"
# gh is auto-configured from GH_TOKEN / GH_REPO. The script guards every gh call and
# always exits 0, so a hiccup (API error, missing permission, fork PR) can never wedge CI
# — and even a total failure here leaves ci.yml's `on: pull_request` path intact.
- name: Trigger CI for the highest-priority PR(s) needing a run
env:
# AUTOUPDATE_TOKEN (a PAT) is REQUIRED here: a CI run dispatched by the built-in
# GITHUB_TOKEN is held for MANUAL approval (`action_required`) and never runs
# un-attended, so the scheduler must dispatch AS the PAT's authorized owner to start
# runs with no approval gate. `|| github.token` keeps this fail-open when the secret is
# absent, but that GITHUB_TOKEN fallback only actually starts CI if repo settings don't
# gate GITHUB_TOKEN-triggered runs — the PAT is the intended path (see #351).
GH_TOKEN: ${{ secrets.AUTOUPDATE_TOKEN || github.token }}
GH_REPO: ${{ github.repository }}
# Poor-man's merge-queue width: at most this many PRs run CI concurrently under the
# scheduler (a P0 emergency bypasses this cap). Kept conservative because each PR
# fans out to the whole E2E matrix (~8 API levels + preview); this is the main knob
# to raise for throughput vs runner budget. The coordinator drives runner allocation.
MAX_INFLIGHT_RUNS: "2"
# The workflow file the scheduler enumerates runs for and dispatches.
CI_WORKFLOW_FILE: "ci.yml"
run: python3 .github/scripts/traffic_control.py --mode trigger
+611 -281
View File
File diff suppressed because it is too large Load Diff
+100
View File
@@ -0,0 +1,100 @@
# SPDX-License-Identifier: GPL-3.0-or-later
# Extracted verbatim from .github/workflows/ci.yml (where it used to run first and gate the
# heavy jobs via `needs: traffic-control`). It is being mothballed pending a rebuild as a
# published GitHub Action and will be disabled after this merges; the heavy CI jobs no longer
# depend on it. The decision core it drives (.github/scripts/traffic_control.py) is unchanged
# and still unit-tested by the `traffic-control-tests` job in ci.yml.
name: Traffic control (runner priority)
# The job reads github.event.pull_request.number, so it needs PR context.
on: pull_request
jobs:
# ── Priority-based runner orchestration ─────────────────────────────
# Runs FIRST (the heavy jobs below all `needs: traffic-control`). It reads THIS
# PR's P0–P9 label, `broken` label, and draft state to order runner access. The
# decision logic lives in .github/scripts/traffic_control.py — a pure, unit-tested
# core (see .github/scripts/test_traffic_control.py) plus a thin gh-I/O shell; this
# step just checks out the repo and runs it.
#
# Effective priority: a `broken` OR `draft` PR => 10 (BOTTOM, below P9), overriding
# any P0–P9; else the lowest-numbered P0–P9 label present (P0 = highest); else P5.
#
# • P0 = EMERGENCY ONLY (app broken in production / emergency security update).
# P0 PREEMPTS: it cancels the in-progress / queued CI runs of ALL strictly-
# LOWER-priority OTHER open PRs to grab their runners immediately. A preempted
# PR simply re-runs on its next push / autoupdate rebase. P0 is the ONLY
# priority that preempts a *normal* lower run — P1–P9 never bump those (a
# higher PR may still reclaim a broken/draft lower run — see below).
#
# • P1–P9 = YIELD WITHOUT BUMPING a *normal* lower run. They do NOT cancel a
# normal lower-priority run already going — a higher-priority PR does not evict
# it, it just takes the next free slot (it MAY still reclaim a broken/draft
# lower run — see below). Mechanism: a bounded hold-back. This job defers (up to
# HOLD_BACK_BUDGET_SECONDS, kept well under timeout-minutes) while any strictly-
# higher-priority OTHER open PR still has an active/queued CI run, and — within
# its OWN priority level — while any peer is ordered ahead of it (an in-flight
# run keeps its place; then oldest createdAt first). It proceeds the moment it
# is at the front, or when the budget elapses (a PR never blocks itself).
#
# • `broken` / `draft` = BOTTOM (effective P10). Always yields, never preempts —
# and because its run is wasted (a broken PR can't merge; a draft isn't merge-
# ready), ANY higher-priority PR (not just P0) MAY cancel that run to reclaim
# its runner (still the strictly-lower rule: broken/draft is the bottom, so any
# ready PR outranks it). A maintainer marks a stuck/failing PR `broken` to drop
# it below everything so others aren't blocked behind it AND may reclaim its
# runner; a draft behaves the same until it is marked ready for review.
#
# Hard safety invariants, enforced in the script:
# • never cancels a run on main / a push event (the gh query filters
# --event pull_request and drops headBranch == main);
# • never cancels THIS PR's own run (skips self by PR number + run id);
# • never cancels an equal-or-higher-priority PR (only strictly-lower, prio > self);
# • P1–P9 never bump a *normal* lower run (they only reclaim broken/draft) —
# otherwise they just wait (bounded), then proceed.
#
# Honest limitation: GitHub Actions has no native priority queue and assigns
# runners roughly FIFO, so the hold-back is a BEST-EFFORT head-start, not a hard
# guarantee — under sustained contention the bounded wait can expire before a
# higher-priority PR drains. The waiting job also occupies a (cheap, short-lived)
# runner meanwhile, which is exactly why the wait is kept bounded.
#
# It is deliberately NOT a merge-gate check: it is absent from `ci-passed`'s
# needs, every API call is guarded, the script always exits 0, and the step is
# `continue-on-error` — so a hiccup (API error, missing permission, fork PR) can
# never fail or block CI. The heavy jobs only *order* after it via `needs`; if it
# were ever skipped/failed they'd be skipped, which `ci-passed` treats as a gate
# failure (fail-safe: blocks merge, never spuriously passes).
traffic-control:
name: Traffic control (runner priority)
runs-on: ubuntu-latest
timeout-minutes: 6 # hard backstop; the P1–P9 hold-back budget stays well under this
permissions:
contents: read # check out .github/scripts/traffic_control.py
actions: write # cancel lower-priority runs (P0 emergencies + broken/draft reclaim)
pull-requests: read # read PR P0–P9 labels + draft state
env:
GH_TOKEN: ${{ github.token }}
GH_REPO: ${{ github.repository }}
# On a `pull_request` run this is the PR number and the script does its full in-run
# runner-priority orchestration. On a scheduler `workflow_dispatch` run (issue #349)
# the event is not `pull_request`, so the script no-ops here (`--mode orchestrate`
# only acts on pull_request events) — priority was ALREADY applied at trigger time by
# ci-trigger.yml, so re-doing the in-run hold-back would just waste runner time. The
# `|| inputs.pr` keeps the number in the log for a dispatched run.
SELF_PR: ${{ github.event.pull_request.number || inputs.pr }}
# P1–P9 bounded hold-back knobs, read by traffic_control.py. BUDGET must stay
# comfortably below timeout-minutes so the poll loop always exits 0 before the
# hard job timeout fires — a timed-out job would skip the heavy jobs and fail
# `ci-passed`.
HOLD_BACK_BUDGET_SECONDS: "180"
HOLD_BACK_POLL_SECONDS: "15"
steps:
- name: Check out source
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
# gh is auto-configured from GH_TOKEN / GH_REPO; python3 is preinstalled on the
# runner. The script guards every API call and always exits 0 (belt-and-braces
# with continue-on-error), so it can never fail or block CI.
- name: Apply runner priority (P0/broken/draft preempt; P1–P9 hold back)
continue-on-error: true
run: python3 .github/scripts/traffic_control.py
+6
View File
@@ -30,6 +30,8 @@ secrets.properties
# Log Files
*.log
# ...but keep the device-testing parser fixtures (verbatim logcat slices used as test inputs)
!scripts/device-testing/tests/fixtures/*.log
# Android Studio / IntelliJ
.idea/
@@ -42,5 +44,9 @@ captures/
# Kotlin
.kotlin/
# Python (dev/CI helper scripts under .github/scripts, .claude/…)
__pycache__/
*.pyc
# Claude Code — personal settings (the shared settings.json IS committed)
.claude/settings.local.json
+146
View File
@@ -0,0 +1,146 @@
# SPDX-License-Identifier: GPL-3.0-or-later
#
# ============================================================================
# Mergify configuration — PHASE 1: serial merge queue (issue #409).
#
# Spec: docs/ci/mergify-integration-spec.md + docs/ci/mergify.yml.proposed
# (issues #407 / #408).
# Schema: https://docs.mergify.com/configuration/file-format/
# Verified against the LIVE Mergify docs on 2026-07-06 (queue rules,
# the queue action, priority rules, parallel checks, batches, setup and
# lifecycle pages) — the config format evolves, so this is not from memory.
# ============================================================================
#
# WHAT THIS DOES
# A SERIAL merge queue that ends the manual serial-bump grind and supersedes the
# hand-rolled "poor-man's merge queue" (autoupdate.yml + ci-trigger.yml +
# traffic-control.yml — all already `disabled_manually`). Mergify updates each
# queued PR onto the latest `main`, re-runs CI, and merges it with a MERGE COMMIT
# when the single required gate — the "CI passed" check — is green. One PR at a
# time, in P0–P9 priority order.
#
# HARD INVARIANTS (do NOT relax without the trilemma decision recorded in the spec):
# * require-up-to-date STAYS ON. This is Phase 1 = batch_size 1 + merge_method:
# merge — the ONLY trilemma combination that keeps GitHub's "Require branches to
# be up to date before merging" LITERALLY enabled AND preserves the merge-commit
# policy. Mergify honours it by updating each PR onto the latest `main` and
# re-running CI before merging ("Updates PRs against the latest main before
# merging" — docs.mergify.com/merge-queue/setup). NO batching: batching would
# require turning that checkbox OFF (docs.mergify.com/merge-queue/batches) and is
# the blocked Phase 2 / issue #410 — explicitly OUT OF SCOPE here.
# * The single required status check stays "CI passed" — the exact `name:` of the
# `ci-passed` job in .github/workflows/ci.yml. NOT "ci-passed". A wrong name means
# PRs queue but never merge.
# ---------------------------------------------------------------------------
# queue_rules — how a queued PR is validated and merged.
# ---------------------------------------------------------------------------
queue_rules:
- name: default
# Final merge gate. Merge ONLY when the single required context is green (the exact
# same check branch protection requires), the PR is not a draft, has no merge
# conflicts, and is not flagged `broken`. NOTE: branch protection requires 0
# approvals here (the active repository ruleset sets required_approving_review_count
# = 0), so there is deliberately NO `#approved-reviews-by` condition — adding one
# would wedge the solo-maintainer flow, where nobody can approve their own PR.
merge_conditions:
- check-success = CI passed
- -draft
- -conflict
- label != broken
# SERIAL: exactly one PR per merge. No batching (Phase 2 / #410). One merge commit
# per PR, which is what lets require-up-to-date stay literally ON.
batch_size: 1
# Merge commit — never squash / rebase / fast-forward (repo policy: merges use
# merge commits, never squash).
merge_method: merge
# ---------------------------------------------------------------------------
# merge_queue — queue-wide options.
# ---------------------------------------------------------------------------
merge_queue:
# Validate ONE PR at a time — true serial, no speculative parallel checks. This is the
# strictest, unambiguously require-up-to-date-compatible setting: Mergify updates the
# REAL PR branch onto the latest `main`, runs CI on that branch, and merges on the real
# green "CI passed" — with no speculative temp-branch/real-branch check mismatch to
# reason about. It also caps the expensive, wedge-prone ~15-min E2E matrix at a single
# concurrent run. Raising this (speculative parallelism) is a throughput optimisation to
# weigh alongside the Phase 2 / #410 batching decision — not part of serial Phase 1.
max_parallel_checks: 1
# ---------------------------------------------------------------------------
# priority_rules — map the repo's P0–P9 labels onto queue priority.
# Higher number merges first (Mergify keywords: low=1000 / medium=2000 / high=3000;
# numeric range 1–10000). P0 is emergency-only and outranks everything. PRs with no P-label
# fall to Mergify's default `medium` (2000).
# ---------------------------------------------------------------------------
priority_rules:
- name: p0-emergency
conditions:
- label = P0
priority: 10000
allow_checks_interruption: true
- name: p1
conditions:
- label = P1
priority: 9000
allow_checks_interruption: true
- name: p2
conditions:
- label = P2
priority: 8000
allow_checks_interruption: true
- name: p3
conditions:
- label = P3
priority: 7000
allow_checks_interruption: true
- name: p4
conditions:
- label = P4
priority: 6000
allow_checks_interruption: true
- name: p5
conditions:
- label = P5
priority: 5000
allow_checks_interruption: true
- name: p6
conditions:
- label = P6
priority: 4000
allow_checks_interruption: true
- name: p7
conditions:
- label = P7
priority: 3000
allow_checks_interruption: true
- name: p8
conditions:
- label = P8
priority: 2000
- name: p9
conditions:
- label = P9
priority: 1000
# ---------------------------------------------------------------------------
# pull_request_rules — WHICH PRs enter the queue.
# The `queue` action is what actually ADDS a PR to the merge queue: per
# docs.mergify.com/merge-queue/lifecycle, queue_conditions alone do NOT auto-queue a PR —
# a queue action (or an `@mergifyio queue` command / auto_merge) is required, otherwise the
# "Mergify Merge Queue" check sits permanently pending. A PR is queued as soon as it is green
# on "CI passed", targets `main`, is not a draft, has no conflicts, and is not flagged
# `broken`. `broken` / `draft` PRs are never queued.
# ---------------------------------------------------------------------------
pull_request_rules:
- name: Queue green, non-draft, non-conflicting PRs targeting main
conditions:
- base = main
- -draft
- -conflict
- label != broken
- check-success = CI passed
actions:
queue:
name: default
+38 -14
View File
@@ -27,24 +27,38 @@ or via Gradle Managed Devices `./gradlew e2eGroupDebugAndroidTest` (whole matrix
`app/build.gradle.kts` must stay in lockstep with the E2E matrix in `.github/workflows/ci.yml`.
**Before treating a change as done**, run the fast CI gate: `assembleDebug` +
`testDebugUnitTest` + `compileDebugAndroidTestKotlin` + `lintDebug` + `ktlintCheck` +
`detekt` + the top-of-matrix emulator E2E `api35DebugAndroidTest` + `api36DebugAndroidTest` +
the API 37 preview E2E via `python3 .claude/skills/preflight/api37_e2e.py` (the `/preflight`
skill does all of this). `compileDebugAndroidTestKotlin` compiles the `androidTest` source set
`testDebugUnitTest` + `jacocoTestCoverageVerification` + `compileDebugAndroidTestKotlin` +
`lintDebug` + `ktlintCheck` + `detekt` + the local emulator E2E — the instrumented test class(es)
you changed via `python3 .claude/skills/preflight/local_instrumented.py <classes>` + the API 37
preview E2E via `python3 .claude/skills/preflight/api37_e2e.py` (the `/preflight` skill does all
of this). `jacocoTestCoverageVerification` runs right after `testDebugUnitTest` (it reads that
task's JVM exec data) and enforces the whole-app **no-regression line-coverage floor (currently
0.84)**, catching coverage regressions locally instead of only in CI (the exact class of failure
that reached CI on #367). `compileDebugAndroidTestKotlin` compiles the `androidTest` source set
that the static part of the gate skips, catching E2E/instrumented-test compile errors before
they surface only in CI. `ktlintCheck`/`detekt` cover the `test`/`androidTest` source sets that
`lintDebug` skips, so they catch style violations that would otherwise fail CI's Static analysis
gate. `api35DebugAndroidTest` and `api36DebugAndroidTest` run the instrumented/E2E suite on the
top two stable API levels in the E2E matrix, each via its own Gradle Managed Device (Gradle boots
and tears down each emulator automatically). API 37 (preview) has no Gradle Managed Device — its
gate. The local E2E does **not** use Gradle Managed Devices (`apiXXDebugAndroidTest`): GMD's
snapshot step fails locally under the AEHD 2.2 hypervisor. Instead `local_instrumented.py`
cold-boots one existing AVD by hand (no GMD, no snapshot) and runs `connectedDebugAndroidTest`
filtered to the class(es) you pass — run the ones you changed; the full ~114-test suite wedges
mid-run locally. API 37 (preview) has no Gradle Managed Device — its
only image is the nonstandard `android-37.0` / `google_apis_ps16k` pairing (see the comment above
`testOptions.managedDevices` in `app/build.gradle.kts`) — so `api37_e2e.py` hand-provisions it,
mirroring CI's `e2e-preview` job (same image + emulator flags, except it uses host-GPU
`-gpu auto-no-window` locally vs CI's headless `-gpu swiftshader_indirect`), boots it headless,
runs `connectedDebugAndroidTest`, and tears it down. Running all three levels locally is required; the
rest of the multi-API matrix (API 29–34) stays CI's job. Emulators need a free hardware
runs `connectedDebugAndroidTest`, and tears it down. Both local E2E steps are required; the
full multi-API matrix (API 29–37) stays CI's job. Emulators need a free hardware
hypervisor (VT-x/WHPX), so shut down VirtualBox/other VMs first or the AVD hangs at 0% CPU.
**Dev scripts: prefer Python (stdlib).** Auxiliary dev / CI-helper scripts — like the preflight
E2E runners (`.claude/skills/preflight/local_instrumented.py`, `api37_e2e.py`) and
`.claude/hooks/check-spdx.py` — are written in **Python 3, standard library only**, for
cross-platform portability. The primary dev box is Windows, where bash-only helpers need Git Bash
and hit gaps (`jq` missing, `taskkill` vs `kill`, path/quoting). **Do not add new bash-only
(`.sh`) or PowerShell-only dev scripts**; write new helpers in Python (or extend the existing
ones). Scope is auxiliary tooling only — product code stays Kotlin and Gradle stays Kotlin DSL.
## Build-config gotchas
- **Built-in Kotlin (AGP 9.x).** Kotlin compilation is handled by AGP's built-in Kotlin;
@@ -77,11 +91,21 @@ pulled in as a real dependency for unit tests because `android.jar`'s version is
A change is not done until it ships with passing **unit tests** and **E2E/instrumented tests**
that exercise the new or changed behaviour. Writing and committing that E2E/instrumented test
is a required part of every task — and the test must actually **run and pass**, not merely
compile: preflight runs the top-of-matrix emulator E2E locally — `api35DebugAndroidTest` +
`api36DebugAndroidTest` (Gradle Managed Devices) plus the API 37 preview via
`api37_e2e.py` (hand-provisioned, mirroring CI's `e2e-preview` job) — and all three must be
green before the change is done. CI then runs the full multi-API matrix plus the API 37 preview
job.
compile: preflight runs the changed instrumented test class(es) on a locally cold-booted
emulator via `local_instrumented.py` (no Gradle Managed Devices — they fail locally) plus the
API 37 preview via `api37_e2e.py` (hand-provisioned, mirroring CI's `e2e-preview` job) — and both
must be green before the change is done. CI then runs the full multi-API matrix plus the API 37
preview job.
No **app source-code** change is complete without **appropriate logging** added at its key
points — lifecycle transitions, error/fallback paths, significant state changes — so behaviour is
diagnosable from a user's debug report. Log through the `AppLog` facade
(`org.libremail.reporting.AppLog`), which mirrors to Logcat **and** the in-memory
`RingLogBuffer` that feeds a `DebugReport` — never raw `android.util.Log` (a detekt guard forbids
it). Logging must be **PII-free**: never log emails, server hosts, message content, or
credentials — use `accountLogRef(account.id)` for account references; throwables passed to
`AppLog` are auto-scrubbed. This applies to app source changes; pure test/config/doc changes
don't need new logging.
## Repo etiquette
+288 -131
View File
@@ -1,4 +1,6 @@
// SPDX-License-Identifier: GPL-3.0-or-later
import org.gradle.testing.jacoco.plugins.JacocoTaskExtension
import org.gradle.testing.jacoco.tasks.JacocoCoverageVerification
import org.gradle.testing.jacoco.tasks.JacocoReport
import java.util.Properties
@@ -64,6 +66,12 @@ android {
buildConfigField("String", "OUTLOOK_OAUTH_CLIENT_ID", "\"$outlookOAuthClientId\"")
buildConfigField("String", "OUTLOOK_OAUTH_REDIRECT_URI", "\"$outlookRedirectScheme://oauth2redirect\"")
buildConfigField("String", "DEBUG_REPORT_ENDPOINT", "\"$debugReportEndpoint\"")
// IMAP connection reuse (issue #357 Part 2, wiring the #125 spike): keep one authenticated
// IMAP connection warm per account instead of paying a cold CONNECT+TLS+LOGIN on every
// operation — the fix for Gmail throttling LibreMail's connect-per-operation traffic. ON by
// default; this is the safety switch: flip to "false" here (a build-config change, no Kotlin
// edit) to fall back to connect-per-operation if a server misbehaves with a kept-alive socket.
buildConfigField("Boolean", "IMAP_CONNECTION_REUSE", "true")
// AppAuth's bundled manifest requires this placeholder; it registers the redirect scheme on
// RedirectUriReceiverActivity so the Outlook sign-in redirect returns to the app.
manifestPlaceholders["appAuthRedirectScheme"] = outlookRedirectScheme
@@ -140,6 +148,12 @@ android {
}
testOptions {
// Robolectric-backed Compose UI unit tests (issue #373) need the merged Android resources
// (drawables, strings, the compiled resource table) on the JVM unit-test classpath so
// `stringResource(...)` and Material3 theming resolve without an emulator. Off by default in
// AGP; JVM tests that don't touch resources are unaffected.
unitTests.isIncludeAndroidResources = true
// Gradle Managed Devices define the per-API E2E matrix as config-as-code: one virtual
// device per supported Android version (a rolling ~7-year window, API 29 → latest stable).
// Run the whole matrix with `./gradlew e2eGroupDebugAndroidTest`, or one level with e.g.
@@ -199,10 +213,181 @@ jacoco {
toolVersion = libs.versions.jacoco.get()
}
// Unit-test coverage report (issue #192). Reads the exec data the base `jacoco` plugin records for
// the JVM `testDebugUnitTest` task, mapped against the debug variant's compiled Kotlin classes and
// the hand-written main sources. Produces machine-readable XML + human-readable HTML under
// build/reports/jacoco/jacocoTestReport/. Instrumented/E2E coverage is out of scope (issue #192).
// The Robolectric-backed JVM Compose UI tests (#373) load the classes-under-test through
// Robolectric's sandbox classloader, which presents them to the JaCoCo agent WITHOUT a code-source
// location. JaCoCo skips no-location classes by default, so on-the-fly coverage for every composable
// exercised only by a Robolectric test would silently record as zero — the file would be removed
// from `jacocoNonJvmTestableSurface` yet contribute nothing but missed lines, dragging the bundle
// ratio DOWN instead of up. `isIncludeNoLocationClasses = true` makes the agent keep that coverage;
// `jdk.internal.*` is excluded because instrumenting those JDK classes breaks under JDK 17+.
tasks.withType<Test>().configureEach {
configure<JacocoTaskExtension> {
isIncludeNoLocationClasses = true
excludes = listOf("jdk.internal.*")
}
}
// Unit-test coverage (issue #192). Two tasks share ONE scoping so they can never measure different
// surfaces: `jacocoTestReport` (XML+HTML under build/reports/jacoco/jacocoTestReport/) and
// `jacocoTestCoverageVerification` (the no-regression gate, further down). Both read the exec data
// the base `jacoco` plugin records for the JVM `testDebugUnitTest` task, mapped against the debug
// variant's compiled Kotlin classes and the hand-written main sources. Instrumented/E2E coverage is
// out of scope (issue #192).
// Strip generated code from the denominator so the % reflects hand-written Kotlin. Verified
// against an actual compileDebugKotlin output tree: Room's KSP-generated `_Impl` DAOs/database
// and the Compose compiler's per-file ComposableSingletons holders are the only generated code
// that actually lands in classDirectories below (Room's KSP output is added as an extra Kotlin
// source root on the *same* compile task, so it comes out the same door as hand-written code).
// Hilt/Dagger's generated Java (Hilt_*, Dagger*_HiltComponents*, *_GeneratedInjector, *_Factory,
// *_MembersInjector, hilt_aggregated_deps) and AGP's BuildConfig/R/Manifest are compiled by a
// separate javac task (hiltJavaCompileDebug / compileDebugJavaWithJavac) into a directory this
// report never reads, so those patterns are conventional belt-and-suspenders in case that ever
// changes. DataBinding isn't enabled in this module (no buildFeatures.dataBinding/viewBinding),
// so there's nothing generated for it to exclude; if it's turned on later, add "**/BR.class",
// "**/DataBinderMapperImpl*.class" and "**/*Binding.class".
//
// Deliberately NOT excluded: Kotlin's own `$$inlined$` synthetic classes (e.g. for
// `Flow.map { ... }` in the repositories) — those hold real hand-written transform logic, not
// generated boilerplate, so stripping them would silently shrink the measured surface.
val jacocoGeneratedExcludes = listOf(
"**/R.class",
"**/R\$*.class",
"**/BuildConfig.*",
"**/Manifest*.*",
"**/Hilt_*.class",
"**/Dagger*.class",
"**/*_Hilt*",
"**/*_GeneratedInjector.class",
"**/hilt_aggregated_deps/**",
"**/dagger/**",
"**/*_Factory*",
"**/*_MembersInjector*",
"**/*_Provide*",
"**/*_Impl*",
"**/ComposableSingletons*",
)
// Scope the denominator to the JVM-testable surface (issues #290/#292, following the Phase-2 coverage
// audit): unlike `jacocoGeneratedExcludes` above, none of this is generated code — it is hand-written
// but structurally unreachable from a JVM unit test, so counting it against the metric just measures
// how much Compose/framework glue exists rather than how well the logic is tested. Four buckets:
// 1. Compose screen/component render code. Historically only exercisable via an emulator, so it was
// excluded here. Issue #373 changes that: Robolectric runs the Android framework on the JVM, so a
// `createComposeRule()` test in the `test` source set now gives these files real JVM coverage
// without an emulator. This bucket therefore SHRINKS one screen at a time — each glob is deleted
// in the same PR that adds that screen's Robolectric JVM Compose test. AddAnotherAccountScreen was
// the first (see AddAnotherAccountScreenJvmTest) and has been removed below; the rest are tracked
// as per-area conversion tickets under #373. The coverage-floor re-ratchet is deferred until the
// whole conversion is done and stable (#373) — do NOT raise it in a conversion PR.
// 2. Android framework entry points the OS instantiates directly (Activity/Service/Application/
// BackupAgent) rather than the app's own code constructing them.
// 3. Hilt DI modules — `@Provides`/`@Binds` one-liners with no branching logic.
// 4. The `src/debug` cold-open probe (issue #221), a `ContentProvider` that only runs in a forked
// instrumented process (see its kdoc) and is never packaged in a release build anyway.
//
// KEPT IN SCOPE — this corrects #292, which excluded `**/*Worker*`: the six WorkManager workers
// (SyncWorker, BackfillWorker, PruneWorker, SendWorker, ReportPurgeWorker, ReportUploadWorker) are
// all directly unit-tested today (construct-the-worker-and-call-doWork(), e.g. SyncWorkerTest,
// SendWorkerTest), so they carry real tested logic and belong in BOTH the numerator and denominator.
// Only their Hilt wiring (WorkManagerModule) is excluded, and that falls under `**/di/**` below — so
// there is intentionally no `**/*Worker*` glob in the list.
//
// Also deliberately NOT excluded, even though each sits in a package/pattern above and renders UI:
// files that carry plain, unit-tested logic alongside their `@Composable` functions. JaCoCo has no
// finer granularity than a class file, and Kotlin compiles every top-level function in a .kt file —
// `@Composable` or not — into the SAME facade class (`<File>Kt.class`); excluding that class would
// silently zero out the tested function's coverage too, not just the render code's. Confirmed
// against these files' own dedicated tests before leaving them out of the list below:
// - ui/compose/RichTextEditor.kt (RichTextEditorTest) — the AnnotatedString<->RichTextContent
// editor-op functions (applyStyle/applyBlock/applyLink/toRichContent/toAnnotatedString/...).
// - ui/settings/AccountReorderList.kt (AccountReorderListTest) — commitDrag's reorder maths.
// - ui/reader/HtmlBody.kt (HtmlBodyTest, InlineImageResolverTest) — cidKey/resolveInlineImage/
// wrapHtml/toCssHex.
// - ui/reporting/ReportReviewScreen.kt (ReportReviewClipboardTest) — copyReportPayloadToClipboard.
// (ui/compose/format/FontRegistry.kt and ui/mailbox/FolderLabels.kt are plain logic files with no
// `@Composable` at all — never at risk — but sit right next to excluded files below.) For the same
// reason this list names each Screen/component file individually rather than a package-wide
// "**/ui/**": a blanket pattern can't carve the four files above back out, and would also reach
// every `*ViewModel*`.
val jacocoNonJvmTestableSurface = listOf(
// --- Compose UI render code: one glob per screen/component file (see the exceptions above) ---
// LibreMailApp KEPT excluded (#384, the acceptable exception): the composable is a real NavHost whose
// non-onboarding start destinations call hiltViewModel(), and standing the graph up needs owners a
// plain JVM compose rule can't surface — so graph-level nav stays on the instrumented OnboardingFlowTest.
// Its JVM-tractable parts (LibreMailBottomBar, StartupCrashPrompt, the cold-start hold guards) ARE
// exercised by LibreMailAppJvmTest, but the file's compiled facade (LibreMailAppKt) stays excluded.
"**/LibreMailApp*",
// AccountPickerScreen, AppPasswordSetupScreen & ManualSetupScreen converted to Robolectric JVM
// Compose tests (#378) — now JVM-covered.
// ComposeScreen (the email editor) converted to a Robolectric JVM Compose test (#382) — now
// JVM-covered.
// ColorSwatch(Row), FontPicker, FontSizePicker & ParagraphAlignmentControl converted to
// Robolectric JVM Compose tests (#376) — now JVM-covered.
// DraftsScreen, OutboxScreen & ProblemReportsScreen converted to Robolectric JVM Compose tests
// (#379) — now JVM-covered.
// LockScreen converted to a Robolectric JVM Compose test (#377) — now JVM-covered.
// AppLockGateHost converted to a Robolectric JVM Compose test (#384) — now JVM-covered.
// FolderDrawer & MailboxScreen (the Paging 3 mailbox list + folder drawer) converted to
// Robolectric JVM Compose tests (#383) — now JVM-covered.
// AddAnotherAccountScreen (#373) plus the onboarding welcome/license and contacts/battery steps
// (#377) converted to Robolectric JVM Compose tests — now JVM-covered.
// ReaderScreen converted to a Robolectric JVM Compose test (#381) — now JVM-covered. Its HTML body
// renders through HtmlBody, a hardened WebView that Robolectric can only present as a non-rendering
// shadow, so ReaderScreenJvmTest asserts the chrome (top bar, star/delete/reply actions, attachment
// accordion) and the loading/plain-text/empty/error/remote-images-banner branches — never the
// WebView's rendered HTML. HtmlBody.kt stays in scope covered by HtmlBodyTest/InlineImageResolverTest.
// SettingsScreen (+ ContactAutocompleteRow), AccountSettingsScreen, SettingsComponents (SwitchRow/
// ClickRow/RadioRow/RetentionSection), SignaturesScreen & SignatureEditScreen converted to
// Robolectric JVM Compose tests (#380) — now JVM-covered.
// CacheEncryptionGate.kt (issue #359/#367 fail-closed encryption gate) is pure render: the gate
// composable, its blank cover, the error screen, and the ephemeral report-review screen — no plain
// top-level logic. Spelled out to "...GateKt*" (the file's compiled facade class), NOT the bare
// "**/CacheEncryptionGate*" this list otherwise uses, because unlike every Screen/ViewModel pair
// above, CacheEncryptionGateViewModel's name literally starts with "CacheEncryptionGate" — a bare
// wildcard would also swallow the (94%-covered, dedicated-tested) ViewModel and its sealed
// CacheEncryptionGateState. CacheEncryptionGateViewModel and CacheEncryptionUnavailableException
// stay in scope (both have JVM tests: CacheEncryptionGateViewModelTest, DatabaseProvisionerTest).
"**/CacheEncryptionGateKt*",
// --- Android framework entry points (OS-instantiated). NB: Workers are intentionally NOT here
// --- (they are unit-tested — see the KEPT IN SCOPE note above).
"**/*Activity*",
"**/*Service*",
"**/LibreMailApplication*",
"**/*BackupAgent*",
// --- Hilt DI wiring (includes WorkManagerModule) ---
"**/di/**",
// --- src/debug cold-open probe (issue #221) ---
"**/data/local/coldopen/**",
// --- src/debug fetch-gate receiver (issue #393): a BroadcastReceiver that only runs on-device
// --- (adb-driven), covered by an instrumented test, never packaged into a release build. Its
// --- pure collaborators DebugFetchGate/FetchScope stay IN scope (unit-tested by DebugFetchGateTest).
"**/debug/FetchGateReceiver*",
)
// Classes = the debug variant's compiled Kotlin (AGP 9 built-in Kotlin output), with the generated
// code and the non-JVM-testable surface above stripped out. All hand-written code here is Kotlin, so
// the javac output (purely Hilt/Dagger/BuildConfig generated) is omitted. Hoisted to a shared val so
// the report and the verification gate always run against the identical denominator.
val jacocoDebugKotlinClasses = layout.buildDirectory.dir(
"intermediates/built_in_kotlinc/debug/compileDebugKotlin/classes",
)
val jacocoClassDirectories = fileTree(jacocoDebugKotlinClasses) {
exclude(jacocoGeneratedExcludes + jacocoNonJvmTestableSurface)
}
// Sources = hand-written main Kotlin.
val jacocoSourceDirectories = files("src/main/kotlin")
// Exec data written by the instrumented testDebugUnitTest task. Accept the base `jacoco` plugin's
// default location and AGP's enableUnitTestCoverage location so the wiring is robust either way.
val jacocoExecutionData = fileTree(layout.buildDirectory) {
include(
"jacoco/testDebugUnitTest.exec",
"outputs/unit_test_code_coverage/debugUnitTest/testDebugUnitTest.exec",
)
}
tasks.register<JacocoReport>("jacocoTestReport") {
// Ensure the unit tests (and thus their coverage exec data) have run first.
dependsOn("testDebugUnitTest")
@@ -214,137 +399,91 @@ tasks.register<JacocoReport>("jacocoTestReport") {
html.required.set(true)
}
// Strip generated code from the denominator so the % reflects hand-written Kotlin. Verified
// against an actual compileDebugKotlin output tree: Room's KSP-generated `_Impl` DAOs/database
// and the Compose compiler's per-file ComposableSingletons holders are the only generated code
// that actually lands in classDirectories below (Room's KSP output is added as an extra Kotlin
// source root on the *same* compile task, so it comes out the same door as hand-written code).
// Hilt/Dagger's generated Java (Hilt_*, Dagger*_HiltComponents*, *_GeneratedInjector, *_Factory,
// *_MembersInjector, hilt_aggregated_deps) and AGP's BuildConfig/R/Manifest are compiled by a
// separate javac task (hiltJavaCompileDebug / compileDebugJavaWithJavac) into a directory this
// report never reads, so those patterns are conventional belt-and-suspenders in case that ever
// changes. DataBinding isn't enabled in this module (no buildFeatures.dataBinding/viewBinding),
// so there's nothing generated for it to exclude; if it's turned on later, add "**/BR.class",
// "**/DataBinderMapperImpl*.class" and "**/*Binding.class".
//
// Deliberately NOT excluded: Kotlin's own `$$inlined$` synthetic classes (e.g. for
// `Flow.map { ... }` in the repositories) — those hold real hand-written transform logic, not
// generated boilerplate, so stripping them would silently shrink the measured surface.
val generated = listOf(
"**/R.class",
"**/R\$*.class",
"**/BuildConfig.*",
"**/Manifest*.*",
"**/Hilt_*.class",
"**/Dagger*.class",
"**/*_Hilt*",
"**/*_GeneratedInjector.class",
"**/hilt_aggregated_deps/**",
"**/dagger/**",
"**/*_Factory*",
"**/*_MembersInjector*",
"**/*_Provide*",
"**/*_Impl*",
"**/ComposableSingletons*",
)
classDirectories.setFrom(jacocoClassDirectories)
sourceDirectories.setFrom(jacocoSourceDirectories)
executionData.setFrom(jacocoExecutionData)
}
// Scope the denominator to the JVM-testable surface (issue #290, following the Phase-2 coverage
// audit): unlike `generated` above, none of this is generated code — it is hand-written but
// structurally unreachable from a JVM unit test, so counting it against the metric just measures
// how much Compose/framework glue exists rather than how well the logic is tested. Four buckets:
// 1. Compose screen/component render code — only exercisable via a Compose UI test or an emulator.
// 2. Android framework entry points the OS instantiates directly (Activity/Service/Worker/
// Application/BackupAgent) rather than the app's own code constructing them.
// 3. Hilt DI modules — `@Provides`/`@Binds` one-liners with no branching logic.
// 4. The `src/debug` cold-open probe (issue #221), a `ContentProvider` that only runs in a forked
// instrumented process (see its kdoc) and is never packaged in a release build anyway.
//
// Deliberately NOT excluded, even though each sits in a package/pattern above and renders UI: files
// that carry plain, unit-tested logic alongside their `@Composable` functions. JaCoCo has no finer
// granularity than a class file, and Kotlin compiles every top-level function in a .kt file —
// `@Composable` or not — into the SAME facade class (`<File>Kt.class`); excluding that class would
// silently zero out the tested function's coverage too, not just the render code's. Confirmed
// against these files' own dedicated tests before leaving them out of the list below:
// - ui/compose/RichTextEditor.kt (RichTextEditorTest) — the AnnotatedString<->RichTextContent
// editor-op functions (applyStyle/applyBlock/applyLink/toRichContent/toAnnotatedString/...).
// - ui/settings/AccountReorderList.kt (AccountReorderListTest) — commitDrag's reorder maths.
// - ui/reader/HtmlBody.kt (HtmlBodyTest, InlineImageResolverTest) — cidKey/resolveInlineImage/
// wrapHtml/toCssHex.
// - ui/reporting/ReportReviewScreen.kt (ReportReviewClipboardTest) — copyReportPayloadToClipboard.
// (ui/compose/format/FontRegistry.kt and ui/mailbox/FolderLabels.kt are plain logic files with no
// `@Composable` at all — never at risk — but sit right next to excluded files below.) For the same
// reason this list names each Screen/component file individually rather than a package-wide
// "**/ui/**": a blanket pattern can't carve the four files above back out, and would also reach
// every `*ViewModel*`.
//
// Tradeoff called out for review rather than silently applied: `**/*Worker*` excludes SyncWorker,
// BackfillWorker, PruneWorker, SendWorker, ReportPurgeWorker and ReportUploadWorker as framework
// entry points, per issue #290 — but all six are directly unit-tested today (construct-the-worker-
// and-call-doWork(), e.g. SyncWorkerTest, SendWorkerTest), so this also removes that already-tested
// coverage from both the numerator and the denominator, not just untested render/glue code.
val nonJvmTestableSurface = listOf(
// --- Compose UI render code: one glob per screen/component file (see the exceptions above) ---
"**/LibreMailApp*",
"**/AccountPickerScreen*",
"**/AppPasswordSetupScreen*",
"**/ManualSetupScreen*",
"**/ComposeScreen*",
"**/ColorSwatch*",
"**/FontPicker*",
"**/FontSizePicker*",
"**/ParagraphAlignmentControl*",
"**/DraftsScreen*",
"**/LockScreen*",
"**/AppLockGateHost*",
"**/FolderDrawer*",
"**/MailboxScreen*",
"**/AddAnotherAccountScreen*",
"**/BatteryOptimizationScreen*",
"**/ContactsAccessScreen*",
"**/LicenseScreen*",
"**/OnboardingWelcomeScreen*",
"**/OutboxScreen*",
"**/ReaderScreen*",
"**/ProblemReportsScreen*",
"**/AccountSettingsScreen*",
"**/SettingsScreen*",
"**/SettingsComponents*",
"**/SignatureEditScreen*",
"**/SignaturesScreen*",
// --- Android framework entry points ---
"**/*Activity*",
"**/*Service*",
"**/*Worker*",
"**/LibreMailApplication*",
"**/*BackupAgent*",
// --- Hilt DI wiring ---
"**/di/**",
// --- src/debug cold-open probe (issue #221) ---
"**/data/local/coldopen/**",
)
// No-regression coverage gate (closes #251; scoping from #290/#292). Fails `check` / CI when the
// overall LINE coverage of the scoped surface above drops below `jacocoLineCoverageFloor`. This is a
// FLOOR, not an absolute 95% target — the maintainer chose a ratchet over a fixed goal. The floor is
// normally set a hair (~0.5–1%) below the measured baseline so ordinary run-to-run noise doesn't
// red-flag it, while a real regression still fails the build. Manual ratchet FOR NOW: when coverage
// rises materially, bump this number up in the SAME PR so the floor tracks reality (there is no
// auto-ratchet yet).
//
// Re-ratcheted for #386 (final step of the Robolectric Compose epic #373, once infra/PoC #375 and
// conversion batches #376-384 had all landed and proven stable): new baseline 87.89% line
// (7994/9095), floor 0.84 — a wider ~3.9% headroom than the usual ~0.5-1%, chosen deliberately
// conservative for this first post-epic measurement; the maintainer can tighten it further in a
// follow-up PR. Prior baseline: 80.21% line (4838/6032), floor 0.79 (~1.2% headroom).
val jacocoLineCoverageFloor = "0.84"
tasks.register<JacocoCoverageVerification>("jacocoTestCoverageVerification") {
// Same inputs as jacocoTestReport (shared vals above) so the gate enforces exactly what the
// report shows. Depend on the unit tests so the exec data exists before verifying.
dependsOn("testDebugUnitTest")
group = "verification"
description = "Fails the build if scoped JVM unit-test LINE coverage regresses below the floor."
// Classes = the debug variant's compiled Kotlin (AGP 9 built-in Kotlin output). All hand-written
// code here is Kotlin, so the javac output (purely Hilt/Dagger/BuildConfig generated) is omitted.
val debugKotlinClasses = layout.buildDirectory.dir(
"intermediates/built_in_kotlinc/debug/compileDebugKotlin/classes",
)
classDirectories.setFrom(
fileTree(debugKotlinClasses) { exclude(generated + nonJvmTestableSurface) },
)
classDirectories.setFrom(jacocoClassDirectories)
sourceDirectories.setFrom(jacocoSourceDirectories)
executionData.setFrom(jacocoExecutionData)
// Sources = hand-written main Kotlin.
sourceDirectories.setFrom(files("src/main/kotlin"))
violationRules {
rule {
element = "BUNDLE"
limit {
counter = "LINE"
value = "COVEREDRATIO"
minimum = jacocoLineCoverageFloor.toBigDecimal()
}
}
}
}
// Exec data written by the instrumented testDebugUnitTest task. Accept the base `jacoco` plugin's
// default location and AGP's enableUnitTestCoverage location so the wiring is robust either way.
executionData.setFrom(
fileTree(layout.buildDirectory) {
include(
"jacoco/testDebugUnitTest.exec",
"outputs/unit_test_code_coverage/debugUnitTest/testDebugUnitTest.exec",
)
},
)
// Make the aggregate `check` lifecycle task enforce the no-regression floor locally too, so a
// coverage regression is caught by `./gradlew check` and not only in CI.
tasks.named("check") {
dependsOn("jacocoTestCoverageVerification")
}
// --- Robolectric android-all offline resolution (issue #373) ------------------------------------
// Robolectric runs the real Android framework on the JVM from a large `android-all-instrumented`
// jar. By default it resolves that jar LAZILY AT TEST TIME by downloading it from Maven Central
// (org.robolectric.internal.dependency.MavenDependencyResolver -> MavenArtifactFetcher). That
// runtime download is unreliable on CI runners and failed the JVM Compose PoC in CI with
// `java.lang.AssertionError at MavenArtifactFetcher ... Caused by: java.io.IOException` ("Failed to
// fetch maven artifact"). Fix: resolve the jar through Gradle instead — reliable, cached, and
// persisted by the CI Gradle cache, using the same repositories as every other dependency — then
// hand it to Robolectric in OFFLINE mode so it never touches the network at test time.
//
// A DEDICATED resolvable configuration (deliberately NOT testImplementation/testRuntimeOnly) keeps
// the ~200 MB instrumented framework jar OFF the JVM unit-test classpath: it must be loaded only by
// Robolectric's sandbox classloader, never flattened onto the app's test classpath where it would
// collide with the stub `android.jar`. `syncRobolectricAndroidAll` stages the resolved jar under
// its Maven filename (android-all-instrumented-<version>.jar) — exactly what Robolectric's
// LocalDependencyResolver looks up as <artifactId>-<version>.jar — and the two system properties
// below switch Robolectric onto that offline directory (see LegacyDependencyResolver). Every
// Robolectric test pins @Config(sdk = 36) (app/src/test/resources/robolectric.properties), so the
// single sdk=36 jar covers them all; a test on a different SDK must add that android-all version to
// this configuration too. The offline properties are inert for non-Robolectric JVM tests.
val robolectricAndroidAll: Configuration = configurations.create("robolectricAndroidAll") {
isCanBeConsumed = false
isCanBeResolved = true
}
val robolectricDepsDir = layout.buildDirectory.dir("robolectric-android-all")
val syncRobolectricAndroidAll = tasks.register<Sync>("syncRobolectricAndroidAll") {
description = "Stages Robolectric's android-all-instrumented jar for offline resolution (issue #373)."
from(robolectricAndroidAll)
into(robolectricDepsDir)
}
tasks.withType<Test>().configureEach {
dependsOn(syncRobolectricAndroidAll)
systemProperty("robolectric.offline", "true")
systemProperty("robolectric.dependency.dir", robolectricDepsDir.get().asFile.absolutePath)
}
dependencies {
@@ -407,6 +546,24 @@ dependencies {
// The real org.json for unit tests (android.jar ships a stubbed, no-op version).
testImplementation("org.json:json:20231013")
// Robolectric-backed JVM Compose UI tests (issue #373): Robolectric runs the Android framework
// on the JVM so `createComposeRule()` can drive composables without an emulator, bringing screen
// render code into the JaCoCo JVM-testable surface. The Compose test artifacts come from the same
// BOM as the app (aligned versions) and reuse the ui-test-junit4 / ui-test-manifest aliases the
// androidTest source set already declares — here in `test` (JVM), not `androidTest`. Robolectric
// sources Android's real org.json from its sandbox, so it does not clash with the stub-replacing
// org.json above (that is for the plain, non-Robolectric JVM tests).
testImplementation(libs.robolectric)
// The android-all-instrumented framework jar Robolectric loads into its sandbox — resolved via
// Gradle and staged for offline use by syncRobolectricAndroidAll above so no flaky test-time
// download happens in CI (issue #373). On its own dedicated configuration, NOT the test
// classpath — see that block for why. The artifact has no transitive dependencies (verified from
// its POM), so it resolves to exactly the one staged jar.
"robolectricAndroidAll"(libs.robolectric.android.all.instrumented)
testImplementation(platform(libs.androidx.compose.bom))
testImplementation(libs.androidx.compose.ui.test.junit4)
testImplementation(libs.androidx.compose.ui.test.manifest)
androidTestImplementation(libs.androidx.junit)
androidTestImplementation(libs.androidx.espresso.core)
androidTestImplementation(libs.androidx.espresso.intents)
@@ -0,0 +1,506 @@
{
"formatVersion": 1,
"database": {
"version": 20,
"identityHash": "8264768635364869a347064a0864df9c",
"entities": [
{
"tableName": "messages",
"createSql": "CREATE TABLE IF NOT EXISTS `${TABLE_NAME}` (`id` TEXT NOT NULL, `accountId` TEXT NOT NULL, `sender` TEXT NOT NULL, `senderEmail` TEXT NOT NULL, `subject` TEXT NOT NULL, `snippet` TEXT NOT NULL, `body` TEXT NOT NULL, `isHtml` INTEGER NOT NULL, `timestampMillis` INTEGER NOT NULL, `isRead` INTEGER NOT NULL, `isStarred` INTEGER NOT NULL, `folder` TEXT NOT NULL DEFAULT 'INBOX', `inInbox` INTEGER NOT NULL, `bodyFetched` INTEGER NOT NULL, `uid` INTEGER NOT NULL DEFAULT 0, `senderFold` TEXT NOT NULL DEFAULT '', `senderEmailFold` TEXT NOT NULL DEFAULT '', `subjectFold` TEXT NOT NULL DEFAULT '', `snippetFold` TEXT NOT NULL DEFAULT '', PRIMARY KEY(`id`))",
"fields": [
{
"fieldPath": "id",
"columnName": "id",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "accountId",
"columnName": "accountId",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "sender",
"columnName": "sender",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "senderEmail",
"columnName": "senderEmail",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "subject",
"columnName": "subject",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "snippet",
"columnName": "snippet",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "body",
"columnName": "body",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "isHtml",
"columnName": "isHtml",
"affinity": "INTEGER",
"notNull": true
},
{
"fieldPath": "timestampMillis",
"columnName": "timestampMillis",
"affinity": "INTEGER",
"notNull": true
},
{
"fieldPath": "isRead",
"columnName": "isRead",
"affinity": "INTEGER",
"notNull": true
},
{
"fieldPath": "isStarred",
"columnName": "isStarred",
"affinity": "INTEGER",
"notNull": true
},
{
"fieldPath": "folder",
"columnName": "folder",
"affinity": "TEXT",
"notNull": true,
"defaultValue": "'INBOX'"
},
{
"fieldPath": "inInbox",
"columnName": "inInbox",
"affinity": "INTEGER",
"notNull": true
},
{
"fieldPath": "bodyFetched",
"columnName": "bodyFetched",
"affinity": "INTEGER",
"notNull": true
},
{
"fieldPath": "uid",
"columnName": "uid",
"affinity": "INTEGER",
"notNull": true,
"defaultValue": "0"
},
{
"fieldPath": "senderFold",
"columnName": "senderFold",
"affinity": "TEXT",
"notNull": true,
"defaultValue": "''"
},
{
"fieldPath": "senderEmailFold",
"columnName": "senderEmailFold",
"affinity": "TEXT",
"notNull": true,
"defaultValue": "''"
},
{
"fieldPath": "subjectFold",
"columnName": "subjectFold",
"affinity": "TEXT",
"notNull": true,
"defaultValue": "''"
},
{
"fieldPath": "snippetFold",
"columnName": "snippetFold",
"affinity": "TEXT",
"notNull": true,
"defaultValue": "''"
}
],
"primaryKey": {
"autoGenerate": false,
"columnNames": [
"id"
]
},
"indices": [
{
"name": "index_messages_accountId",
"unique": false,
"columnNames": [
"accountId"
],
"orders": [],
"createSql": "CREATE INDEX IF NOT EXISTS `index_messages_accountId` ON `${TABLE_NAME}` (`accountId`)"
},
{
"name": "index_messages_timestampMillis",
"unique": false,
"columnNames": [
"timestampMillis"
],
"orders": [],
"createSql": "CREATE INDEX IF NOT EXISTS `index_messages_timestampMillis` ON `${TABLE_NAME}` (`timestampMillis`)"
},
{
"name": "index_messages_accountId_folder_uid",
"unique": false,
"columnNames": [
"accountId",
"folder",
"uid"
],
"orders": [],
"createSql": "CREATE INDEX IF NOT EXISTS `index_messages_accountId_folder_uid` ON `${TABLE_NAME}` (`accountId`, `folder`, `uid`)"
},
{
"name": "index_messages_folder_inInbox_timestampMillis",
"unique": false,
"columnNames": [
"folder",
"inInbox",
"timestampMillis"
],
"orders": [],
"createSql": "CREATE INDEX IF NOT EXISTS `index_messages_folder_inInbox_timestampMillis` ON `${TABLE_NAME}` (`folder`, `inInbox`, `timestampMillis`)"
}
]
},
{
"tableName": "attachments",
"createSql": "CREATE TABLE IF NOT EXISTS `${TABLE_NAME}` (`messageId` TEXT NOT NULL, `partIndex` INTEGER NOT NULL, `filename` TEXT NOT NULL, `mimeType` TEXT NOT NULL, `sizeBytes` INTEGER NOT NULL, `contentId` TEXT, PRIMARY KEY(`messageId`, `partIndex`), FOREIGN KEY(`messageId`) REFERENCES `messages`(`id`) ON UPDATE NO ACTION ON DELETE CASCADE )",
"fields": [
{
"fieldPath": "messageId",
"columnName": "messageId",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "partIndex",
"columnName": "partIndex",
"affinity": "INTEGER",
"notNull": true
},
{
"fieldPath": "filename",
"columnName": "filename",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "mimeType",
"columnName": "mimeType",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "sizeBytes",
"columnName": "sizeBytes",
"affinity": "INTEGER",
"notNull": true
},
{
"fieldPath": "contentId",
"columnName": "contentId",
"affinity": "TEXT"
}
],
"primaryKey": {
"autoGenerate": false,
"columnNames": [
"messageId",
"partIndex"
]
},
"indices": [
{
"name": "index_attachments_messageId",
"unique": false,
"columnNames": [
"messageId"
],
"orders": [],
"createSql": "CREATE INDEX IF NOT EXISTS `index_attachments_messageId` ON `${TABLE_NAME}` (`messageId`)"
}
],
"foreignKeys": [
{
"table": "messages",
"onDelete": "CASCADE",
"onUpdate": "NO ACTION",
"columns": [
"messageId"
],
"referencedColumns": [
"id"
]
}
]
},
{
"tableName": "outbox",
"createSql": "CREATE TABLE IF NOT EXISTS `${TABLE_NAME}` (`id` TEXT NOT NULL, `accountId` TEXT NOT NULL, `toAddresses` TEXT NOT NULL, `ccAddresses` TEXT NOT NULL, `bccAddresses` TEXT NOT NULL DEFAULT '', `subject` TEXT NOT NULL, `body` TEXT NOT NULL, `createdAt` INTEGER NOT NULL, `lastError` TEXT, `bodyHtml` TEXT, `attachments` TEXT NOT NULL DEFAULT '', PRIMARY KEY(`id`))",
"fields": [
{
"fieldPath": "id",
"columnName": "id",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "accountId",
"columnName": "accountId",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "toAddresses",
"columnName": "toAddresses",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "ccAddresses",
"columnName": "ccAddresses",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "bccAddresses",
"columnName": "bccAddresses",
"affinity": "TEXT",
"notNull": true,
"defaultValue": "''"
},
{
"fieldPath": "subject",
"columnName": "subject",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "body",
"columnName": "body",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "createdAt",
"columnName": "createdAt",
"affinity": "INTEGER",
"notNull": true
},
{
"fieldPath": "lastError",
"columnName": "lastError",
"affinity": "TEXT"
},
{
"fieldPath": "bodyHtml",
"columnName": "bodyHtml",
"affinity": "TEXT"
},
{
"fieldPath": "attachments",
"columnName": "attachments",
"affinity": "TEXT",
"notNull": true,
"defaultValue": "''"
}
],
"primaryKey": {
"autoGenerate": false,
"columnNames": [
"id"
]
}
},
{
"tableName": "drafts",
"createSql": "CREATE TABLE IF NOT EXISTS `${TABLE_NAME}` (`id` TEXT NOT NULL, `accountId` TEXT, `toAddresses` TEXT NOT NULL, `ccAddresses` TEXT NOT NULL, `bccAddresses` TEXT NOT NULL DEFAULT '', `subject` TEXT NOT NULL, `body` TEXT NOT NULL, `updatedAt` INTEGER NOT NULL, `attachments` TEXT NOT NULL, `bodyHtml` TEXT, PRIMARY KEY(`id`))",
"fields": [
{
"fieldPath": "id",
"columnName": "id",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "accountId",
"columnName": "accountId",
"affinity": "TEXT"
},
{
"fieldPath": "toAddresses",
"columnName": "toAddresses",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "ccAddresses",
"columnName": "ccAddresses",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "bccAddresses",
"columnName": "bccAddresses",
"affinity": "TEXT",
"notNull": true,
"defaultValue": "''"
},
{
"fieldPath": "subject",
"columnName": "subject",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "body",
"columnName": "body",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "updatedAt",
"columnName": "updatedAt",
"affinity": "INTEGER",
"notNull": true
},
{
"fieldPath": "attachments",
"columnName": "attachments",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "bodyHtml",
"columnName": "bodyHtml",
"affinity": "TEXT"
}
],
"primaryKey": {
"autoGenerate": false,
"columnNames": [
"id"
]
}
},
{
"tableName": "folders",
"createSql": "CREATE TABLE IF NOT EXISTS `${TABLE_NAME}` (`accountId` TEXT NOT NULL, `fullName` TEXT NOT NULL, `displayName` TEXT NOT NULL, `role` TEXT NOT NULL, `selectable` INTEGER NOT NULL, `sortOrder` INTEGER NOT NULL, `specialUse` INTEGER NOT NULL DEFAULT 0, `hierarchyDelimiter` TEXT, PRIMARY KEY(`accountId`, `fullName`))",
"fields": [
{
"fieldPath": "accountId",
"columnName": "accountId",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "fullName",
"columnName": "fullName",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "displayName",
"columnName": "displayName",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "role",
"columnName": "role",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "selectable",
"columnName": "selectable",
"affinity": "INTEGER",
"notNull": true
},
{
"fieldPath": "sortOrder",
"columnName": "sortOrder",
"affinity": "INTEGER",
"notNull": true
},
{
"fieldPath": "specialUse",
"columnName": "specialUse",
"affinity": "INTEGER",
"notNull": true,
"defaultValue": "0"
},
{
"fieldPath": "hierarchyDelimiter",
"columnName": "hierarchyDelimiter",
"affinity": "TEXT"
}
],
"primaryKey": {
"autoGenerate": false,
"columnNames": [
"accountId",
"fullName"
]
}
},
{
"tableName": "backfill_progress",
"createSql": "CREATE TABLE IF NOT EXISTS `${TABLE_NAME}` (`accountId` TEXT NOT NULL, `folder` TEXT NOT NULL, `nextBeforeUid` INTEGER NOT NULL, `complete` INTEGER NOT NULL, PRIMARY KEY(`accountId`, `folder`))",
"fields": [
{
"fieldPath": "accountId",
"columnName": "accountId",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "folder",
"columnName": "folder",
"affinity": "TEXT",
"notNull": true
},
{
"fieldPath": "nextBeforeUid",
"columnName": "nextBeforeUid",
"affinity": "INTEGER",
"notNull": true
},
{
"fieldPath": "complete",
"columnName": "complete",
"affinity": "INTEGER",
"notNull": true
}
],
"primaryKey": {
"autoGenerate": false,
"columnNames": [
"accountId",
"folder"
]
}
}
],
"setupQueries": [
"CREATE TABLE IF NOT EXISTS room_master_table (id INTEGER PRIMARY KEY,identity_hash TEXT)",
"INSERT OR REPLACE INTO room_master_table (id,identity_hash) VALUES(42, '8264768635364869a347064a0864df9c')"
]
}
}
@@ -13,6 +13,7 @@ import kotlinx.coroutines.runBlocking
import org.json.JSONObject
import org.junit.After
import org.junit.Assert.assertEquals
import org.junit.Assert.assertFalse
import org.junit.Assert.assertNotNull
import org.junit.Assert.assertNull
import org.junit.Assert.assertTrue
@@ -21,6 +22,8 @@ import org.junit.Rule
import org.junit.Test
import org.junit.runner.RunWith
import org.libremail.data.local.entity.CredentialEntity
import org.libremail.reporting.AppLog
import org.libremail.reporting.RingLogBuffer
/**
* The one-time move performed by [AccountDataMigrator] (issue #111): copying accounts / credentials /
@@ -142,6 +145,26 @@ class AccountDataMigratorTest {
}
}
@Test
fun copyEmitsANonPiiAppLogBreadcrumbNamingOnlyTheMovedTables() = runBlocking<Unit> {
seedVersion14Cache()
val buffer = RingLogBuffer()
AppLog.install(buffer)
AccountDataMigrator.copyAccountTables(cacheFile, cachePassphrase = "", accountsFile = accountsFile)
val entry = buffer.snapshot()
.single { it.message.startsWith("moved account tables into the account database") }
assertEquals("the migration breadcrumb is a debug line", 'D', entry.level)
listOf("accounts", "credentials", "account_settings", "signatures").forEach { table ->
assertTrue("breadcrumb must name the moved table $table", entry.message.contains(table))
}
// The breadcrumb carries only table names — never the seeded email, secret, or passphrase.
assertFalse(entry.message.contains("ada@example.org"))
assertFalse(entry.message.contains("sealed-secret"))
assertFalse(entry.message.contains(passphrase))
}
@Test
fun reRunningTheCopyIsIdempotentAndKeepsLaterEdits() = runBlocking<Unit> {
seedVersion14Cache()
@@ -17,6 +17,8 @@ import org.junit.Before
import org.junit.Test
import org.junit.runner.RunWith
import org.libremail.data.local.entity.MessageEntity
import org.libremail.reporting.AppLog
import org.libremail.reporting.RingLogBuffer
import java.io.File
/**
@@ -147,7 +149,43 @@ class DatabaseEncryptionTest {
} finally {
encrypted.close()
}
assertEquals("Room's schema version must survive the plaintext -> encrypted conversion", 19, version)
assertEquals("Room's schema version must survive the plaintext -> encrypted conversion", 20, version)
}
@Test
fun conversionEmitsNonPiiAppLogBreadcrumbs() = runBlocking<Unit> {
val buffer = RingLogBuffer()
AppLog.install(buffer)
// The seeded row carries an email address so the PII assertions below are meaningful.
openPlaintext().apply {
messageDao().insertNew(listOf(message("acct:1")))
close()
}
DatabaseEncryption.ensureEncrypted(dbFile, passphrase)
val afterEncrypt = buffer.snapshot()
val converting = afterEncrypt.single { it.message.startsWith("converting local cache database") }
assertEquals("the start breadcrumb is informational", 'I', converting.level)
assertEquals("converting local cache database (targetEncrypted=true)", converting.message)
val convertedAfterEncrypt = afterEncrypt.single { it.message == "local cache database converted" }
assertEquals('D', convertedAfterEncrypt.level)
// Converting back to plaintext logs the same pair with the flag flipped.
buffer.clear()
DatabaseEncryption.ensurePlaintext(dbFile, passphrase)
val afterDecrypt = buffer.snapshot()
assertTrue(
afterDecrypt.any { it.message == "converting local cache database (targetEncrypted=false)" },
)
assertTrue(afterDecrypt.any { it.message == "local cache database converted" })
// Neither conversion's breadcrumbs may leak the passphrase, the on-disk path, or account PII.
(afterEncrypt + afterDecrypt).forEach { entry ->
assertFalse("must not leak the passphrase", entry.message.contains(passphrase))
assertFalse("must not leak the db file path", entry.message.contains(dbFile.absolutePath))
assertFalse("must not leak the seeded email", entry.message.contains("ada@example.org"))
}
}
private fun openPlaintext(): LibreMailDatabase =
@@ -142,6 +142,34 @@ class DatabaseProvisionerInstrumentedTest {
}
}
/**
* Issue #359: with `encryptCache` on and NO cache yet (a fresh install enabling encryption), the
* provisioner must load SQLCipher's native library, report [CacheOpenMode.Encrypted], and a real keyed
* open must then create and read the encrypted cache — i.e. `libsqlcipher.so` actually loads and runs.
*
* On a 16 KB memory-page device/image (Android 15+, and the CI API-37 preview `google_apis_ps16k`
* E2E image) an `.so` not aligned for 16 KB pages fails exactly here with `UnsatisfiedLinkError` at
* `SQLiteConnection.nativeOpen`. Running this on that image makes the 16 KB native-lib load a tested
* invariant, so a dependency bump that regressed alignment is caught in CI rather than on-device.
*/
@Test
fun freshEncryptOnStartLoadsThe16KbNativeLibAndOpensKeyedWithoutCrashing() = runBlocking<Unit> {
every { settingsRepository.settings } returns flowOf(AppSettings(encryptCache = true, appLock = false))
assertFalse("precondition: no cache file exists yet", dbFile.exists())
val mode = provisioner().prepareCache()
assertEquals(CacheOpenMode.Encrypted(passphrase), mode)
// The keyed open must actually succeed on real SQLCipher — loading and using libsqlcipher.so on
// whatever ABI / page size this device or emulator image uses.
openEncrypted().apply {
messageDao().insertNew(listOf(message("acct:1")))
assertEquals(listOf("acct:1"), messageDao().observeSummaries().first().map { it.id })
close()
}
assertTrue("the fresh cache was created in SQLCipher (encrypted) form", DatabaseEncryption.isEncrypted(dbFile))
}
@Test
fun encryptionTurnedOffDecryptsAnEncryptedCacheToPlaintext() = runBlocking<Unit> {
every { settingsRepository.settings } returns flowOf(AppSettings(encryptCache = false, appLock = false))
@@ -64,6 +64,56 @@ class MigrationTest {
}
}
/**
* v19 -> v20 (issue #187): the unified-inbox covering index appears over exactly
* `(folder, inInbox, timestampMillis)`, cached rows survive, and the paged "All inboxes" query now
* plans as a bounded `SEARCH` on that index with no temp B-tree sort (instead of a whole-table
* `SCAN`). Asserting the query plan on the real Android SQLite proves the index is genuinely
* covering the filter+order, not merely present.
*/
@Test
fun migrate19To20_addsUnifiedInboxCoveringIndexUsedByTheSummaryScan() {
helper.createDatabase(TEST_DB, 19).apply {
// A representative spread: two synced INBOX rows, plus a transient search hit (inInbox = 0)
// — all must survive the pure additive index migration. Fold columns default to ''.
execSQL(
"INSERT INTO messages (id, accountId, sender, senderEmail, subject, snippet, body, isHtml, " +
"timestampMillis, isRead, isStarred, folder, inInbox, bodyFetched, uid) VALUES " +
"('a:INBOX:2', 'a', 'Ada', 'ada@example.org', 'Hi', '', '', 0, 2000, 0, 0, 'INBOX', 1, 1, 2), " +
"('b:INBOX:1', 'b', 'Bob', 'bob@example.org', 'Yo', '', '', 0, 1000, 0, 0, 'INBOX', 1, 1, 1), " +
"('a:INBOX:9', 'a', 'Cy', 'cy@example.org', 'Q', '', '', 0, 3000, 0, 0, 'INBOX', 0, 0, 9)",
)
close()
}
val db = helper.runMigrationsAndValidate(TEST_DB, 20, true, MIGRATION_19_20)
// The index exists over exactly (folder, inInbox, timestampMillis), in that order.
assertEquals(
"19->20 must create the (folder, inInbox, timestampMillis) unified-inbox covering index",
listOf("folder", "inInbox", "timestampMillis"),
db.indexColumns("index_messages_folder_inInbox_timestampMillis"),
)
// The cached rows are untouched by the additive migration.
assertEquals("19->20 must not touch the mail cache", 3, db.count("messages"))
// The production pagingUnifiedFolderSummaries query now SEARCHes the new index and drops the
// temp B-tree sort (before this index it SCANned index_messages_timestampMillis whole-table).
val plan = db.queryPlan(
"SELECT id, accountId, sender, senderEmail, subject, snippet, timestampMillis, isRead, " +
"isStarred, folder, inInbox, bodyFetched FROM messages " +
"WHERE folder = 'INBOX' AND inInbox = 1 ORDER BY timestampMillis DESC",
)
assertTrue(
"the unified-inbox summary query must SEARCH the covering index, not SCAN; plan was $plan",
plan.any { it.contains("SEARCH") && it.contains("index_messages_folder_inInbox_timestampMillis") },
)
assertTrue(
"the covering index must supply the ordering (no temp B-tree sort); plan was $plan",
plan.none { it.contains("TEMP B-TREE") },
)
db.close()
}
/** v11 -> v12 (PR #54): `folders.specialUse` appears defaulting to 0 and existing data survives. */
@Test
fun migrate11To12_defaultsExistingFoldersToNotSpecialUse() {
@@ -576,6 +626,22 @@ class MigrationTest {
c.getInt(0)
}
/** Column names of [index], in index (seqno) order — empty if the index does not exist. */
private fun SupportSQLiteDatabase.indexColumns(index: String): List<String> =
query("PRAGMA index_info(`$index`)").use { c ->
buildList {
// PRAGMA index_info rows are (seqno, cid, name); the cursor yields them in seqno order.
while (c.moveToNext()) add(c.getString(2))
}
}
/** The human-readable `detail` step of each `EXPLAIN QUERY PLAN [sql]` row (the last column). */
private fun SupportSQLiteDatabase.queryPlan(sql: String): List<String> = query("EXPLAIN QUERY PLAN $sql").use { c ->
buildList {
while (c.moveToNext()) add(c.getString(c.columnCount - 1))
}
}
private companion object {
const val TEST_DB = "migration-test.db"
@@ -0,0 +1,205 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.data.local
import android.content.Context
import android.os.Build
import android.system.Os
import android.system.OsConstants
import androidx.room.Room
import androidx.test.core.app.ApplicationProvider
import androidx.test.ext.junit.runners.AndroidJUnit4
import kotlinx.coroutines.flow.first
import kotlinx.coroutines.runBlocking
import net.zetetic.database.sqlcipher.SQLiteDatabase
import net.zetetic.database.sqlcipher.SupportOpenHelperFactory
import org.junit.After
import org.junit.Assert.assertEquals
import org.junit.Assert.assertTrue
import org.junit.Before
import org.junit.Test
import org.junit.runner.RunWith
import org.libremail.data.local.entity.MessageEntity
import org.libremail.reporting.AppLog
import java.io.File
import java.io.PrintWriter
import java.io.StringWriter
/**
* SPIKE for issue #359 — characterizes, on a real device, whether opening the opt-in SQLCipher-encrypted
* cache throws `UnsatisfiedLinkError` because the bundled `libsqlcipher.so` is not compatible with the
* device's memory **page size**. Android 15+/SDK-37 devices may run **16 KB pages**; a native `.so` not
* built/aligned for 16 KB pages fails to load (`dlopen`) or to bind its JNI methods, surfacing as
* `UnsatisfiedLinkError` at `System.loadLibrary("sqlcipher")` or at `SQLiteConnection.nativeOpen`.
*
* This is investigation-only: it does not change production behaviour. It runs the exact #359 open path in
* three independent stages so a failing run tells us **which** stage breaks, and it records the device page
* size ([Os.sysconf] `_SC_PAGESIZE`) so a pass/fail can be tied to 4 KB vs 16 KB pages. It is meant to be
* run twice on the SAME device — once in 4 KB mode, once in 16 KB mode (Pixel Developer Options toggle) —
* to give a definitive A/B: if every stage passes at 4 KB and fails at 16 KB, 16 KB pages are the cause.
*
* On success each stage asserts the encrypted DB genuinely opens and round-trips a row. On failure each
* stage re-raises the full throwable — class, message (which `.so`), whether a [LinkageError] is in the
* cause chain, the page size, and the complete stack trace — so the A/B report captures the real cause
* rather than a bare assertion. All data is synthetic; nothing logged or asserted is PII.
*/
@RunWith(AndroidJUnit4::class)
class SqlCipherOpenSpikeTest {
private val context = ApplicationProvider.getApplicationContext<Context>()
private val dbName = "sqlcipher_spike_test.db"
private val dbFile: File get() = context.getDatabasePath(dbName)
private val probeDbName = "sqlcipher_spike_probe.db"
private val probeDbFile: File get() = context.getDatabasePath(probeDbName)
// 64 hex chars == a 32-byte SQLCipher passphrase, matching DatabaseKeyStore's format.
private val passphrase = "0123456789abcdef".repeat(4)
@Before
@After
fun clean() {
listOf(dbName, probeDbName).forEach { name ->
context.deleteDatabase(name)
context.getDatabasePath(name).parentFile
?.listFiles { f -> f.name.startsWith(name) }
?.forEach { it.delete() }
}
}
/**
* Stage A — load the SQLCipher native library the way production does
* ([DatabaseEncryption.ensureNativeLibraryLoaded] -> `System.loadLibrary("sqlcipher")`). This is the
* first place a 16 KB-incompatible `.so` can fail (`dlopen` rejects an unaligned library).
*/
@Test
fun stageA_sqlCipherNativeLibraryLoads() {
AppLog.i(TAG, "stageA start: $environment")
try {
DatabaseEncryption.ensureNativeLibraryLoaded()
} catch (t: Throwable) {
surface("A/loadLibrary(\"sqlcipher\")", t)
}
AppLog.i(TAG, "stageA PASS: SQLCipher native library loaded; $environment")
}
/**
* Stage B — reach `SQLiteConnection.nativeOpen`: after loading the library (as production does), open a
* keyed SQLCipher database and round-trip a row through the cipher. This is the exact call site named in
* the #359 crash (`UnsatisfiedLinkError … SQLiteConnection.nativeOpen`).
*/
@Test
fun stageB_keyedNativeOpenSucceeds() {
AppLog.i(TAG, "stageB start: $environment")
try {
DatabaseEncryption.ensureNativeLibraryLoaded()
val db = SQLiteDatabase.openOrCreateDatabase(
probeDbFile.absolutePath,
passphrase.toByteArray(Charsets.US_ASCII),
null, // no CursorFactory
null, // no DatabaseErrorHandler
)
try {
db.execSQL("CREATE TABLE IF NOT EXISTS spike(x INTEGER)")
db.execSQL("INSERT INTO spike(x) VALUES (42)")
db.rawQuery("SELECT x FROM spike LIMIT 1", null).use { cursor ->
assertTrue("keyed DB returned no row", cursor.moveToFirst())
assertEquals("keyed DB round-trip mismatch", 42, cursor.getInt(0))
}
} finally {
db.close()
}
} catch (t: Throwable) {
surface("B/SQLiteConnection.nativeOpen (keyed open)", t)
}
AppLog.i(TAG, "stageB PASS: keyed nativeOpen + round-trip OK; $environment")
}
/**
* Stage C — the full #359 production path: create a plaintext Room cache with one row, convert it to
* SQLCipher ciphertext ([DatabaseEncryption.ensureEncrypted]), load the library, then reopen the cache
* through Room's [SupportOpenHelperFactory] (exactly [org.libremail.di.DatabaseModule]'s encrypted open
* lambda) and read the seeded row back.
*/
@Test
fun stageC_encryptedRoomCacheOpensThroughProductionFactory() {
AppLog.i(TAG, "stageC start: $environment")
try {
Room.databaseBuilder(context, LibreMailDatabase::class.java, dbName).build().apply {
runBlocking { messageDao().insertNew(listOf(message("acct:1"))) }
close()
}
DatabaseEncryption.ensureEncrypted(dbFile, passphrase)
assertTrue("precondition: fixture must be genuinely encrypted", DatabaseEncryption.isEncrypted(dbFile))
DatabaseEncryption.ensureNativeLibraryLoaded()
val database = Room.databaseBuilder(context, LibreMailDatabase::class.java, dbName)
.openHelperFactory(SupportOpenHelperFactory(passphrase.toByteArray(Charsets.US_ASCII), null, false))
.build()
try {
val ids = runBlocking { database.messageDao().observeSummaries().first().map { it.id } }
assertEquals("encrypted cache did not read the seeded row back", listOf("acct:1"), ids)
} finally {
database.close()
}
} catch (t: Throwable) {
surface("C/Room encrypted cache open (SupportOpenHelperFactory)", t)
}
AppLog.i(TAG, "stageC PASS: encrypted Room cache opened through production factory; $environment")
}
/** A one-line, PII-free description of the device + page size every stage stamps into its log/report. */
private val environment: String
get() = "PAGE_SIZE=${pageSizeBytes()} bytes (16384 => 16 KB pages), SDK=${Build.VERSION.SDK_INT}, " +
"release=${Build.VERSION.RELEASE}, abis=${Build.SUPPORTED_ABIS.joinToString(",")}"
private fun pageSizeBytes(): Long = Os.sysconf(OsConstants._SC_PAGESIZE)
/**
* Fails the stage while surfacing the complete cause so the on-device A/B report captures the real
* `UnsatisfiedLinkError` (which `.so`, full stack trace, page size) instead of a bare assertion.
*/
private fun surface(stage: String, t: Throwable): Nothing {
val stack = StringWriter().also { t.printStackTrace(PrintWriter(it)) }.toString()
val chain = buildString {
var current: Throwable? = t
while (current != null) {
append("\n - ").append(current.javaClass.name).append(": ").append(current.message)
current = current.cause
}
}
val diagnostic = buildString {
append("SQLCipher spike stage '").append(stage).append("' FAILED on this device.")
append("\n ").append(environment)
append("\n LinkageError in cause chain = ").append(hasLinkageError(t))
append("\n cause chain:").append(chain)
append("\n full stack trace:\n").append(stack)
}
AppLog.e(TAG, "SQLCipher spike stage '$stage' FAILED; $environment", t)
throw AssertionError(diagnostic, t)
}
private fun hasLinkageError(throwable: Throwable): Boolean {
var current: Throwable? = throwable
while (current != null) {
if (current is LinkageError) return true
current = current.cause
}
return false
}
private fun message(id: String) = MessageEntity(
id = id,
accountId = "acct",
sender = "Ada",
senderEmail = "ada@example.org",
subject = "Hi",
snippet = "",
body = "",
timestampMillis = 1_000L,
isRead = false,
isStarred = false,
)
private companion object {
const val TAG = "SqlCipherOpenSpike"
}
}
@@ -0,0 +1,182 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.debug
import android.content.BroadcastReceiver
import android.content.ComponentName
import android.content.Context
import android.content.Intent
import androidx.test.core.app.ApplicationProvider
import androidx.test.ext.junit.runners.AndroidJUnit4
import androidx.work.ListenableWorker.Result
import androidx.work.WorkerFactory
import androidx.work.WorkerParameters
import androidx.work.testing.TestListenableWorkerBuilder
import dagger.Lazy
import io.mockk.coEvery
import io.mockk.every
import io.mockk.mockk
import io.mockk.unmockkAll
import io.mockk.verify
import kotlinx.coroutines.flow.flowOf
import kotlinx.coroutines.runBlocking
import kotlinx.coroutines.withTimeout
import org.junit.After
import org.junit.Assert.assertEquals
import org.junit.Assert.assertFalse
import org.junit.Assert.assertTrue
import org.junit.Before
import org.junit.Test
import org.junit.runner.RunWith
import org.libremail.data.security.EncryptedCacheGuard
import org.libremail.data.security.PassphraseSession
import org.libremail.data.settings.AppSettings
import org.libremail.data.settings.SettingsRepository
import org.libremail.data.sync.BackfillWorker
import org.libremail.data.sync.DebugFetchGate
import org.libremail.data.sync.FetchScope
import org.libremail.data.sync.MailBackfiller
import java.util.concurrent.CountDownLatch
import java.util.concurrent.TimeUnit
/**
* On-device proof of the debug-only fetch gate (issue #393): the adb-reachable [FetchGateReceiver]
* updates [DebugFetchGate] and returns the resulting state as ordered-broadcast result data (exactly
* what `adb shell am broadcast ... FETCH_GATE` prints back to the harness), and a gated proactive path
* ([BackfillWorker]) genuinely defers while an un-gated path keeps running. The broadcast is sent
* ordered — the same delivery mode `am broadcast` uses — so [BroadcastReceiver.getResultData] on the
* final receiver reads back what the gate set, with no logcat race.
*
* The worker-deferral cases reuse `WorkerCacheLockDeferralInstrumentedTest`'s approach: build a
* [BackfillWorker] with a real, never-unlocked-or-off [EncryptedCacheGuard] and a `Lazy` [MailBackfiller]
* whose resolution is observable, so "the gate deferred before touching the DB" is proven by the `Lazy`
* never being resolved.
*/
@RunWith(AndroidJUnit4::class)
class FetchGateReceiverInstrumentedTest {
private val context: Context = ApplicationProvider.getApplicationContext()
// A fresh, real, never-unlocked session per test — so the real guard reports UNLOCKED only because
// app-lock is off (see [unlockedGuard]), never because of leftover auth state.
private val session = PassphraseSession()
@Before
@After
fun resetGate() {
DebugFetchGate.reset()
unmockkAll()
}
@Test
fun pauseUpdatesTheGateAndReturnsTheReadBack() {
val data = sendGateBroadcast(FetchGateReceiver.ACTION_PAUSE, "backfill,prefetch")
assertEquals("paused=[backfill,prefetch]", data)
assertTrue(DebugFetchGate.isPaused(FetchScope.BACKFILL))
assertTrue(DebugFetchGate.isPaused(FetchScope.PREFETCH))
}
@Test
fun resumeAllClearsTheGateAndReturnsAnEmptyReadBack() {
sendGateBroadcast(FetchGateReceiver.ACTION_PAUSE, "all")
val data = sendGateBroadcast(FetchGateReceiver.ACTION_RESUME, "all")
assertEquals("paused=[]", data)
assertFalse(DebugFetchGate.isPaused(FetchScope.BACKFILL))
assertFalse(DebugFetchGate.isPaused(FetchScope.PREFETCH))
}
@Test
fun queryReadsBackTheStateWithoutMutatingIt() {
sendGateBroadcast(FetchGateReceiver.ACTION_PAUSE, "backfill")
val data = sendGateBroadcast(FetchGateReceiver.ACTION_QUERY, scope = null)
assertEquals("paused=[backfill]", data)
assertTrue(DebugFetchGate.isPaused(FetchScope.BACKFILL))
assertFalse(DebugFetchGate.isPaused(FetchScope.PREFETCH))
}
@Test
fun aBackfillPausedGateDefersTheBackfillWorkerWithoutResolvingTheBackfiller() = runBlocking<Unit> {
sendGateBroadcast(FetchGateReceiver.ACTION_PAUSE, "backfill")
val lazyBackfiller = mockk<Lazy<MailBackfiller>>()
val worker = TestListenableWorkerBuilder<BackfillWorker>(context)
.setWorkerFactory(backfillWorkerFactory(lazyBackfiller, unlockedGuard()))
.build()
val result = withTimeout(TIMEOUT_MS) { worker.doWork() }
assertEquals(Result.retry(), result)
// The gate deferred BEFORE any DB-backed work — the Lazy was never resolved.
verify(exactly = 0) { lazyBackfiller.get() }
}
@Test
fun aPrefetchOnlyPauseLeavesTheBackfillWorkerRunning() = runBlocking<Unit> {
// The worker gate honours BACKFILL only; pausing PREFETCH must NOT defer history paging — the
// on-device analogue of "on-demand open and header sync stay live while prefetch is paused".
sendGateBroadcast(FetchGateReceiver.ACTION_PAUSE, "prefetch")
val backfiller = mockk<MailBackfiller> { coEvery { runBackfill(any()) } returns false }
val lazyBackfiller = mockk<Lazy<MailBackfiller>> { every { get() } returns backfiller }
val worker = TestListenableWorkerBuilder<BackfillWorker>(context)
.setWorkerFactory(backfillWorkerFactory(lazyBackfiller, unlockedGuard()))
.build()
val result = withTimeout(TIMEOUT_MS) { worker.doWork() }
assertEquals(Result.success(), result)
verify { lazyBackfiller.get() }
}
/**
* Sends the [FetchGateReceiver.ACTION] broadcast to the receiver by explicit component (mirroring
* `am broadcast -n`), ordered, and returns the result data the receiver set (the harness read-back).
*/
private fun sendGateBroadcast(action: String, scope: String?): String {
val latch = CountDownLatch(1)
val readBack = arrayOfNulls<String>(1)
val intent = Intent(FetchGateReceiver.ACTION).apply {
component = ComponentName(context, FetchGateReceiver::class.java)
putExtra(FetchGateReceiver.EXTRA_ACTION, action)
if (scope != null) putExtra(FetchGateReceiver.EXTRA_SCOPE, scope)
}
context.sendOrderedBroadcast(
intent,
null,
object : BroadcastReceiver() {
override fun onReceive(c: Context, i: Intent) {
readBack[0] = resultData
latch.countDown()
}
},
null,
0,
null,
null,
)
assertTrue("gate broadcast timed out", latch.await(TIMEOUT_MS, TimeUnit.MILLISECONDS))
return requireNotNull(readBack[0]) { "receiver set no result data" }
}
/** A real [EncryptedCacheGuard] reporting UNLOCKED (app-lock off) — so only the gate can defer. */
private fun unlockedGuard(): EncryptedCacheGuard {
val settingsRepository = mockk<SettingsRepository>()
every { settingsRepository.settings } returns flowOf(AppSettings(appLock = false, encryptCache = true))
return EncryptedCacheGuard(settingsRepository, session)
}
private fun backfillWorkerFactory(lazyBackfiller: Lazy<MailBackfiller>, cacheGuard: EncryptedCacheGuard) =
object : WorkerFactory() {
override fun createWorker(
appContext: Context,
workerClassName: String,
workerParameters: WorkerParameters,
) = BackfillWorker(appContext, workerParameters, lazyBackfiller, cacheGuard)
}
private companion object {
const val TIMEOUT_MS = 5_000L
}
}
@@ -0,0 +1,109 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.di
import android.content.Context
import android.content.ContextWrapper
import androidx.test.core.app.ApplicationProvider
import androidx.test.ext.junit.runners.AndroidJUnit4
import io.mockk.Runs
import io.mockk.coEvery
import io.mockk.every
import io.mockk.just
import io.mockk.mockk
import io.mockk.mockkObject
import io.mockk.unmockkAll
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.flow.flowOf
import kotlinx.coroutines.runBlocking
import org.junit.After
import org.junit.Assert.assertEquals
import org.junit.Before
import org.junit.Test
import org.junit.runner.RunWith
import org.libremail.data.local.AccountDataMigrator
import org.libremail.data.local.DatabaseEncryption
import org.libremail.data.local.DatabaseFiles
import org.libremail.data.local.DatabaseProvisioner
import org.libremail.data.security.DatabaseKeyStore
import org.libremail.data.settings.AppSettings
import org.libremail.data.settings.SettingsRepository
import java.io.File
/**
* Pins the fail-closed contract's ONE resilience exception (issue #359): the plaintext account store is
* never encrypted and never uses SQLCipher, so a cache-encryption native-load failure — which the
* provisioner surfaces as `CacheEncryptionUnavailableException` — must NOT brick it. If it did, the app
* couldn't read accounts to render the encryption error gate or assemble the PII-free problem report.
*
* Mirrors [DatabaseModuleInstrumentedTest]'s style: MockK collaborators, a real [ContextWrapper] (never
* `mockk<Context>()`, which trips an ART parameter-annotation mismatch on API 31/32), and a real
* [DatabaseProvisioner] whose encryption gate is forced to fail via a spied [DatabaseEncryption].
*/
@RunWith(AndroidJUnit4::class)
class AccountDatabaseModuleInstrumentedTest {
private val appContext = ApplicationProvider.getApplicationContext<Context>()
private val cacheDbName = "accountmodule_cache_test.db"
private val accountsDbName = "accountmodule_accounts_test.db"
private val cacheFile: File get() = appContext.getDatabasePath(cacheDbName)
private val accountsFile: File get() = appContext.getDatabasePath(accountsDbName)
// 64 hex chars == a 32-byte SQLCipher passphrase.
private val passphrase = "0123456789abcdef".repeat(4)
private val keyStore = mockk<DatabaseKeyStore>()
private val settingsRepository = mockk<SettingsRepository>()
private val migrator = mockk<AccountDataMigrator>()
// Route the provisioner's cache lookup and Room's account-store lookup to this test's private files,
// never the app's real databases.
private val context: Context = object : ContextWrapper(appContext) {
override fun getDatabasePath(name: String): File = when (name) {
DatabaseFiles.NAME -> cacheFile
DatabaseFiles.ACCOUNTS_NAME -> accountsFile
else -> super.getDatabasePath(name)
}
}
@Before
fun setUp() {
clean()
coEvery { keyStore.isClearPending() } returns false
coEvery { keyStore.resolvePassphrase(any()) } returns passphrase
coEvery { migrator.migrateIfNeeded() } just Runs
}
@After
fun tearDown() {
unmockkAll()
clean()
}
private fun clean() {
listOf(cacheDbName, accountsDbName).forEach { name ->
appContext.deleteDatabase(name)
appContext.getDatabasePath(name).parentFile?.listFiles { f -> f.name.startsWith(name) }
?.forEach { it.delete() }
}
}
private fun provisioner() = DatabaseProvisioner(context, keyStore, settingsRepository, migrator, Dispatchers.IO)
@Test
fun accountStoreStillOpensWhenTheCacheEncryptionLibraryFailsToLoad() = runBlocking<Unit> {
every { settingsRepository.settings } returns flowOf(AppSettings(encryptCache = true, appLock = false))
mockkObject(DatabaseEncryption) // spy: real impls run except the forced failure below
// Fault injection: the encrypted-cache gate can't load SQLCipher, so prepareCache() fails closed
// with CacheEncryptionUnavailableException — the exact condition provideAccountDatabase tolerates.
every { DatabaseEncryption.ensureNativeLibraryLoaded() } throws
UnsatisfiedLinkError("dlopen failed: libsqlcipher.so is not loadable")
val database = AccountDatabaseModule.provideAccountDatabase(context, provisioner())
try {
// The plaintext account store opens and a query succeeds despite the cache-encryption failure.
assertEquals(emptyList<Any>(), database.accountDao().getAll())
} finally {
database.close()
}
}
}
@@ -0,0 +1,87 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.push
import android.app.Notification
import android.app.Service
import android.content.Context
import androidx.test.core.app.ApplicationProvider
import androidx.test.ext.junit.runners.AndroidJUnit4
import org.junit.Assert.assertEquals
import org.junit.Assert.assertFalse
import org.junit.Assert.assertSame
import org.junit.Assert.assertTrue
import org.junit.Test
import org.junit.runner.RunWith
import org.libremail.R
import org.libremail.data.sync.PushMode
/**
* On-device coverage of the #354 dataSync-FGS degrade path that [IdleService.onStartCommand] routes
* through [IdleForegroundStarter]. When a foreground start is rejected — the runtime-cap
* `ForegroundServiceStartNotAllowedException`, surfaced as its [IllegalStateException] supertype — the
* seam must catch it, skip IDLE watching, and degrade to periodic sync plus the degraded
* ("instant delivery paused") notification, never propagating. This drives the same decision seam the
* service uses and builds the real degraded notification with a real application `Context` (a
* `ContextWrapper`, never a mocked `Context`), mirroring `PushStatusNotificationInstrumentedTest`; it
* stands up no foreground service, Hilt graph, or network, so it is deterministic — and unlike a JVM
* unit test it exercises the real `Notification` build (the unit-test `android.jar`'s
* `NotificationCompat` is a no-op stub).
*/
@RunWith(AndroidJUnit4::class)
class IdleServiceForegroundStartInstrumentedTest {
private val context = ApplicationProvider.getApplicationContext<Context>()
@Test
fun rejectedForegroundStart_degradesToPeriodicSyncWithPausedNotification_andSkipsWatching() {
val rejection = IllegalStateException(
"Time limit already exhausted for foreground service type dataSync",
)
var watchingStarted = false
var periodicSyncScheduled = false
var degradedNotification: Notification? = null
val result = IdleForegroundStarter.startForegroundOrDegrade(
capActive = false,
enterForeground = { throw rejection },
onStarted = { watchingStarted = true },
onDegraded = { cause ->
assertSame("the runtime-cap rejection must reach the degrade path", rejection, cause)
// Mirror IdleService.degradeToPeriodicSync on a real Context: (re)assert periodic sync and
// build the degraded status notification the service would post.
periodicSyncScheduled = true
PushStatusNotification.ensureChannel(context)
degradedNotification = PushStatusNotification.build(context, PushMode.POLLING, timedOut = true)
},
)
assertEquals(Service.START_NOT_STICKY, result)
assertFalse("a rejected dataSync FGS start must not begin IDLE watching", watchingStarted)
assertTrue("the degrade path must (re)assert the 15-minute periodic sync fallback", periodicSyncScheduled)
val notification = requireNotNull(degradedNotification) { "the degrade path must build a status notification" }
assertEquals(
"the degraded notification must show the instant-delivery-paused text",
context.getString(R.string.notif_push_status_text_timed_out),
notification.extras.getCharSequence(Notification.EXTRA_TEXT).toString(),
)
}
@Test
fun activeCapWindow_skipsForegroundStartAttempt_andStillDegrades() {
var enterForegroundAttempted = false
var watchingStarted = false
var degraded = false
val result = IdleForegroundStarter.startForegroundOrDegrade(
capActive = true,
enterForeground = { enterForegroundAttempted = true },
onStarted = { watchingStarted = true },
onDegraded = { degraded = true },
)
assertEquals(Service.START_NOT_STICKY, result)
assertFalse("must not attempt a dataSync FGS start while still inside the cap window", enterForegroundAttempted)
assertFalse(watchingStarted)
assertTrue("must fall back to periodic sync while capped", degraded)
}
}
@@ -0,0 +1,95 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.reporting
import androidx.test.ext.junit.runners.AndroidJUnit4
import androidx.test.platform.app.InstrumentationRegistry
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Dispatchers
import org.junit.After
import org.junit.Assert.assertEquals
import org.junit.Assert.assertFalse
import org.junit.Assert.assertTrue
import org.junit.Before
import org.junit.Test
import org.junit.runner.RunWith
import org.libremail.data.security.KeystoreCrypto
import java.io.File
/**
* On-device proof for issue #369: with at-rest encryption ON, [ReportStore] persists a report as real
* Android Keystore ciphertext — no report content in plaintext on disk — and reads it back intact;
* with it OFF the file stays plaintext JSON. This closes the gap between the JVM-tested ReportStore
* branching (which fakes the cipher) and the device-only [KeystoreCrypto] the branching drives in
* production, using the same non-auth master key that lets a crash-while-locked report still be sealed.
*/
@RunWith(AndroidJUnit4::class)
class ReportStoreEncryptionInstrumentedTest {
private val context =
InstrumentationRegistry.getInstrumentation().targetContext.applicationContext
private val dir = File(context.cacheDir, "report-encryption-test")
@Before
fun setUp() {
dir.deleteRecursively()
}
@After
fun tearDown() {
dir.deleteRecursively()
}
@Test
fun encryptedReportIsCiphertextOnDiskAndReadsBack() {
val store = store(enabled = true)
store.save(report("enc"))
val raw = File(dir, "enc.json").readText()
// The distinctive plaintext token must NOT be on disk — the report is Keystore-sealed at rest.
assertFalse("report content must not be persisted in plaintext", raw.contains(SENTINEL))
assertFalse("a sealed report is not plaintext JSON", raw.startsWith("{"))
// A fresh store over the same directory (same master key) unseals and reads it back intact.
assertEquals(SENTINEL, store(enabled = true).find("enc")?.logs?.single())
}
@Test
fun plaintextReportWhenEncryptionOff() {
val store = store(enabled = false)
store.save(report("plain"))
val raw = File(dir, "plain.json").readText()
assertTrue("with encryption off the report stays plaintext JSON", raw.contains(SENTINEL))
assertTrue(raw.startsWith("{"))
}
private fun store(enabled: Boolean): ReportStore {
val crypto = KeystoreCrypto()
val encryption = object : ReportEncryption {
override fun enabled(): Boolean = enabled
override fun encrypt(plaintext: String): String = crypto.encrypt(plaintext)
override fun decrypt(encoded: String): String = crypto.decrypt(encoded)
}
return ReportStore(dir, CoroutineScope(Dispatchers.Unconfined), encryption)
}
private fun report(id: String) = DebugReport(
id = id,
createdAtMillis = 1_000L,
kind = ReportKind.CRASH,
appVersionName = "1.0",
appVersionCode = 1L,
androidRelease = "14",
androidSdkInt = 34,
deviceManufacturer = "Test",
deviceModel = "Model",
stackTrace = SENTINEL,
settings = emptyMap(),
logs = listOf(SENTINEL),
)
private companion object {
const val SENTINEL = "SENTINEL-PLAINTEXT-TOKEN-369"
}
}
@@ -15,6 +15,7 @@ import androidx.compose.ui.test.onNodeWithText
import androidx.compose.ui.test.performClick
import androidx.compose.ui.test.performTextInput
import androidx.lifecycle.SavedStateHandle
import androidx.lifecycle.ViewModelStore
import androidx.room.Room
import androidx.test.ext.junit.runners.AndroidJUnit4
import androidx.test.platform.app.InstrumentationRegistry
@@ -60,10 +61,20 @@ class ComposeScreenTest {
private var db: AccountDatabase? = null
// Holds the real ComposeViewModel built by hand in setContent() below, so closeDb() can clear()
// it (triggering ViewModel.onCleared()) before closing the DB.
private val viewModelStore = ViewModelStore()
private fun string(resId: Int) = composeTestRule.activity.getString(resId)
@After
fun closeDb() {
// Clear the store (→ ViewModel.onCleared() → cancels viewModelScope) BEFORE closing the DB.
// ComposeViewModel's init block launches a viewModelScope coroutine that reads the real
// accountSettings/signature Room repositories (applySignature()); without this, that read can
// still be in flight when the DB closes, racing a SQLITE_MISUSE ("connection is closed") —
// the same class of teardown race fixed in SignaturesScreenTest/AccountSettingsScreenTest.
viewModelStore.clear()
db?.close()
}
@@ -92,6 +103,7 @@ class ComposeScreenTest {
signatureRepository = SignatureRepository(database.signatureDao()),
settingsRepository = SettingsRepository(context),
)
viewModelStore.put("compose", viewModel)
composeTestRule.setContent {
LibreMailTheme(darkTheme = false, dynamicColor = false) {
ComposeScreen(onBack = onBack, viewModel = viewModel)
@@ -0,0 +1,55 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.ui.security
import androidx.activity.ComponentActivity
import androidx.compose.ui.test.assertIsDisplayed
import androidx.compose.ui.test.junit4.createAndroidComposeRule
import androidx.compose.ui.test.onNodeWithText
import androidx.compose.ui.test.performClick
import androidx.test.ext.junit.runners.AndroidJUnit4
import org.junit.Rule
import org.junit.Test
import org.junit.runner.RunWith
import org.libremail.R
import org.libremail.ui.theme.LibreMailTheme
/**
* UI coverage for the fail-closed encrypted-cache error screen (issue #359). [CacheEncryptionErrorScreen]
* is presentational (its report action is wired by the caller), so it is exercised in isolation: the
* verbatim error message shows, and tapping "Report a problem" reports back.
*/
@RunWith(AndroidJUnit4::class)
class CacheEncryptionErrorScreenTest {
@get:Rule
val composeTestRule = createAndroidComposeRule<ComponentActivity>()
private fun string(resId: Int) = composeTestRule.activity.getString(resId)
private fun setContent(onReportProblem: () -> Unit = {}) {
composeTestRule.setContent {
LibreMailTheme(darkTheme = false, dynamicColor = false) {
CacheEncryptionErrorScreen(onReportProblem = onReportProblem)
}
}
}
@Test
fun showsTheVerbatimErrorMessageAndReportAction() {
setContent()
// The exact maintainer-specified message must render, unchanged.
composeTestRule.onNodeWithText(string(R.string.cache_encryption_error_message)).assertIsDisplayed()
composeTestRule.onNodeWithText(string(R.string.cache_encryption_report_action)).assertIsDisplayed()
}
@Test
fun tappingReportProblem_invokesCallback() {
var reported = false
setContent(onReportProblem = { reported = true })
composeTestRule.onNodeWithText(string(R.string.cache_encryption_report_action)).performClick()
composeTestRule.waitUntil(5_000) { reported }
}
}
@@ -7,11 +7,13 @@ import androidx.compose.ui.test.junit4.createAndroidComposeRule
import androidx.compose.ui.test.onNodeWithText
import androidx.compose.ui.test.performClick
import androidx.lifecycle.SavedStateHandle
import androidx.lifecycle.ViewModelStore
import androidx.room.Room
import androidx.test.core.app.ApplicationProvider
import androidx.test.ext.junit.runners.AndroidJUnit4
import androidx.work.WorkManager
import kotlinx.coroutines.runBlocking
import org.junit.After
import org.junit.Rule
import org.junit.Test
import org.junit.runner.RunWith
@@ -54,13 +56,27 @@ class AccountSettingsScreenTest {
private var manageSignaturesClicked = false
private lateinit var db: AccountDatabase
// Holds the real AccountSettingsViewModel built by hand in setContent() below, so tearDown() can
// clear() it (triggering ViewModel.onCleared()) before closing the DB.
private val viewModelStore = ViewModelStore()
@After
fun tearDown() {
// Clear the store (→ ViewModel.onCleared() → cancels viewModelScope) BEFORE closing the DB.
// The ViewModel's `settings`/`signatureCount`/`defaultSignatureName`/`account` Room
// InvalidationTracker Flows are kept alive by stateIn(WhileSubscribed(5_000)): without this,
// a collector can still be live up to 5s after the UI detaches, so a re-query lands on the
// just-closed in-memory DB and throws SQLITE_MISUSE ("connection is closed") — an intermittent
// teardown race, not a real bug. (Previously worked around by never closing the DB at all.)
viewModelStore.clear()
db.close()
}
private fun setContent(): AccountSettingsRepository {
val context = ApplicationProvider.getApplicationContext<Context>()
// Intentionally not closed in an @After: the ViewModel's `settings` Room Flow (kept alive by
// stateIn/WhileSubscribed) keeps querying after the test body, so closing the in-memory DB out
// from under it races and crashes ("connection pool has been closed"). The DB is reclaimed with
// the test process.
val db = Room.inMemoryDatabaseBuilder(context, AccountDatabase::class.java).build()
db = Room.inMemoryDatabaseBuilder(context, AccountDatabase::class.java).build()
val repository = AccountSettingsRepository(db.accountSettingsDao())
runBlocking {
db.accountDao().upsert(account.toEntity()) // FK parent for the account_settings row
@@ -74,6 +90,7 @@ class AccountSettingsScreenTest {
syncScheduler = SyncScheduler(Provider { WorkManager.getInstance(context) }),
settingsRepository = SettingsRepository(context),
)
viewModelStore.put("account-settings", viewModel)
composeTestRule.setContent {
LibreMailTheme(darkTheme = false, dynamicColor = false) {
AccountSettingsScreen(
@@ -14,6 +14,7 @@ import androidx.compose.ui.test.onAllNodesWithText
import androidx.compose.ui.test.onNodeWithText
import androidx.compose.ui.test.performClick
import androidx.lifecycle.SavedStateHandle
import androidx.lifecycle.ViewModelStore
import androidx.room.Room
import androidx.test.ext.junit.runners.AndroidJUnit4
import androidx.test.platform.app.InstrumentationRegistry
@@ -50,6 +51,10 @@ class SignaturesScreenTest {
private lateinit var db: AccountDatabase
private lateinit var repository: SignatureRepository
// Holds the real SignaturesViewModel built by hand in setContent() below, so tearDown() can
// clear() it (triggering ViewModel.onCleared()) before closing the DB.
private val viewModelStore = ViewModelStore()
@Before
fun setUp() {
db = Room.inMemoryDatabaseBuilder(context, AccountDatabase::class.java).build()
@@ -70,7 +75,15 @@ class SignaturesScreenTest {
}
@After
fun tearDown() = db.close()
fun tearDown() {
// Clear the store (→ ViewModel.onCleared() → cancels viewModelScope) BEFORE closing the DB.
// SignaturesViewModel.signatures is a Room InvalidationTracker Flow kept alive by
// stateIn(WhileSubscribed(5_000)): without this, the collector can still be live up to 5s
// after the UI detaches, so a re-query lands on the just-closed in-memory DB and throws
// SQLITE_MISUSE ("connection is closed") — an intermittent teardown race, not a real bug.
viewModelStore.clear()
db.close()
}
private fun string(resId: Int) = composeTestRule.activity.getString(resId)
@@ -81,6 +94,7 @@ class SignaturesScreenTest {
SavedStateHandle(mapOf(Routes.SIGNATURES_ARG_ACCOUNT to accountId)),
repository,
)
viewModelStore.put("signatures", viewModel)
composeTestRule.setContent {
LibreMailTheme(darkTheme = false, dynamicColor = false) {
SignaturesScreen(onBack = {}, onEdit = {}, onAdd = {}, viewModel = viewModel)
+17 -1
View File
@@ -1,5 +1,6 @@
<!-- SPDX-License-Identifier: GPL-3.0-or-later -->
<manifest xmlns:android="http://schemas.android.com/apk/res/android">
<manifest xmlns:android="http://schemas.android.com/apk/res/android"
xmlns:tools="http://schemas.android.com/tools">
<!--
Debug-only test harness for issue #221 (never merged into a release APK — this manifest belongs to
@@ -16,6 +17,21 @@
android:authorities="${applicationId}.coldopen"
android:exported="false"
android:process=":coldopen" />
<!--
Debug-only fetch-gate receiver (issue #393; also never merged into a release APK — this
manifest belongs to the debug source set). Lets the on-device perf harness pause proactive
fetch (backfill + body prefetch) via `adb shell am broadcast` so a genuine uncached
message-open can be measured. Must be exported="true" so the adb `shell` UID can reach it by
explicit component (`-n`); it targets the debug BuildConfig.DEBUG-guarded DebugFetchGate only,
carries no PII, and — being debug-only — can never ship. tools:ignore suppresses the
exported-without-permission lint note: a signature permission would (by design) also lock out
the shell UID this hook exists to serve.
-->
<receiver
android:name="org.libremail.debug.FetchGateReceiver"
android:exported="true"
tools:ignore="ExportedReceiver" />
</application>
</manifest>
@@ -0,0 +1,71 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.debug
import android.content.BroadcastReceiver
import android.content.Context
import android.content.Intent
import org.libremail.data.sync.DebugFetchGate
import org.libremail.data.sync.FetchScope
import org.libremail.reporting.AppLog
/**
* Debug-only [BroadcastReceiver] (issue #393) that lets an adb-driven perf harness pause/resume the
* proactive-fetch activities tracked by [DebugFetchGate], so a genuinely uncached message-open can be
* measured (add an account, let headers sync, then open a message that must hit the network instead of a
* warmed cache). Declared **only** in `app/src/debug/AndroidManifest.xml`, so it is physically absent
* from every release APK — the same source-set guarantee `ColdOpenCacheProbe` (#221) relies on.
*
* Driven by (component targeted with `-n`, so it needs no `<intent-filter>`):
* ```
* adb shell am broadcast -a org.libremail.debug.FETCH_GATE \
* -n org.libremail.app/org.libremail.debug.FetchGateReceiver \
* --es action <pause|resume|query> --es scope <backfill,prefetch|all>
* ```
* `am broadcast` delivers this **ordered**, so the receiver returns the resulting state as result data
* (e.g. `data=paused=[backfill,prefetch]`) which `am` prints — a synchronous, race-free read-back for
* the harness. A `query` reports the current state without changing it. The new state is logged via the
* PII-free [AppLog] (scope names only — never an email, host, or message content).
*/
class FetchGateReceiver : BroadcastReceiver() {
override fun onReceive(context: Context, intent: Intent) {
val action = intent.getStringExtra(EXTRA_ACTION)?.trim()?.lowercase()
val scopes = FetchScope.parse(intent.getStringExtra(EXTRA_SCOPE))
when (action) {
ACTION_PAUSE -> {
DebugFetchGate.pause(scopes)
AppLog.i(TAG, "fetch gate pause -> ${DebugFetchGate.pausedResult()}")
}
ACTION_RESUME -> {
DebugFetchGate.resume(scopes)
AppLog.i(TAG, "fetch gate resume -> ${DebugFetchGate.pausedResult()}")
}
ACTION_QUERY -> AppLog.i(TAG, "fetch gate query -> ${DebugFetchGate.pausedResult()}")
else -> AppLog.w(TAG, "fetch gate: unknown action")
}
// Return the gate state as ordered-broadcast result data for a synchronous read-back. Guarded so
// a non-ordered send (which has no result receiver) can't crash the receiver.
if (isOrderedBroadcast) {
resultCode = RESULT_CODE
resultData = DebugFetchGate.pausedResult()
}
}
companion object {
/** The broadcast action the harness sends (kept for parity with the adb command; delivery is by `-n`). */
const val ACTION = "org.libremail.debug.FETCH_GATE"
/** `--es action <pause|resume|query>`. */
const val EXTRA_ACTION = "action"
/** `--es scope <comma-list|all>` (see [FetchScope.parse]). */
const val EXTRA_SCOPE = "scope"
const val ACTION_PAUSE = "pause"
const val ACTION_RESUME = "resume"
const val ACTION_QUERY = "query"
private const val TAG = "FetchGateReceiver"
private const val RESULT_CODE = 0
}
}
@@ -13,6 +13,7 @@ import kotlinx.coroutines.flow.combine
import kotlinx.coroutines.flow.distinctUntilChanged
import kotlinx.coroutines.flow.map
import kotlinx.coroutines.launch
import org.libremail.data.security.KeystoreReportEncryption
import org.libremail.data.settings.SettingsRepository
import org.libremail.data.sync.SyncScheduler
import org.libremail.domain.repository.AccountRepository
@@ -48,6 +49,8 @@ class LibreMailApplication :
@Inject lateinit var diagnosticsCollector: DiagnosticsCollector
@Inject lateinit var reportEncryption: KeystoreReportEncryption
private val appScope = CoroutineScope(SupervisorJob() + Dispatchers.Default)
/** Whether the IDLE push service should currently be running (push enabled AND an account exists). */
@@ -74,6 +77,10 @@ class LibreMailApplication :
// Warm the settings cache so a later crash report can include non-PII settings without
// touching DataStore on the crashing thread.
appScope.launch { runCatching { diagnosticsCollector.warmSettingsCache() } }
// Mirror the encryptCache setting so a crash-time report save (synchronous, on the crashing
// thread) can seal the report at rest without touching DataStore (#369). Collects for the
// process lifetime, so a mid-session toggle takes effect on the next report write.
appScope.launch { runCatching { reportEncryption.observeEncryptCacheSetting() } }
syncScheduler.schedulePeriodicSync()
// Full-history backfill (#12) and device-only retention pruning (#13) run as their own bounded,
// resumable background jobs so they never block foreground sync / pull-to-refresh.
@@ -21,6 +21,7 @@ import org.libremail.ui.LibreMailApp
import org.libremail.ui.compose.ComposePrefill
import org.libremail.ui.compose.IntentComposeParser
import org.libremail.ui.lock.AppLockGateHost
import org.libremail.ui.security.CacheEncryptionGate
import org.libremail.ui.theme.LibreMailTheme
import javax.inject.Inject
@@ -73,12 +74,18 @@ class MainActivity : FragmentActivity() {
// Gate the whole app behind the screen-lock when app-lock is enabled. When it is off
// the gate resolves straight to the content, so this is a no-op for most users.
AppLockGateHost {
LibreMailApp(
pendingCompose = pendingCompose.value,
onComposeHandled = { pendingCompose.value = null },
pendingOpenMessageId = pendingOpenMessageId.value,
onOpenMessageHandled = { pendingOpenMessageId.value = null },
)
// Inside the app-lock gate (so the auth-bound passphrase is already unlocked): fail
// closed if the encrypted cache's SQLCipher library won't load (#359), showing the
// error gate instead of ever opening the cache unencrypted. Resolves straight to the
// content when the cache is openable, so it is a no-op for most users.
CacheEncryptionGate {
LibreMailApp(
pendingCompose = pendingCompose.value,
onComposeHandled = { pendingCompose.value = null },
pendingOpenMessageId = pendingOpenMessageId.value,
onOpenMessageHandled = { pendingOpenMessageId.value = null },
)
}
}
}
}
@@ -106,12 +106,16 @@ class OutlookAuthManager @Inject constructor(@ApplicationContext private val con
val authState = AuthState(response, exception).apply { update(tokenResponse, null) }
val email = emailFromIdToken(tokenResponse.idToken)
?: throw IllegalStateException("Could not read the account email from the token")
// Mint an Exchange Online token so the caller can verify the account over IMAP.
val outlook = refreshForScope(authState, OUTLOOK_SCOPE)
// The code exchange above already named the Exchange Online resource ($OUTLOOK_SCOPE), so
// this access token is an outlook.office.com token the caller can verify over IMAP directly.
// Don't re-refresh for the same scope: that second round-trip only rotates the just-issued
// refresh token and adds a needless onboarding failure point. The Graph token is a different
// resource and is minted on demand later (freshGraphToken); the durable AuthState — refresh
// token plus this token's expiry — is serialized here for those later refreshes.
return OAuthResult(
email = email,
accessToken = outlook.accessToken,
authStateJson = outlook.authStateJson,
accessToken = tokenResponse.accessToken.orEmpty(),
authStateJson = authState.jsonSerializeString(),
)
} finally {
service.dispose()
@@ -2,7 +2,6 @@
package org.libremail.data.local
import android.content.Context
import android.util.Log
import androidx.datastore.core.DataStore
import androidx.datastore.preferences.core.Preferences
import androidx.datastore.preferences.core.booleanPreferencesKey
@@ -15,6 +14,7 @@ import kotlinx.coroutines.withContext
import net.zetetic.database.sqlcipher.SQLiteDatabase
import org.libremail.data.security.DatabaseKeyStore
import org.libremail.data.settings.SettingsRepository
import org.libremail.reporting.AppLog
import java.io.File
import javax.inject.Inject
import javax.inject.Singleton
@@ -181,7 +181,7 @@ class AccountDataMigrator @Inject constructor(
"(SELECT COUNT(*) FROM `accounts` AS ranked WHERE ranked.`email` < `accounts`.`email`)",
)
}
Log.d(TAG, "moved account tables into the account database: $present")
AppLog.d(TAG, "moved account tables into the account database: $present")
} finally {
db.rawExecSQL("DETACH DATABASE cache;")
}
@@ -0,0 +1,24 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.data.local
/**
* Raised by [DatabaseProvisioner] when the opt-in encrypted cache cannot be opened because SQLCipher's
* native library failed to load or link on this device (issue #359 — e.g. an `.so` the platform
* rejects, surfacing as an `UnsatisfiedLinkError`/`LinkageError` at `SQLiteConnection.nativeOpen` or
* from [DatabaseEncryption.ensureNativeLibraryLoaded]).
*
* The app must **fail closed**: it must NOT fall back to an unencrypted cache (that would silently
* defeat the user's opt-in encryption), NOT wipe the on-disk ciphertext, and NOT touch the
* `encryptCache` setting. Instead this distinct, expected signal is surfaced so the startup UI
* (`CacheEncryptionGate`) can show the encryption error gate — not the mailbox, and not a crash.
*
* Deliberately a dedicated type (not a bare [LinkageError]) so only this precise condition is treated
* as "encryption unavailable"; any other failure still propagates. The provisioner never memoizes it,
* so a later launch — where the library may load, e.g. after an app update — re-attempts and recovers
* automatically.
*/
class CacheEncryptionUnavailableException(cause: Throwable) :
Exception(
"Encrypted cache unavailable: the SQLCipher native library failed to load on this device",
cause,
)
@@ -1,8 +1,8 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.data.local
import android.util.Log
import net.zetetic.database.sqlcipher.SQLiteDatabase
import org.libremail.reporting.AppLog
import java.io.File
/**
@@ -40,6 +40,7 @@ object DatabaseEncryption {
* tables but not that pragma, and a reset version would make Room attempt a bogus migration.
*/
private fun migrate(dbFile: File, sourcePassphrase: String, targetPassphrase: String) {
AppLog.i(TAG, "converting local cache database (targetEncrypted=${targetPassphrase.isNotEmpty()})")
ensureNativeLibraryLoaded()
val dir = dbFile.parentFile ?: error("database file has no parent directory")
val tmp = File(dir, dbFile.name + ".migrate").apply { delete() }
@@ -88,7 +89,7 @@ object DatabaseEncryption {
tmp.copyTo(dbFile, overwrite = true)
tmp.delete()
}
Log.d(TAG, "local cache database converted")
AppLog.d(TAG, "local cache database converted")
}
private fun startsWithSqliteHeader(dbFile: File): Boolean {
@@ -10,7 +10,10 @@ import kotlinx.coroutines.sync.Mutex
import kotlinx.coroutines.sync.withLock
import kotlinx.coroutines.withContext
import org.libremail.data.security.DatabaseKeyStore
import org.libremail.data.settings.AppSettings
import org.libremail.data.settings.SettingsRepository
import org.libremail.reporting.AppLog
import java.io.File
import javax.inject.Inject
import javax.inject.Singleton
@@ -115,6 +118,35 @@ class DatabaseProvisioner internal constructor(
// resolvePassphrase waits on PassphraseSession until the user authenticates — which is why this
// must never run on the main thread while the cache is locked (issue #93).
val settings = settingsRepository.settings.first()
return try {
resolveOpenMode(settings, dbFile)
} catch (nativeLoadFailure: LinkageError) {
// FAIL CLOSED (issue #359, security rework of #367). SQLCipher's native library could not be
// loaded/linked (e.g. UnsatisfiedLinkError at SQLiteConnection.nativeOpen or from
// ensureNativeLibraryLoaded), so the encrypted cache cannot be opened OR converted. We must
// NOT silently degrade to an unencrypted cache (that would defeat the user's opt-in
// encryption), so we deliberately do NOT: open plaintext, wipe the on-disk ciphertext, or
// write the encryptCache setting. Instead raise a distinct signal the startup UI catches to
// show the encryption error gate. This throw is NOT memoized (it skips prepareCache's
// `.also { prepared = it }`), so the next launch re-attempts and recovers automatically if
// the library later loads.
AppLog.w(
TAG,
"SQLCipher native library failed to load; failing closed (encrypted cache unavailable)",
nativeLoadFailure,
)
throw CacheEncryptionUnavailableException(nativeLoadFailure)
}
}
/**
* The encryption gate (step 3 of [runStartupSequence]): convert the on-disk cache to the form the
* `encryptCache` setting asks for and report how Room must open it. Split out so a native-library
* load failure on either the encrypt or the decrypt-on-disable path (both need SQLCipher's `.so`) is
* caught in one place — see [runStartupSequence]'s handler, which fails closed by raising
* [CacheEncryptionUnavailableException] rather than degrading to an unencrypted cache.
*/
private suspend fun resolveOpenMode(settings: AppSettings, dbFile: File): CacheOpenMode {
val appLock = settings.appLock
return when {
settings.encryptCache -> {
@@ -140,4 +172,8 @@ class DatabaseProvisioner internal constructor(
else -> CacheOpenMode.Plaintext
}
}
private companion object {
const val TAG = "DatabaseProvisioner"
}
}
@@ -35,7 +35,7 @@ import org.libremail.data.local.entity.OutboxEntity
FolderEntity::class,
BackfillProgressEntity::class,
],
version = 19,
version = 20,
exportSchema = true,
)
abstract class LibreMailDatabase : RoomDatabase() {
@@ -380,3 +380,25 @@ val MIGRATION_18_19 = object : Migration(18, 19) {
)
}
}
/**
* v19 -> v20: covering index for the unified-inbox summary scan (issue #187; preserves existing data
* — a pure additive index, no column/table change or data transformation). The paged "All inboxes"
* query [org.libremail.data.local.dao.MessageDao.pagingUnifiedFolderSummaries] filters
* `folder = ? AND inInbox = 1 ORDER BY timestampMillis DESC`, but no index led with `folder`, so the
* planner walked the whole table via `index_messages_timestampMillis` in timestamp order and filtered
* `folder`/`inInbox` per row (a full `SCAN`, verified via `EXPLAIN QUERY PLAN`). The
* `(folder, inInbox, timestampMillis)` index turns the two equality predicates into an index seek and
* supplies the `timestampMillis` ordering, so the scan becomes a bounded `SEARCH … USING INDEX
* index_messages_folder_inInbox_timestampMillis (folder=? AND inInbox=?)` with no temp B-tree sort.
* `CREATE INDEX IF NOT EXISTS` is idempotent, and the name/columns match the Room `@Index` on
* [org.libremail.data.local.entity.MessageEntity] so the migrated schema validates against 20.json.
*/
val MIGRATION_19_20 = object : Migration(19, 20) {
override fun migrate(db: SupportSQLiteDatabase) {
db.execSQL(
"CREATE INDEX IF NOT EXISTS `index_messages_folder_inInbox_timestampMillis` " +
"ON `messages` (`folder`, `inInbox`, `timestampMillis`)",
)
}
}
@@ -11,7 +11,19 @@ import androidx.room.PrimaryKey
// The (accountId, folder, uid) index serves the folder-scoped UID probes the backfill/reconcile
// hot paths run on every page/sync: MIN(uid) (lowestSyncedUid) and the uid >= window bound
// (deleteSyncedInWindowNotIn / syncedIdsBeyondCountInFolder).
indices = [Index("accountId"), Index("timestampMillis"), Index("accountId", "folder", "uid")],
//
// The (folder, inInbox, timestampMillis) index serves the unified-inbox summary scan (issue #187):
// MessageDao.pagingUnifiedFolderSummaries filters `folder = ? AND inInbox = 1 ORDER BY
// timestampMillis DESC` with no folder-leading index, so it SCANned the whole table via
// index_messages_timestampMillis and filtered per row. This index makes the two equalities an
// index seek and supplies the timestampMillis ordering, turning the SCAN into a bounded SEARCH
// with no temp B-tree sort (verified via EXPLAIN QUERY PLAN).
indices = [
Index("accountId"),
Index("timestampMillis"),
Index("accountId", "folder", "uid"),
Index("folder", "inInbox", "timestampMillis"),
],
)
data class MessageEntity(
@PrimaryKey val id: String,
@@ -41,6 +41,7 @@ import org.libremail.data.settings.AccountSettingsRepository
import org.libremail.data.settings.SignatureRepository
import org.libremail.data.sync.MailConnectionFactory
import org.libremail.data.sync.SendScheduler
import org.libremail.data.sync.logSafeFolderLabel
import org.libremail.domain.model.Attachment
import org.libremail.domain.model.Draft
import org.libremail.domain.model.Folder
@@ -56,6 +57,8 @@ import org.libremail.domain.model.UnreadCount
import org.libremail.domain.model.sanitizeAttachmentName
import org.libremail.domain.repository.MailRepository
import org.libremail.mail.ImapClient
import org.libremail.reporting.AppLog
import org.libremail.reporting.accountLogRef
import java.io.File
import java.util.UUID
import javax.inject.Inject
@@ -153,11 +156,15 @@ class MailRepositoryImpl @Inject constructor(
override suspend fun openMessage(id: String): Result<Message> = withContext(Dispatchers.IO) {
runCatching {
// Time the whole open so a debug report shows what the reader's spinner is waiting on — a
// cached open is a local read; a first open blocks on the IMAP body fetch below (issue #358).
val startNanos = System.nanoTime()
// Route on the body-less projection: a cached, already-read message needs no account, no
// credentials, and no network, so it skips the Keystore decrypt + DataStore read that
// resolving connection params costs (issue #186). Only the fetch / SEEN-push branches below
// pull the account and resolve params, and each does so lazily right where it is needed.
val routing = messageDao.getRouting(id) ?: error("Message not found")
val fetchedBody = !routing.bodyFetched
if (!routing.bodyFetched || !routing.isRead) {
val account = accountDao.getById(routing.accountId)?.toDomain()
if (account != null && !routing.bodyFetched) {
@@ -177,7 +184,14 @@ class MailRepositoryImpl @Inject constructor(
}
}
// The single full-body read, reserved for the value the reader actually renders (issue #186).
messageDao.getById(id)?.toDomain() ?: error("Message not found")
val message = messageDao.getById(id)?.toDomain() ?: error("Message not found")
// PII-free: hashed account ref, system-folder label only, plus the branch taken and elapsed ms.
AppLog.i(
READER_TAG,
"openMessage ${accountLogRef(routing.accountId)} folder=${logSafeFolderLabel(routing.folder)} " +
"fetchedBody=$fetchedBody took=${(System.nanoTime() - startNanos) / NANOS_PER_MS}ms",
)
message
}
}
@@ -535,6 +549,10 @@ class MailRepositoryImpl @Inject constructor(
private const val SEARCH_LIMIT = 50
/** Perf-breadcrumb tag and ns→ms divisor for the reader-open timing (issue #358). */
private const val READER_TAG = "MailReader"
private const val NANOS_PER_MS = 1_000_000L
/** Rows per page for the unified inbox (issue #124) — a page is a few screenfuls of message rows. */
private const val MAILBOX_PAGE_SIZE = 40
@@ -1,9 +1,12 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.data.security
import android.os.Build
import android.security.keystore.KeyGenParameterSpec
import android.security.keystore.KeyProperties
import android.security.keystore.StrongBoxUnavailableException
import android.util.Base64
import org.libremail.reporting.AppLog
import java.security.GeneralSecurityException
import java.security.KeyStore
import javax.crypto.AEADBadTagException
@@ -115,23 +118,52 @@ abstract class AesGcmKeystoreCipher(private val alias: String, private val gener
// Synchronized so two concurrent first-run encrypts can't both generate a key under the same
// alias — the second would overwrite the first, leaving the first secret undecryptable.
protected open fun getOrCreateKey(): SecretKey = synchronized(keyLock) {
existingKey()?.let { return it }
existingKey() ?: generateKeyWithStrongBoxFallback()
}
/**
* Mint the key, preferring the hardware **StrongBox** secure element (a dedicated tamper-resistant
* chip) so the non-exportable key is bound to the strongest keystore available. Devices without
* StrongBox report [StrongBoxUnavailableException] at generation time; we then regenerate a
* TEE-backed key so key creation still succeeds on every device. Both the master ([KeystoreCrypto])
* and auth-bound ([DatabaseKeyCipher]) keys inherit this through the shared base.
*/
private fun generateKeyWithStrongBoxFallback(): SecretKey = try {
generateKey(strongBox = true)
} catch (e: StrongBoxUnavailableException) {
// Expected on devices with no StrongBox — not an error. PII-free (a device-capability fact).
AppLog.i(TAG, "StrongBox unavailable for Keystore alias '$alias'; using a TEE-backed key: ${e.message}")
generateKey(strongBox = false)
}
/**
* Test seam over the raw Android Keystore key generation for a given [strongBox] preference. The
* real [KeyGenerator] is device-only, so JVM unit tests override this to exercise the StrongBox
* fallback in [generateKeyWithStrongBoxFallback] without a Keystore.
*/
protected open fun generateKey(strongBox: Boolean): SecretKey {
val generator = KeyGenerator.getInstance(KeyProperties.KEY_ALGORITHM_AES, ANDROID_KEYSTORE)
generator.init(keySpec())
generator.generateKey()
generator.init(keySpec(strongBox))
return generator.generateKey()
}
/** The alias-bound [KeyGenParameterSpec] for this key; subclasses extend [keySpecBuilder]. */
protected abstract fun keySpec(): KeyGenParameterSpec
protected abstract fun keySpec(strongBox: Boolean): KeyGenParameterSpec
/** The common AES-256-GCM builder (encrypt + decrypt, GCM, no padding, 256-bit) to extend. */
protected fun keySpecBuilder(): KeyGenParameterSpec.Builder = KeyGenParameterSpec.Builder(
protected fun keySpecBuilder(strongBox: Boolean): KeyGenParameterSpec.Builder = KeyGenParameterSpec.Builder(
alias,
KeyProperties.PURPOSE_ENCRYPT or KeyProperties.PURPOSE_DECRYPT,
)
.setBlockModes(KeyProperties.BLOCK_MODE_GCM)
.setEncryptionPaddings(KeyProperties.ENCRYPTION_PADDING_NONE)
.setKeySize(AES_KEY_SIZE_BITS)
.apply {
// Bind the key to the StrongBox secure element when requested and supported (API 28+; minSdk
// is 29, so the guard is defensive). If the device has no StrongBox, generateKey() catches
// StrongBoxUnavailableException and retries with strongBox = false for a TEE-backed key.
if (strongBox && Build.VERSION.SDK_INT >= Build.VERSION_CODES.P) setIsStrongBoxBacked(true)
}
private companion object {
const val ANDROID_KEYSTORE = "AndroidKeyStore"
@@ -139,5 +171,6 @@ abstract class AesGcmKeystoreCipher(private val alias: String, private val gener
const val IV_LENGTH = 12
const val TAG_BITS = 128
const val AES_KEY_SIZE_BITS = 256
const val TAG = "AesGcmKeystoreCipher"
}
}
@@ -5,7 +5,7 @@ import android.os.Build
import android.security.keystore.KeyGenParameterSpec
import android.security.keystore.KeyPermanentlyInvalidatedException
import android.security.keystore.UserNotAuthenticatedException
import android.util.Log
import org.libremail.reporting.AppLog
import javax.inject.Inject
import javax.inject.Singleton
@@ -43,7 +43,7 @@ class DatabaseKeyCipher @Inject constructor() :
override fun encrypt(plaintext: String): String = try {
super.encrypt(plaintext)
} catch (e: KeyPermanentlyInvalidatedException) {
Log.d(TAG, "replacing invalidated auth-bound key before sealing", e)
AppLog.d(TAG, "replacing invalidated auth-bound key before sealing", e)
deleteKey()
super.encrypt(plaintext)
}
@@ -60,16 +60,16 @@ class DatabaseKeyCipher @Inject constructor() :
initEncryptCipher(key)
false
} catch (e: KeyPermanentlyInvalidatedException) {
Log.d(TAG, "auth-bound database key invalidated", e)
AppLog.d(TAG, "auth-bound database key invalidated", e)
true
} catch (e: UserNotAuthenticatedException) {
// Valid key, just outside its time-bound auth window — not invalidated.
Log.d(TAG, "auth-bound key outside its auth window; not invalidated", e)
AppLog.d(TAG, "auth-bound key outside its auth window; not invalidated", e)
false
} catch (e: Exception) {
// Never let a validity probe crash the foreground pass; a real decrypt later surfaces any
// genuine problem. Treat an unknown probe failure as "not invalidated" (don't wipe).
Log.d(TAG, "auth-bound key validity probe failed; treating as valid", e)
AppLog.d(TAG, "auth-bound key validity probe failed; treating as valid", e)
false
}
}
@@ -82,8 +82,8 @@ class DatabaseKeyCipher @Inject constructor() :
/** A missing auth-bound key means it was invalidated; surface that instead of regenerating. */
override fun onMissingDecryptionKey(): Nothing = error("auth-bound database key is missing")
override fun keySpec(): KeyGenParameterSpec {
val builder = keySpecBuilder()
override fun keySpec(strongBox: Boolean): KeyGenParameterSpec {
val builder = keySpecBuilder(strongBox)
.setUserAuthenticationRequired(true)
.setInvalidatedByBiometricEnrollment(true)
if (Build.VERSION.SDK_INT >= Build.VERSION_CODES.R) {
@@ -18,7 +18,7 @@ import javax.inject.Singleton
@Singleton
class KeystoreCrypto @Inject constructor() : AesGcmKeystoreCipher(alias = KEY_ALIAS, generateKeyOnDecrypt = true) {
override fun keySpec(): KeyGenParameterSpec = keySpecBuilder().build()
override fun keySpec(strongBox: Boolean): KeyGenParameterSpec = keySpecBuilder(strongBox).build()
private companion object {
const val KEY_ALIAS = "libremail.master.key"
@@ -0,0 +1,60 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.data.security
import kotlinx.coroutines.flow.distinctUntilChanged
import kotlinx.coroutines.flow.map
import org.libremail.data.settings.SettingsRepository
import org.libremail.reporting.AppLog
import org.libremail.reporting.ReportEncryption
import javax.inject.Inject
import javax.inject.Singleton
/**
* On-device [ReportEncryption]: seals a persisted report's JSON with the non-auth Keystore master key
* ([KeystoreCrypto]) so at-rest report storage honours the opt-in `encryptCache` setting (issue #369).
* Reuses the vetted AES-256-GCM crypto rather than rolling new; encryption is `Base64(iv || ciphertext)`.
*
* The **master** key (not the auth-bound cache key) is deliberate: it is usable without a user-presence
* prompt, so a crash that occurs while the app is locked can still seal and persist its report — the
* ticket requires crash reports to survive, encrypted, even then.
*
* [enabled] is answered from an in-memory mirror of the `encryptCache` setting, never a live DataStore
* read: a crash-time [org.libremail.reporting.ReportStore.save] runs synchronously on the crashing
* thread and must not touch DataStore (#296). [observeEncryptCacheSetting], launched once at startup,
* keeps that mirror current so a mid-session toggle takes effect on the next report write. The mirror
* defaults to `false` (plaintext) until the first settings value lands — the same brief unwarmed
* startup window [org.libremail.reporting.DiagnosticsCollector] accepts for a crash report's settings.
*/
@Singleton
class KeystoreReportEncryption @Inject constructor(
private val crypto: KeystoreCrypto,
private val settingsRepository: SettingsRepository,
) : ReportEncryption {
@Volatile
private var encryptionEnabled: Boolean = false
override fun enabled(): Boolean = encryptionEnabled
override fun encrypt(plaintext: String): String = crypto.encrypt(plaintext)
override fun decrypt(encoded: String): String = crypto.decrypt(encoded)
/**
* Mirrors the `encryptCache` setting into [encryptionEnabled] for the process lifetime. Collects
* forever, so launch it once from application startup. PII-free — only the on/off state is logged.
*/
suspend fun observeEncryptCacheSetting() {
settingsRepository.settings
.map { it.encryptCache }
.distinctUntilChanged()
.collect { enabled ->
encryptionEnabled = enabled
AppLog.i(TAG, "Report at-rest encryption is now ${if (enabled) "ON" else "OFF"}")
}
}
private companion object {
const val TAG = "ReportEncryption"
}
}
@@ -9,7 +9,9 @@ import dagger.Lazy
import dagger.assisted.Assisted
import dagger.assisted.AssistedInject
import kotlinx.coroutines.CancellationException
import org.libremail.BuildConfig
import org.libremail.data.security.EncryptedCacheGuard
import org.libremail.reporting.AppLog
/**
* Runs one bounded slice of the full-history backfill (issue #12). Cancellable (WorkManager stops it
@@ -29,9 +31,22 @@ class BackfillWorker @AssistedInject constructor(
) : CoroutineWorker(appContext, workerParams) {
override suspend fun doWork(): Result {
// Debug-only fetch gate (issue #393): a test harness can pause backfill via an adb broadcast so a
// genuinely uncached message-open can be measured (proactive backfill would otherwise warm the
// cache first). Skip-and-reschedule exactly like the cache-lock deferral below; WorkManager
// retries and picks up from the persisted per-folder boundary once the gate resumes. The whole
// branch is compiled out of release: BuildConfig.DEBUG is a compile-time `false` there, so R8
// strips it (and DebugFetchGate with it).
if (BuildConfig.DEBUG && DebugFetchGate.isPaused(FetchScope.BACKFILL)) {
AppLog.i(TAG, "backfill deferred: fetch-gate paused")
return Result.retry()
}
// Can't open the encrypted DB without the user present — retry later rather than parking a
// WorkManager thread (which also wedges the shared serial executor) on an unsatisfiable await.
if (cacheGuard.isCacheLocked()) return Result.retry()
if (cacheGuard.isCacheLocked()) {
AppLog.i(TAG, "backfill deferred: cache locked")
return Result.retry()
}
return runCatching {
// Chain bounded slices back-to-back while history remains, so a large mailbox isn't limited
// to one slice per periodic run. runBackfill() returns true while any folder still has pages
@@ -39,8 +54,19 @@ class BackfillWorker @AssistedInject constructor(
val mailBackfiller = backfiller.get()
while (mailBackfiller.runBackfill() && !isStopped) { /* page the next slice */ }
}.fold(
onSuccess = { Result.success() },
onFailure = { error -> if (error is CancellationException) throw error else Result.retry() },
onSuccess = {
AppLog.i(TAG, "backfill worker: success")
Result.success()
},
onFailure = { error ->
if (error is CancellationException) throw error
AppLog.w(TAG, "backfill worker: retry", error)
Result.retry()
},
)
}
private companion object {
const val TAG = "BackfillWorker"
}
}
@@ -0,0 +1,97 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.data.sync
/**
* The proactive-fetch activities a debug harness can pause via [DebugFetchGate] (issue #393). Only the
* two *proactive* activities are gateable: full-history [BACKFILL] paging ([BackfillWorker]) and the
* post-sync body [PREFETCH] ([MailSyncer]/[MailBackfiller] `prefetchIfEnabled`). Header sync and the
* on-demand message open are deliberately absent — they are **never** gated, so a paused gate can defer
* background caching without ever blocking new mail arriving or a user-triggered (uncached) open. The
* `all` wire alias ([ALL_ALIAS]) expands to every entry here.
*/
enum class FetchScope(val wireName: String) {
BACKFILL("backfill"),
PREFETCH("prefetch"),
;
companion object {
/** The `all` scope alias accepted on the adb wire — expands to every [FetchScope]. */
const val ALL_ALIAS = "all"
/**
* Parses the comma-separated `scope` extra of the debug broadcast (e.g. `"backfill,prefetch"`
* or `"all"`) into the set of scopes it names. Case- and whitespace-insensitive; the [ALL_ALIAS]
* expands to every scope; unrecognised or blank tokens are ignored; a null/blank input yields
* the empty set. Declaration order is preserved so the read-back string is stable.
*/
fun parse(raw: String?): Set<FetchScope> {
if (raw.isNullOrBlank()) return emptySet()
val tokens = raw.split(',').map { it.trim().lowercase() }.filter { it.isNotEmpty() }
if (tokens.contains(ALL_ALIAS)) return entries.toSet()
return entries.filterTo(LinkedHashSet()) { it.wireName in tokens }
}
}
}
/**
* In-memory, thread-safe holder of the currently-paused proactive-fetch [FetchScope]s — a **debug-only**
* test hook (issue #393) that lets an adb-driven perf harness pause background body caching so a genuine
* uncached message-open can be measured (the harness adds an account, lets headers sync, then opens a
* message that must hit the network rather than a warmed cache).
*
* Lives in `src/main` so the workers/syncer/backfiller can reference it, but **every read is wrapped in
* `if (BuildConfig.DEBUG && ...)`**. `BuildConfig.DEBUG` is a compile-time `false` in release, so R8
* dead-code-eliminates each such branch, leaving this object unreferenced and stripping it (and
* [FetchScope]) from the release APK entirely — verified by issue #393's release-exclusion check. The
* writer, `FetchGateReceiver`, lives wholly in `src/debug` and is never packaged into release either;
* this is the same source-set guarantee `ColdOpenCacheProbe` (#221) relies on.
*
* Defaults to **nothing paused**, so the gate is inert until a debug broadcast pauses a scope. Reads are
* lock-free (a `@Volatile` snapshot of an immutable set, cheap enough for the fetch hot path); the rare
* writes swap the reference under a lock.
*/
object DebugFetchGate {
private val writeLock = Any()
@Volatile
private var paused: Set<FetchScope> = emptySet()
/** Whether [scope]'s proactive fetch is currently paused. */
fun isPaused(scope: FetchScope): Boolean = scope in paused
/** The scopes currently paused, in [FetchScope] declaration order. */
fun pausedScopes(): Set<FetchScope> {
val snapshot = paused
return FetchScope.entries.filterTo(LinkedHashSet()) { it in snapshot }
}
/** Pauses [scopes] (union with whatever is already paused). No-op for an empty set. */
fun pause(scopes: Set<FetchScope>) {
if (scopes.isEmpty()) return
synchronized(writeLock) { paused = paused + scopes }
}
/** Resumes [scopes] (removes them from the paused set). No-op for an empty set. */
fun resume(scopes: Set<FetchScope>) {
if (scopes.isEmpty()) return
synchronized(writeLock) { paused = paused - scopes }
}
/** Clears every pause, restoring the default not-paused state. Used to isolate tests. */
fun reset() {
synchronized(writeLock) { paused = emptySet() }
}
/**
* The synchronous read-back string the debug receiver returns as ordered-broadcast result data —
* e.g. `"paused=[backfill,prefetch]"` (declaration order) or `"paused=[]"` when nothing is paused.
*/
fun pausedResult(): String {
val snapshot = paused
return FetchScope.entries.filter { it in snapshot }.joinToString(
separator = ",",
prefix = "paused=[",
postfix = "]",
) { it.wireName }
}
}
@@ -9,6 +9,7 @@ import kotlinx.coroutines.delay
import kotlinx.coroutines.ensureActive
import kotlinx.coroutines.sync.withLock
import kotlinx.coroutines.withContext
import org.libremail.BuildConfig
import org.libremail.data.local.dao.AccountDao
import org.libremail.data.local.dao.BackfillProgressDao
import org.libremail.data.local.dao.MessageDao
@@ -25,6 +26,8 @@ import org.libremail.domain.model.ImapConnectionParams
import org.libremail.domain.repository.MailRepository
import org.libremail.mail.ImapClient
import org.libremail.power.BatteryStatusProvider
import org.libremail.reporting.AppLog
import org.libremail.reporting.accountLogRef
import javax.inject.Inject
import javax.inject.Singleton
@@ -65,13 +68,19 @@ class MailBackfiller @Inject constructor(
* count as more work — it is retried on a future scheduled run instead of spun on back-to-back.
*/
suspend fun runBackfill(maxBatches: Int = DEFAULT_MAX_BATCHES): Boolean = maintenanceGate.mutex.withLock {
AppLog.i(TAG, "backfill slice: maxBatches=$maxBatches")
var remaining = maxBatches
var moreWork = false
for (account in accountDao.getAll().map { it.toDomain() }) {
accounts@ for (account in accountDao.getAll().map { it.toDomain() }) {
val params = runCatching { connectionFactory.imapParamsFor(account) }.getOrNull() ?: continue
val policy = accountSettingsRepository.effectiveRetention(settingsRepository, account.id)
for (folder in messageDao.syncedFolders(account.id)) {
if (remaining <= 0) return@withLock true
if (remaining <= 0) {
// The budget ran out before every folder was visited, so there is very likely more
// work left even though nothing here reported it directly.
moreWork = true
break@accounts
}
// Per-folder failures (e.g. a transient server error) must not abort the whole slice.
val result = runCatching { backfillFolder(account, params, folder, policy, remaining) }
.getOrElse { FolderResult(batches = 0, moreWork = true) }
@@ -79,6 +88,7 @@ class MailBackfiller @Inject constructor(
if (result.moreWork) moreWork = true
}
}
AppLog.i(TAG, "backfill slice done: moreWork=$moreWork")
moreWork
}
@@ -153,6 +163,8 @@ class MailBackfiller @Inject constructor(
delay(BACKFILL_BATCH_DELAY_MS)
}
if (complete) markComplete(account.id, folder, beforeUid)
val folderLabel = logSafeFolderLabel(folder)
AppLog.d(TAG, "backfill ${accountLogRef(account.id)} folder=$folderLabel pages=$batches complete=$complete")
return FolderResult(batches, moreWork = !complete && !stalled)
}
@@ -213,6 +225,14 @@ class MailBackfiller @Inject constructor(
* fetched is filled in lazily when the message is opened.
*/
private suspend fun prefetchIfEnabled(ids: List<String>) {
// Debug-only fetch gate (issue #393): pause proactive body prefetch so a later open is a genuine
// uncached fetch. Header paging above is untouched (its own gate is the BackfillWorker entry), so
// history still lands; a skipped body is filled in lazily on open. Compiled out of release
// (BuildConfig.DEBUG is a compile-time false, so R8 drops the branch).
if (BuildConfig.DEBUG && DebugFetchGate.isPaused(FetchScope.PREFETCH)) {
AppLog.i(TAG, "prefetch skipped: fetch-gate paused")
return
}
val shouldPrefetch = SyncResourcePolicy.shouldPrefetchContent(
policy = settingsRepository.fetchPolicy(),
unmetered = { context.isActiveNetworkUnmetered() },
@@ -226,6 +246,8 @@ class MailBackfiller @Inject constructor(
}
private companion object {
const val TAG = "MailBackfiller"
/** Headers fetched per server page. */
const val BACKFILL_BATCH_SIZE = 50
@@ -13,6 +13,7 @@ import org.libremail.data.settings.AccountSettingsRepository
import org.libremail.data.settings.RetentionPolicy
import org.libremail.data.settings.SettingsRepository
import org.libremail.data.settings.effectiveRetention
import org.libremail.reporting.AppLog
import javax.inject.Inject
import javax.inject.Singleton
@@ -49,6 +50,7 @@ class MailPruner @Inject constructor(
if (policy.isUnlimited) continue
removed += pruneAccount(account.id, policy, nowMillis)
}
AppLog.i(TAG, "prune done: removed=$removed")
removed
}
@@ -82,6 +84,8 @@ class MailPruner @Inject constructor(
}
private companion object {
const val TAG = "MailPruner"
/** Ids per DELETE, kept under SQLite's 999-host-parameter limit on older Android. */
const val DELETE_CHUNK = 500
}
@@ -9,6 +9,7 @@ import kotlinx.coroutines.ensureActive
import kotlinx.coroutines.sync.Mutex
import kotlinx.coroutines.sync.withLock
import kotlinx.coroutines.withContext
import org.libremail.BuildConfig
import org.libremail.data.local.dao.AccountDao
import org.libremail.data.local.dao.MessageDao
import org.libremail.data.local.toDomain
@@ -21,6 +22,8 @@ import org.libremail.domain.repository.MailRepository
import org.libremail.mail.ImapClient
import org.libremail.notifications.MailNotifier
import org.libremail.power.BatteryStatusProvider
import org.libremail.reporting.AppLog
import org.libremail.reporting.accountLogRef
import javax.inject.Inject
import javax.inject.Singleton
@@ -47,6 +50,7 @@ class MailSyncer @Inject constructor(
/** Syncs every account's inbox. Succeeds if at least one account synced (or there are none). */
override suspend fun syncAll(): Result<Int> {
val accounts = accountDao.getAll().map { it.toDomain() }
AppLog.i(TAG, "sync all: ${accounts.size} accounts")
if (accounts.isEmpty()) return Result.success(0)
val result = syncMutex.withLock {
@@ -64,6 +68,8 @@ class MailSyncer @Inject constructor(
}
if (anySuccess || firstError == null) Result.success(total) else Result.failure(firstError)
}
result.onSuccess { total -> AppLog.i(TAG, "sync all done: fetched=$total") }
.onFailure { error -> AppLog.w(TAG, "sync all failed", error) }
if (result.isSuccess) accounts.forEach { prefetchIfEnabled(it, INBOX) }
return result
}
@@ -147,6 +153,8 @@ class MailSyncer @Inject constructor(
notifier.notifyNewMail(account, newMessages.sortedByDescending { it.timestampMillis })
}
}
val folderLabel = logSafeFolderLabel(folder)
AppLog.d(TAG, "sync ${accountLogRef(account.id)} folder=$folderLabel fetched=${fetched.size}")
fetched.size
}
@@ -159,6 +167,14 @@ class MailSyncer @Inject constructor(
* and is cancellable between messages so an IDLE renewal stops it promptly.
*/
private suspend fun prefetchIfEnabled(account: Account, folder: String) {
// Debug-only fetch gate (issue #393): a test harness pauses proactive body prefetch so a later
// open does a genuine uncached fetch. Header sync above already ran, so mail still arrives; the
// skipped prefetch is filled in lazily on open, exactly as the low-battery pause behaves. Compiled
// out of release (BuildConfig.DEBUG is a compile-time false, so R8 drops the branch).
if (BuildConfig.DEBUG && DebugFetchGate.isPaused(FetchScope.PREFETCH)) {
AppLog.i(TAG, "prefetch skipped: fetch-gate paused")
return
}
val shouldPrefetch = SyncResourcePolicy.shouldPrefetchContent(
policy = settingsRepository.fetchPolicy(),
unmetered = { context.isActiveNetworkUnmetered() },
@@ -172,6 +188,7 @@ class MailSyncer @Inject constructor(
}
private companion object {
const val TAG = "MailSyncer"
const val INBOX = "INBOX"
/**
@@ -9,6 +9,7 @@ import dagger.Lazy
import dagger.assisted.Assisted
import dagger.assisted.AssistedInject
import org.libremail.data.security.EncryptedCacheGuard
import org.libremail.reporting.AppLog
/**
* Enforces device-only retention (issue #13) by running [MailPruner]. Purely local — it never
@@ -28,10 +29,23 @@ class PruneWorker @AssistedInject constructor(
override suspend fun doWork(): Result {
// Can't open the encrypted DB without the user present — retry later rather than parking a
// WorkManager thread (which also wedges the shared serial executor) on an unsatisfiable await.
if (cacheGuard.isCacheLocked()) return Result.retry()
if (cacheGuard.isCacheLocked()) {
AppLog.i(TAG, "prune deferred: cache locked")
return Result.retry()
}
return runCatching { pruner.get().prune() }.fold(
onSuccess = { Result.success() },
onFailure = { Result.retry() },
onSuccess = {
AppLog.i(TAG, "prune worker: success")
Result.success()
},
onFailure = { error ->
AppLog.w(TAG, "prune worker: retry", error)
Result.retry()
},
)
}
private companion object {
const val TAG = "PruneWorker"
}
}
@@ -2,7 +2,6 @@
package org.libremail.data.sync
import android.content.Context
import android.util.Log
import androidx.hilt.work.HiltWorker
import androidx.work.CoroutineWorker
import androidx.work.WorkerParameters
@@ -24,6 +23,8 @@ import org.libremail.mail.GraphSendException
import org.libremail.mail.GraphSender
import org.libremail.mail.SendableAttachment
import org.libremail.mail.SmtpSender
import org.libremail.reporting.AppLog
import org.libremail.reporting.accountLogRef
import java.io.File
import kotlin.coroutines.cancellation.CancellationException
@@ -57,66 +58,89 @@ class SendWorker @AssistedInject constructor(
val attachmentUriGrants = this.attachmentUriGrants.get()
val pending = outboxDao.getAll()
if (pending.isEmpty()) return Result.success()
AppLog.i(TAG, "outbox drain: ${pending.size} queued")
var anyFailed = false
for (entity in pending) {
val attachmentDir = File(applicationContext.cacheDir, "outbox/${entity.id}")
val account = accountDao.getById(entity.accountId)?.toDomain()
if (account == null) {
outboxDao.delete(entity.id) // account removed — drop the queued message
attachmentDir.deleteRecursively()
attachmentUriGrants.releaseUnreferenced(entity.attachmentUris())
continue
}
runCatching {
val message = OutgoingMessage(
accountId = entity.accountId,
to = entity.toAddresses,
cc = entity.ccAddresses,
bcc = entity.bccAddresses,
subject = entity.subject,
body = entity.body,
bodyHtml = entity.bodyHtml,
)
val attachments = stagedAttachments(attachmentDir, entity.attachments.toOutgoingAttachments())
if (account.authType == AuthType.OAUTH_OUTLOOK) {
sendOutlook(connectionFactory, account, message, attachments)
} else {
smtpSender.send(
connectionFactory.smtpParamsFor(account),
from = account.email,
message = message,
attachments = attachments,
)
}
}.fold(
onSuccess = {
outboxDao.delete(entity.id)
attachmentDir.deleteRecursively()
// The picked bytes were staged at enqueue; with the row sent, drop the persistable
// grant unless a live draft/outbox row still references the same URI (security review).
attachmentUriGrants.releaseUnreferenced(entity.attachmentUris())
},
onFailure = { e ->
if (e is GraphSendException && e.mayHaveSent) {
// Graph may already have delivered this; auto-retrying (or any other send)
// would duplicate it, so leave it queued with a clear status and let the
// user decide. Not counted as a failure, so WorkManager won't auto-retry.
outboxDao.setError(
entity.id,
"Send status unknown — check your Sent folder, then retry or cancel",
)
} else {
outboxDao.setError(entity.id, e.message)
anyFailed = true
}
},
)
val failed = sendQueued(entity, outboxDao, accountDao, connectionFactory, attachmentUriGrants)
if (failed) anyFailed = true
}
// Retry (with WorkManager backoff) so failed sends are reattempted when conditions improve.
return if (anyFailed) Result.retry() else Result.success()
}
/**
* Sends one queued [entity] — or drops it if its account was removed — updating the outbox row
* and releasing its staged attachment grant. Returns true if the send genuinely failed (should
* count toward a WorkManager retry); the ambiguous "may have sent" Graph case returns false, since
* it is deliberately left queued rather than retried (see [sendOutlook]).
*/
private suspend fun sendQueued(
entity: OutboxEntity,
outboxDao: OutboxDao,
accountDao: AccountDao,
connectionFactory: MailConnectionFactory,
attachmentUriGrants: AttachmentUriGrants,
): Boolean {
val attachmentDir = File(applicationContext.cacheDir, "outbox/${entity.id}")
val account = accountDao.getById(entity.accountId)?.toDomain()
if (account == null) {
outboxDao.delete(entity.id) // account removed — drop the queued message
attachmentDir.deleteRecursively()
attachmentUriGrants.releaseUnreferenced(entity.attachmentUris())
return false
}
var failed = false
runCatching {
val message = OutgoingMessage(
accountId = entity.accountId,
to = entity.toAddresses,
cc = entity.ccAddresses,
bcc = entity.bccAddresses,
subject = entity.subject,
body = entity.body,
bodyHtml = entity.bodyHtml,
)
val attachments = stagedAttachments(attachmentDir, entity.attachments.toOutgoingAttachments())
if (account.authType == AuthType.OAUTH_OUTLOOK) {
sendOutlook(connectionFactory, account, message, attachments)
} else {
smtpSender.send(
connectionFactory.smtpParamsFor(account),
from = account.email,
message = message,
attachments = attachments,
)
}
}.fold(
onSuccess = {
outboxDao.delete(entity.id)
attachmentDir.deleteRecursively()
// The picked bytes were staged at enqueue; with the row sent, drop the persistable
// grant unless a live draft/outbox row still references the same URI (security review).
attachmentUriGrants.releaseUnreferenced(entity.attachmentUris())
val via = if (account.authType == AuthType.OAUTH_OUTLOOK) "Graph" else "SMTP"
AppLog.i(TAG, "sent ${accountLogRef(account.id)} via $via")
},
onFailure = { e ->
if (e is GraphSendException && e.mayHaveSent) {
// Graph may already have delivered this; auto-retrying (or any other send) would
// duplicate it, so leave it queued with a clear status and let the user decide.
// Not counted as a failure, so WorkManager won't auto-retry.
outboxDao.setError(
entity.id,
"Send status unknown — check your Sent folder, then retry or cancel",
)
} else {
outboxDao.setError(entity.id, e.message)
failed = true
AppLog.w(TAG, "send failed for ${accountLogRef(account.id)}; will retry")
}
},
)
return failed
}
/**
* Outlook prefers Microsoft Graph. Fall back to SMTP only when Graph definitely did NOT send
* (a rejection, a pre-send/transport error, or a token failure); never fall back when the Graph
@@ -143,7 +167,7 @@ class SendWorker @AssistedInject constructor(
throw e
} catch (e: Exception) {
// Graph was never reached (e.g. token refresh failed) — SMTP cannot duplicate it.
Log.w(TAG, "Graph send failed for ${account.email}; falling back to SMTP", e)
AppLog.w(TAG, "Graph send failed for ${accountLogRef(account.id)}; falling back to SMTP", e)
smtpSender.send(
connectionFactory.smtpParamsFor(account),
from = account.email,
@@ -0,0 +1,45 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.data.sync
/**
* The folder name safe to write to an [org.libremail.reporting.AppLog] breadcrumb (and so, in turn, a
* submitted [org.libremail.reporting.DebugReport]): the leaf name itself for a known **system** folder
* (INBOX, Sent, Drafts, Trash, Spam/Junk, Archive — including the alternate names real IMAP servers use
* for them), or a fixed placeholder for anything else. A user-created folder or label (e.g. a client
* name or project) can be PII-ish, so only this fixed, closed set of well-known names is ever logged
* verbatim; every other folder logs as the placeholder, regardless of nesting or the server's hierarchy
* delimiter. Matching is name-only — no server SPECIAL-USE attributes are available down here at the
* sync layer — so it is necessarily best-effort in the same way
* [org.libremail.domain.model.FolderRole.roleOf]'s display-name fallback is. That is the safe direction:
* a false negative just logs the placeholder, never a leaked name.
*/
internal fun logSafeFolderLabel(folder: String): String {
val leaf = folder.substringAfterLast('/').substringAfterLast('.').trim()
return if (leaf.lowercase() in SYSTEM_FOLDER_NAMES) leaf else FOLDER_PLACEHOLDER
}
private const val FOLDER_PLACEHOLDER = "<folder>"
/** Case-insensitive leaf names recognized as provider-supplied system folders, never user-created. */
private val SYSTEM_FOLDER_NAMES = setOf(
"inbox",
"sent",
"sent mail",
"sent items",
"sent messages",
"drafts",
"draft",
"junk",
"spam",
"junk e-mail",
"junk email",
"bulk mail",
"trash",
"deleted",
"deleted items",
"deleted messages",
"bin",
"archive",
"archives",
"all mail",
)
@@ -9,6 +9,7 @@ import dagger.Lazy
import dagger.assisted.Assisted
import dagger.assisted.AssistedInject
import org.libremail.data.security.EncryptedCacheGuard
import org.libremail.reporting.AppLog
@HiltWorker
class SyncWorker @AssistedInject constructor(
@@ -23,10 +24,23 @@ class SyncWorker @AssistedInject constructor(
override suspend fun doWork(): Result {
// Can't open the encrypted DB without the user present — retry later rather than parking a
// WorkManager thread (which also wedges the shared serial executor) on an unsatisfiable await.
if (cacheGuard.isCacheLocked()) return Result.retry()
if (cacheGuard.isCacheLocked()) {
AppLog.i(TAG, "sync deferred: cache locked")
return Result.retry()
}
return mailSyncer.get().syncAll().fold(
onSuccess = { Result.success() },
onFailure = { Result.retry() },
onSuccess = {
AppLog.i(TAG, "sync worker: success")
Result.success()
},
onFailure = { error ->
AppLog.w(TAG, "sync worker: retry", error)
Result.retry()
},
)
}
private companion object {
const val TAG = "SyncWorker"
}
}
@@ -12,6 +12,7 @@ import dagger.hilt.components.SingletonComponent
import kotlinx.coroutines.runBlocking
import org.libremail.data.local.ACCOUNT_MIGRATION_1_2
import org.libremail.data.local.AccountDatabase
import org.libremail.data.local.CacheEncryptionUnavailableException
import org.libremail.data.local.DatabaseFiles.ACCOUNTS_NAME
import org.libremail.data.local.DatabaseProvisioner
import org.libremail.data.local.DeferredOpenHelperFactory
@@ -19,6 +20,7 @@ import org.libremail.data.local.dao.AccountDao
import org.libremail.data.local.dao.AccountSettingsDao
import org.libremail.data.local.dao.CredentialDao
import org.libremail.data.local.dao.SignatureDao
import org.libremail.reporting.AppLog
import javax.inject.Singleton
/**
@@ -37,6 +39,13 @@ object AccountDatabaseModule {
* the migrate-before-open ordering the old construction-time dependency on `LibreMailDatabase`
* enforced, now moved OFF the injection path (issue #93). This store always opens unkeyed, so it
* ignores the returned cache open-mode and only awaits the shared sequence.
*
* This store is plaintext and never uses SQLCipher, so a cache-encryption native-load failure
* (issue #359, surfaced as [CacheEncryptionUnavailableException]) must NOT brick it: the app still
* needs accounts/credentials to render the encryption error gate and let the user file a PII-free
* problem report. The wipe + migrate steps run BEFORE the encryption gate that can throw, so the
* migrate-before-open ordering still holds when we tolerate that one specific failure here; any
* other failure still propagates.
*/
@Provides
@Singleton
@@ -47,7 +56,11 @@ object AccountDatabaseModule {
.addMigrations(ACCOUNT_MIGRATION_1_2)
.openHelperFactory(
DeferredOpenHelperFactory { configuration ->
runBlocking { provisioner.prepareCache() }
try {
runBlocking { provisioner.prepareCache() }
} catch (e: CacheEncryptionUnavailableException) {
AppLog.w(TAG, "cache encryption unavailable; opening the plaintext account store anyway", e)
}
FrameworkSQLiteOpenHelperFactory().create(configuration)
},
)
@@ -64,4 +77,6 @@ object AccountDatabaseModule {
@Provides
fun provideSignatureDao(database: AccountDatabase): SignatureDao = database.signatureDao()
private const val TAG = "AccountDatabaseModule"
}
@@ -26,6 +26,7 @@ import org.libremail.data.local.MIGRATION_15_16
import org.libremail.data.local.MIGRATION_16_17
import org.libremail.data.local.MIGRATION_17_18
import org.libremail.data.local.MIGRATION_18_19
import org.libremail.data.local.MIGRATION_19_20
import org.libremail.data.local.MIGRATION_1_2
import org.libremail.data.local.MIGRATION_2_3
import org.libremail.data.local.MIGRATION_3_4
@@ -73,6 +74,7 @@ object DatabaseModule {
MIGRATION_16_17,
MIGRATION_17_18,
MIGRATION_18_19,
MIGRATION_19_20,
)
@Provides
@@ -8,6 +8,7 @@ import dagger.hilt.InstallIn
import dagger.hilt.android.qualifiers.ApplicationContext
import dagger.hilt.components.SingletonComponent
import org.libremail.BuildConfig
import org.libremail.data.security.KeystoreReportEncryption
import org.libremail.reporting.DebugReportEndpoint
import org.libremail.reporting.ReportStore
import java.io.File
@@ -17,10 +18,19 @@ import javax.inject.Singleton
@InstallIn(SingletonComponent::class)
object ReportingModule {
/**
* The file-backed report store. [KeystoreReportEncryption] wires in the opt-in at-rest encryption
* (issue #369): reports are sealed on disk when the `encryptCache` setting is on, plaintext when off.
*/
@Provides
@Singleton
fun provideReportStore(@ApplicationContext context: Context): ReportStore =
ReportStore(File(context.filesDir, "debug_reports"))
fun provideReportStore(
@ApplicationContext context: Context,
reportEncryption: KeystoreReportEncryption,
): ReportStore = ReportStore(
directory = File(context.filesDir, "debug_reports"),
encryption = reportEncryption,
)
/**
* The debug-report ingest endpoint (empty by default — see [DebugReportEndpoint]). Provided as an
@@ -1,7 +1,6 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.mail
import android.util.Log
import jakarta.mail.FetchProfile
import jakarta.mail.Flags
import jakarta.mail.Folder
@@ -31,9 +30,12 @@ import kotlinx.coroutines.launch
import kotlinx.coroutines.withContext
import org.eclipse.angus.mail.imap.IMAPFolder
import org.eclipse.angus.mail.imap.IMAPMessage
import org.libremail.BuildConfig
import org.libremail.domain.model.ImapConnectionParams
import org.libremail.domain.model.MailSecurity
import org.libremail.reporting.AppLog
import java.util.Properties
import java.util.concurrent.atomic.AtomicInteger
import javax.inject.Inject
import javax.inject.Singleton
@@ -98,23 +100,29 @@ data class ReplyContext(
/** Thin IMAP client over Jakarta/Angus Mail. Supports password and XOAUTH2 auth. */
@Singleton
class ImapClient(private val reuseConnections: Boolean) {
class ImapClient internal constructor(
private val reuseConnections: Boolean,
private val reuseIdleTimeoutMillis: Long = DEFAULT_REUSE_IDLE_TIMEOUT_MS,
) {
/**
* Production entry point. Connection reuse is a SPIKE flag (issue #125), **OFF by default** so it
* cannot destabilize the connect-per-operation behaviour on `main`: with it off, [withStore] is
* byte-for-byte today's connect + LOGOUT-per-call. Once real-device validation (see
* `docs/perf/issue-125-connection-reuse-spike.md`) confirms the win, wire this to a setting or
* `BuildConfig`; today only the reuse harness flips it on via the primary constructor.
* Production entry point. Connection reuse (issue #357 Part 2, wiring the #125 spike) is **ON by
* default**, driven by [BuildConfig.IMAP_CONNECTION_REUSE]: instead of a cold `CONNECT + TLS +
* LOGIN` per operation, each account keeps one authenticated connection warm (see
* [ImapConnectionCache]), which is the fix for Gmail throttling LibreMail's connect-per-operation
* traffic (`docs/perf/issue-125-*`). The `BuildConfig` field is the safety switch: flipping it to
* `false` (a build-config change, no code edit) restores connect-per-operation if a server
* misbehaves with a kept-alive socket. The internal constructor is the test/harness seam.
*/
@Inject constructor() : this(reuseConnections = false)
@Inject constructor() : this(reuseConnections = BuildConfig.IMAP_CONNECTION_REUSE)
/**
* Per-account keep-alive cache; allocated only when the spike flag is on, so the default build
* carries neither the state nor the reuse code path.
* Per-account keep-alive cache; allocated only when reuse is enabled, so a reuse-disabled build
* carries neither the state nor the reuse code path (and [withStore] stays byte-for-byte the old
* connect + LOGOUT-per-call).
*/
private val connectionCache: ImapConnectionCache? =
if (reuseConnections) ImapConnectionCache(::openConnectedStore) else null
if (reuseConnections) ImapConnectionCache(::openConnectedStore, reuseIdleTimeoutMillis) else null
/** Connects and returns the account's folders with their SPECIAL-USE attributes. Throws on failure. */
suspend fun listFolders(params: ImapConnectionParams): List<FetchedFolder> = withContext(Dispatchers.IO) {
@@ -190,7 +198,7 @@ class ImapClient(private val reuseConnections: Boolean) {
beforeUid: Long,
limit: Int,
): List<FetchedMessage> = withContext(Dispatchers.IO) {
withStore(params) { store ->
withStore(params, op = "backfill-page") { store ->
val mailbox = store.getFolder(folder)
mailbox.open(Folder.READ_ONLY)
try {
@@ -270,15 +278,29 @@ class ImapClient(private val reuseConnections: Boolean) {
/** Fetches a message body by UID from [folder] and marks it \Seen on the server. */
suspend fun fetchBodyMarkingSeen(params: ImapConnectionParams, folder: String, uid: String): MessageContent =
withContext(Dispatchers.IO) {
withStore(params) { store ->
withStore(params, op = "body-fetch") { store ->
val mailbox = store.getFolder(folder)
val selectStart = System.nanoTime()
mailbox.open(Folder.READ_WRITE)
val selectMs = (System.nanoTime() - selectStart) / NANOS_PER_MS
try {
val message = (mailbox as UIDFolder).getMessageByUID(uid.toLong())
?: error("Message $uid not found")
val fetchStart = System.nanoTime()
val content = (extractBody(message) ?: MessageContent("", isHtml = false))
.copy(attachments = collectAttachments(message))
val fetchMs = (System.nanoTime() - fetchStart) / NANOS_PER_MS
// RFC822.SIZE (server-reported wire size) is the download-budget signal; body char
// count and attachment count round it out. All are numbers — never message content.
val rfc822Bytes = runCatching { message.size }.getOrDefault(-1)
val flagStart = System.nanoTime()
message.setFlag(Flags.Flag.SEEN, true)
val flagMs = (System.nanoTime() - flagStart) / NANOS_PER_MS
AppLog.d(
PERF_TAG,
"body-fetch select=${selectMs}ms body=${fetchMs}ms flag=${flagMs}ms " +
"rfc822=${rfc822Bytes}B chars=${content.body.length} att=${content.attachments.size}",
)
content
} finally {
runCatching { mailbox.close(false) }
@@ -292,7 +314,7 @@ class ImapClient(private val reuseConnections: Boolean) {
*/
suspend fun fetchBodyPeek(params: ImapConnectionParams, folder: String, uid: String): MessageContent =
withContext(Dispatchers.IO) {
withStore(params) { store ->
withStore(params, op = "prefetch-body") { store ->
val mailbox = store.getFolder(folder)
mailbox.open(Folder.READ_ONLY)
try {
@@ -314,7 +336,7 @@ class ImapClient(private val reuseConnections: Boolean) {
uid: String,
partIndex: Int,
): DownloadedAttachment = withContext(Dispatchers.IO) {
withStore(params) { store ->
withStore(params, op = "attachment") { store ->
val mailbox = store.getFolder(folder)
mailbox.open(Folder.READ_ONLY)
try {
@@ -337,7 +359,7 @@ class ImapClient(private val reuseConnections: Boolean) {
suspend fun setFlag(params: ImapConnectionParams, folder: String, uid: String, flag: Flags.Flag, value: Boolean) =
withContext(Dispatchers.IO) {
withStore(params) { store ->
withStore(params, op = "flag") { store ->
val mailbox = store.getFolder(folder)
mailbox.open(Folder.READ_WRITE)
try {
@@ -474,12 +496,12 @@ class ImapClient(private val reuseConnections: Boolean) {
runCatching { store.close() }
throw e
}
Log.d(TAG, "IDLE connected")
AppLog.d(TAG, "IDLE connected")
val pushes = Channel<Unit>(Channel.CONFLATED)
inbox.addMessageCountListener(object : MessageCountAdapter() {
override fun messagesAdded(event: MessageCountEvent) {
Log.d(TAG, "IDLE push: ${event.messages.size} new message(s)")
AppLog.d(TAG, "IDLE push: ${event.messages.size} new message(s)")
pushes.trySend(Unit)
}
})
@@ -596,22 +618,44 @@ class ImapClient(private val reuseConnections: Boolean) {
)
}
/**
* Live count of connect-per-operation IMAP connections currently open through [withStore] (this
* excludes the single long-lived IDLE connection, which does not go through here). Logged per op so
* a debug report can show how close we run to a provider's simultaneous-connection ceiling — Gmail
* allows 15 — under concurrent backfill + interactive load.
*/
private val liveConnectionCount = AtomicInteger(0)
/**
* Runs [block] against a connected [Store]. With the reuse flag OFF (default) this is the original
* behaviour: a fresh, authenticated connection per call, torn down in `finally`. With it ON, the
* call borrows a kept-alive per-account connection from [connectionCache] (established once, reused
* across folder-opens) instead — see issue #125.
*
* [op] is a short, PII-free label for the caller's intent (e.g. `body-fetch`, `backfill-page`) used
* only in the perf breadcrumb below, so a debug report can attribute latency to connection
* establishment (`connect`) vs the operation's own server work (`work`).
*/
private suspend fun <T> withStore(params: ImapConnectionParams, block: (Store) -> T): T {
private suspend fun <T> withStore(params: ImapConnectionParams, op: String = "imap", block: (Store) -> T): T {
val cache = connectionCache
return if (cache != null) {
cache.withStore(params, block)
cache.withStore(params, op, block)
} else {
// Time CONNECT + TLS + LOGIN separately from the op's own work, and record how many
// connect-per-op sockets are live at once, so a slow op can be attributed and the provider
// connection ceiling observed. Logged in `finally` so a thrown/timed-out op is captured too.
val connectStart = System.nanoTime()
val store = openConnectedStore(params)
val connectMs = (System.nanoTime() - connectStart) / NANOS_PER_MS
val live = liveConnectionCount.incrementAndGet()
val workStart = System.nanoTime()
try {
block(store)
} finally {
val workMs = (System.nanoTime() - workStart) / NANOS_PER_MS
AppLog.d(PERF_TAG, "$op connect=${connectMs}ms work=${workMs}ms live=$live")
runCatching { store.close() }
liveConnectionCount.decrementAndGet()
}
}
}
@@ -625,14 +669,24 @@ class ImapClient(private val reuseConnections: Boolean) {
}
/**
* SPIKE hook (issue #125): closes any kept-alive reused connections (`LOGOUT` + teardown), a no-op
* when the reuse flag is OFF. The reuse harness calls this to force settlement; a shipped feature
* would also drive it from an idle-eviction timer and the low-battery push teardown (#88/#89/#90).
* Tears down every kept-alive reused connection (`LOGOUT` + teardown); a no-op when reuse is
* disabled. `IdleService` drives this on the low-battery push-teardown path (#88/#89/#90), mirroring
* the IDLE connection teardown, and the reuse tests call it to force settlement.
*/
suspend fun closeReusedConnections() {
connectionCache?.closeAll()
}
/**
* Closes any reused connection that has sat unused past the reuse idle timeout (issue #357 Part 2);
* a no-op when reuse is disabled or nothing is idle. `IdleService` calls this on a periodic sweep so
* a socket kept warm for latency doesn't linger and drain battery; a connection currently in use is
* skipped.
*/
suspend fun evictIdleReusedConnections() {
connectionCache?.evictIdle()
}
private fun buildProps(protocol: String, params: ImapConnectionParams, reuse: Boolean = false): Properties =
Properties().apply {
put("mail.store.protocol", protocol)
@@ -664,6 +718,13 @@ class ImapClient(private val reuseConnections: Boolean) {
private companion object {
const val TIMEOUT_MS = "15000"
const val TAG = "LibreMailIdle"
const val PERF_TAG = "ImapPerf"
const val NANOS_PER_MS = 1_000_000L
// A reused connection unused for this long is idle-evicted (issue #357 Part 2): long enough to
// stay warm across an active reading session, well under typical server idle timeouts (Gmail
// ~30 min) so eviction, not a server drop, is what usually closes it.
const val DEFAULT_REUSE_IDLE_TIMEOUT_MS = 5 * 60_000L
}
}
@@ -7,83 +7,172 @@ import jakarta.mail.Store
import jakarta.mail.StoreClosedException
import kotlinx.coroutines.sync.Mutex
import kotlinx.coroutines.sync.withLock
import org.eclipse.angus.mail.iap.ConnectionException
import org.libremail.domain.model.ImapConnectionParams
import org.libremail.reporting.AppLog
import java.io.IOException
import java.util.concurrent.ConcurrentHashMap
import java.util.concurrent.atomic.AtomicInteger
/**
* SPIKE (issue #125): a per-account keep-alive cache of authenticated IMAP [Store]s, so folder-opens
* and message operations reuse one already-connected session instead of re-paying
* `CONNECT + TLS + LOGIN` on every call. See `docs/perf/issue-125-connection-reuse-spike.md`.
* A per-account keep-alive cache of authenticated IMAP [Store]s, so folder-opens and message
* operations reuse one already-connected session instead of re-paying `CONNECT + TLS + LOGIN` on
* every call. Wiring the reuse path proven by the #125 spike; the production default is ON (see
* `BuildConfig.IMAP_CONNECTION_REUSE`). This is the fix for the on-device finding that Gmail throttles
* LibreMail's connect-per-operation IMAP traffic — collapsing ~one socket per operation to ~one warm
* socket per account removes the throttle's trigger (issue #357 Part 2, `docs/perf/issue-125-*`).
*
* Prototype stance — deliberately the simplest thing that *proves reuse*, leaving the tuning knobs to a
* measured follow-up:
* Design:
* - **One connection per account, mutex-guarded.** Each account key holds a single [Store] behind its
* own [Mutex]; every operation on that account serializes through it. This is the simplest safe
* design and the one the investigation named as the starting point. Its known cost is head-of-line
* blocking — a quick flag toggle can queue behind a slow body download. A bounded pool would trade
* that for more sockets (and a size cap + eviction); not prototyped here.
* - **Lazy, catch-and-retry-once stale handling.** No periodic `NOOP` probe (that would add a
* round-trip to every reused op, partly defeating the point). An operation runs optimistically; if
* it fails with a dropped-connection signal, the socket is rebuilt once and the operation retried.
* own [Mutex]; every operation on that account serializes through it, so the single socket is only
* ever touched by one caller at a time (IMAP is serial per connection). Its known cost is
* head-of-line blocking — a quick flag toggle can queue behind a slow body download. A bounded pool
* would trade that for more sockets; that (and per-provider connection caps) is a separate effort
* (#356/#360-#364), deliberately NOT in scope here.
* - **Transparent stale-connection recovery.** No periodic `NOOP` probe (that would add a round-trip
* to every reused op). An operation runs optimistically; if it fails with a dropped-connection
* signal (server idle-timeout, NAT rebind, network change) the socket is rebuilt once and the
* operation retried, so the caller never sees a spurious error. A second failure clears the slot so
* the next call reconnects. A non-connection error (e.g. "message not found") is never retried, so a
* working socket is never needlessly torn down and a mutation is never re-issued over a live socket.
* - **Idle eviction.** [evictIdle] closes any connection unused for longer than [idleTimeoutMillis]
* (driven by a periodic sweep in `IdleService`), so a socket kept warm for latency doesn't linger
* and drain battery once the user goes idle. It skips any connection currently in use.
* - **Teardown.** [closeAll] evicts everything (`LOGOUT` + socket teardown); `IdleService` drives it
* on the low-battery push-teardown path (#88/#89/#90), mirroring the IDLE connection teardown.
* - **Keyed by connection identity, not the secret.** The OAuth access token
* ([ImapConnectionParams.secret]) rotates; keying on host/port/user/security/mechanism keeps a token
* refresh from orphaning a live, already-authenticated socket. A refreshed secret only matters when
* we actually reconnect, and [connect] is always handed the current [params].
*
* Coexists with IMAP IDLE: `ImapClient.idle` holds its own dedicated long-lived [Store] (not in this
* cache), so reuse adds at most one more persistent socket per account — well under provider limits
* (Gmail ~15).
*
* Thread-safety: [ImapClient]'s UI operations are not otherwise serialized and prefetch runs outside
* the syncer's mutex, so [withStore] must be safe under concurrent callers for the same account — the
* per-key mutex provides that. Not wired to any lifecycle/battery signal yet: [closeAll] is the only
* eviction and is driven by the harness today; an idle-eviction timer and low-battery teardown
* (#88/#89/#90) are follow-ups.
* the syncer's mutex, so every entry point here is safe under concurrent callers for the same account —
* the per-key [Mutex] provides that, and [evictIdle] takes it non-blockingly so a sweep never stalls
* behind (or interrupts) an in-flight operation.
*
* @param connect builds and authenticates a fresh [Store] for the given params (blocking network I/O).
* @param idleTimeoutMillis how long a cached connection may sit unused before [evictIdle] closes it.
* @param nowNanos monotonic clock source (injected for deterministic idle-eviction tests).
*/
internal class ImapConnectionCache(private val connect: (ImapConnectionParams) -> Store) {
internal class ImapConnectionCache(
private val connect: (ImapConnectionParams) -> Store,
private val idleTimeoutMillis: Long,
private val nowNanos: () -> Long = System::nanoTime,
) {
private class Entry {
/**
* One account's reused connection. [id] is an opaque per-cache ordinal used only for PII-free log
* correlation — it is NOT derived from the host/username/secret, so a log line can attribute an
* event to an account without ever naming it.
*/
private class Entry(val id: Int) {
val mutex = Mutex()
@Volatile
var store: Store? = null
@Volatile
var lastUsedAtNanos: Long = 0L
}
private val entries = ConcurrentHashMap<String, Entry>()
private val nextId = AtomicInteger(0)
/**
* Runs [block] against a reused, authenticated [Store] for [params]'s account: it is established on
* first use and kept open afterwards, so only the first call pays connection setup. Serialized per
* account by the key's [Mutex]. If the operation hits a dropped connection the socket is rebuilt
* once and the operation retried; a second failure clears the slot so the next call reconnects.
* Runs [block] against a reused, authenticated [Store] for [params]'s account: established on first
* use and kept open afterwards, so only the first call pays connection setup. Serialized per account
* by the key's [Mutex]. Transparently reconnects once if the cached socket has been dropped. [op] is
* a short, PII-free intent label (`body-fetch`, `backfill-page`, …) for the perf breadcrumb.
*/
suspend fun <T> withStore(params: ImapConnectionParams, block: (Store) -> T): T {
val entry = entries.computeIfAbsent(key(params)) { Entry() }
return entry.mutex.withLock {
val store = entry.store ?: connect(params).also { entry.store = it }
suspend fun <T> withStore(params: ImapConnectionParams, op: String, block: (Store) -> T): T {
val entry = entries.computeIfAbsent(key(params)) { Entry(nextId.incrementAndGet()) }
return entry.mutex.withLock { runReusing(entry, params, op, block) }
}
/** Establishes-or-reuses the account's [Store], runs [block], and reconnects once on a dropped socket. */
private fun <T> runReusing(entry: Entry, params: ImapConnectionParams, op: String, block: (Store) -> T): T {
val connectMs = ensureConnected(entry, params) // 0ms when the live connection is reused
entry.lastUsedAtNanos = nowNanos()
val workStart = nowNanos()
return try {
block(requireNotNull(entry.store))
} catch (e: Throwable) {
if (!isConnectionDrop(e)) throw e
// Stale socket: rebuild once and retry so the caller never sees the drop.
AppLog.d(TAG, "reuse stale acct=${entry.id}; reconnecting", e)
reconnectAndRetry(entry, params, block)
} finally {
entry.lastUsedAtNanos = nowNanos()
AppLog.d(PERF_TAG, "$op connect=${connectMs}ms work=${elapsedMs(workStart)}ms live=${liveCount()}")
}
}
/** Reuses the live [Store] (0ms) or connects a fresh one, returning the connect cost in ms. */
private fun ensureConnected(entry: Entry, params: ImapConnectionParams): Long {
if (entry.store != null) {
AppLog.d(TAG, "reuse hit acct=${entry.id}")
return 0L
}
val start = nowNanos()
entry.store = connect(params)
val ms = elapsedMs(start)
AppLog.d(TAG, "reuse open acct=${entry.id} connect=${ms}ms live=${liveCount()}")
return ms
}
/** Closes the dropped socket, reconnects once, and retries [block]; a second failure clears the slot. */
private fun <T> reconnectAndRetry(entry: Entry, params: ImapConnectionParams, block: (Store) -> T): T {
runCatching { entry.store?.close() }
entry.store = null
entry.store = connect(params)
entry.lastUsedAtNanos = nowNanos()
AppLog.d(TAG, "reuse reconnected acct=${entry.id} live=${liveCount()}")
return try {
block(requireNotNull(entry.store))
} catch (retry: Throwable) {
runCatching { entry.store?.close() }
entry.store = null
AppLog.w(TAG, "reuse reconnect failed acct=${entry.id}", retry)
throw retry
}
}
/**
* Closes and forgets every connection whose last use is older than [idleTimeoutMillis] (`LOGOUT` +
* teardown). Takes each account's lock non-blockingly, so a connection currently in use is left
* untouched and the sweep never stalls behind a slow operation. No-op when nothing is cached.
*/
suspend fun evictIdle() {
if (entries.isEmpty()) return
val cutoffNanos = idleTimeoutMillis * NANOS_PER_MS
for ((_, entry) in entries) {
if (!entry.mutex.tryLock()) continue // in use — skip, don't interrupt or wait
try {
block(store)
} catch (e: Throwable) {
if (!isConnectionDrop(e)) throw e
// Stale socket (server idle-timeout, NAT rebind, network change): rebuild once and retry.
runCatching { store.close() }
// Forget the dead socket before reconnecting, so a failed connect leaves a clean slot.
entry.store = null
val fresh = connect(params)
entry.store = fresh
try {
block(fresh)
} catch (retry: Throwable) {
runCatching { fresh.close() }
val store = entry.store
if (store != null && nowNanos() - entry.lastUsedAtNanos >= cutoffNanos) {
runCatching { store.close() }
entry.store = null
throw retry
AppLog.d(TAG, "reuse evict idle acct=${entry.id} live=${liveCount()}")
}
} finally {
entry.mutex.unlock()
}
}
}
/** Closes and forgets every cached connection (`LOGOUT` + socket teardown). The only eviction today. */
/** Closes and forgets every cached connection (`LOGOUT` + socket teardown). No-op when empty. */
suspend fun closeAll() {
if (entries.isEmpty()) return
for ((_, entry) in entries) {
entry.mutex.withLock {
entry.store?.let { store -> runCatching { store.close() } }
entry.store?.let { store ->
runCatching { store.close() }
AppLog.d(TAG, "reuse teardown acct=${entry.id}")
}
entry.store = null
}
}
@@ -97,15 +186,42 @@ internal class ImapConnectionCache(private val connect: (ImapConnectionParams) -
private fun key(params: ImapConnectionParams): String =
"${params.host}|${params.port}|${params.security}|${params.username}|${params.useXoauth2}"
/** Count of currently-held reused sockets, for the PII-free log breadcrumb (approximate under races). */
private fun liveCount(): Int = entries.values.count { it.store != null }
private fun elapsedMs(startNanos: Long): Long = (nowNanos() - startNanos) / NANOS_PER_MS
/**
* Whether [error] signals a dropped connection (retry on a fresh socket) rather than a genuine
* protocol/application error (propagate as-is). Deliberately narrow: a plain [MessagingException]
* for a real server error whose connection is still live is NOT retried, so we never re-issue a
* mutation over a working connection.
* protocol/application error (propagate as-is). A server idle-timeout / NAT rebind / network change
* surfaces as a [FolderClosedException], a [StoreClosedException], a raw [IOException] or Angus's own
* [ConnectionException] — or, most commonly for `folder.open()` on a server-dropped socket, a plain
* [MessagingException] *caused by* one of those ("Connection dropped by server?"). Deliberately
* still narrow: a [MessagingException] caused by anything else (a `CommandFailedException` /
* `BadCommandException` — a real server NO on a live connection) is NOT retried, so a working socket
* is never needlessly torn down.
*
* NB (residual, tracked as the deferred mutation-idempotency review — see the #125 spike doc): the
* retry re-runs the whole operation, so a *mutation* (flag/move/expunge) dropped mid-flight is
* at-least-once. Flag sets are idempotent; the dominant real case — a socket the server dropped
* while idle, detected on the next op's first command before any mutation is issued — is safe. A
* copy-then-expunge move interrupted between its two halves is the rare exception left to that review.
*/
private fun isConnectionDrop(error: Throwable): Boolean = when (error) {
is FolderClosedException, is StoreClosedException, is IOException -> true
is MessagingException -> error.cause is IOException
else -> false
private fun isConnectionDrop(error: Throwable): Boolean {
// FolderClosedException/StoreClosedException are themselves MessagingException subtypes, so these
// definite-drop checks must run before the MessagingException guard below — otherwise they'd fall
// into it and get gated on a `.cause` they don't carry, instead of the unconditional `true` below.
when {
error is FolderClosedException || error is StoreClosedException -> return true
error is IOException || error is ConnectionException -> return true
error !is MessagingException -> return false
else -> return error.cause is IOException || error.cause is ConnectionException
}
}
private companion object {
const val TAG = "ImapReuse"
const val PERF_TAG = "ImapPerf"
const val NANOS_PER_MS = 1_000_000L
}
}
@@ -0,0 +1,65 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.push
import android.app.Service
/**
* The dataSync foreground-start decision for [IdleService.onStartCommand] (#354), pulled out of the
* Android [Service] so it is JVM-unit-testable without Robolectric (which this repo does not use).
*
* After the Android 14+ dataSync FGS runtime cap fires (#302) and [IdleService] self-stops, the
* service is (re)started — the app deterministically restarts it via
* `LibreMailApplication.ensurePushStarted()`, and previously the platform also auto-restarted it via
* `START_STICKY` with a null intent. Each restart re-entered `onStartCommand` and unconditionally
* called `startForeground(..., dataSync)` while the rolling-24h budget was still exhausted, so the
* platform rejected it with `ForegroundServiceStartNotAllowedException` (uncaught → crash → sticky
* restart → loop). This seam encodes the fix: skip the start outright while still inside the cap
* window, otherwise attempt it and route the rejection — a `ForegroundServiceStartNotAllowedException`
* (API 31+), caught here via its [IllegalStateException] supertype so no `minSdk`-29 class load is
* needed — into a clean degrade instead of letting it propagate. Any other throwable propagates.
*/
internal object IdleForegroundStarter {
/**
* Enters dataSync foreground state, or degrades to the 15-minute periodic-sync fallback when that
* start is — or would be — illegal.
*
* @param capActive true while still within the runtime-cap window recorded at the last cap event;
* the start is then skipped without attempting it (and without a cause).
* @param enterForeground the real `ServiceCompat.startForeground(..., dataSync)` call; may throw
* `ForegroundServiceStartNotAllowedException` (an [IllegalStateException]) when the runtime cap is
* exhausted or the start raced into the background.
* @param onStarted run after a successful foreground start (proceed with IDLE watchers).
* @param onDegraded run when the start was skipped ([capActive]) or rejected; receives the
* rejection cause, or `null` for the cap-window skip. Must schedule periodic sync and stop the
* service (it was started via `startForegroundService`, so it must stop promptly to avoid the
* "did not call startForeground in time" ANR).
* @return the value `onStartCommand` should return — always [Service.START_NOT_STICKY]: push is
* app-managed, so the platform's sticky null-intent auto-restart is redundant and fires exactly
* in the states that cannot legally start a dataSync FGS.
*/
fun startForegroundOrDegrade(
capActive: Boolean,
enterForeground: () -> Unit,
onStarted: () -> Unit,
onDegraded: (cause: Throwable?) -> Unit,
): Int {
when {
capActive -> onDegraded(null)
else -> {
val entered = try {
enterForeground()
true
} catch (rejected: IllegalStateException) {
// ForegroundServiceStartNotAllowedException (API 31+) extends IllegalStateException;
// catching the supertype funnels the runtime-cap/background rejection into the
// degrade path without a version gate, instead of crash-looping (#354).
onDegraded(rejected)
false
}
if (entered) onStarted()
}
}
return Service.START_NOT_STICKY
}
}
@@ -8,7 +8,7 @@ import android.content.Intent
import android.content.pm.PackageManager
import android.content.pm.ServiceInfo
import android.os.IBinder
import android.util.Log
import android.os.SystemClock
import androidx.core.app.NotificationManagerCompat
import androidx.core.app.ServiceCompat
import androidx.core.content.ContextCompat
@@ -38,6 +38,7 @@ import org.libremail.domain.model.Account
import org.libremail.mail.ImapClient
import org.libremail.power.BatteryStatusProvider
import org.libremail.reporting.AppLog
import org.libremail.reporting.accountLogRef
import javax.inject.Inject
/**
@@ -83,22 +84,72 @@ class IdleService : Service() {
/** Active IDLE watcher per account id, so we can start/stop them as accounts change. */
private val watchers = mutableMapOf<String, Job>()
override fun onStartCommand(intent: Intent?, flags: Int, startId: Int): Int {
startAsForeground(shownMode)
if (!watching) {
watching = true
scope.launch {
// Can't open the encrypted DB without the user present. Defer (stop) and let the app
// restart push after the next unlock, rather than block the service and ANR.
if (cacheGuard.isCacheLocked()) {
Log.i(TAG, "encrypted cache locked; deferring IDLE push until the app is unlocked")
stopSelf()
return@launch
}
reconcileWatchers()
override fun onStartCommand(intent: Intent?, flags: Int, startId: Int): Int =
// Always START_NOT_STICKY (never START_STICKY): push is app-managed — LibreMailApplication's
// settings/account collector and ensurePushStarted() deterministically (re)start the service
// whenever it should run — so the platform's sticky null-intent auto-restart is redundant AND,
// after the dataSync FGS runtime cap (#302), re-enters here exactly when a dataSync foreground
// start is illegal, which was the #354 crash loop. The start/degrade decision lives in the
// JVM-testable IdleForegroundStarter seam.
IdleForegroundStarter.startForegroundOrDegrade(
capActive = capWindowActive(),
enterForeground = { startAsForeground(shownMode) },
onStarted = ::startWatchingIfNeeded,
onDegraded = ::degradeAfterBlockedForegroundStart,
)
/** After a successful foreground start, begin watching accounts for IDLE (once per service life). */
private fun startWatchingIfNeeded() {
if (watching) return
watching = true
scope.launch {
// Can't open the encrypted DB without the user present. Defer (stop) and let the app
// restart push after the next unlock, rather than block the service and ANR.
if (cacheGuard.isCacheLocked()) {
AppLog.i(TAG, "encrypted cache locked; deferring IDLE push until the app is unlocked")
stopSelf()
return@launch
}
reconcileWatchers()
}
// The reuse cache (issue #357 Part 2) keeps interactive/sync IMAP connections warm; sweep
// them so a socket that has gone idle past the reuse timeout is closed rather than left
// draining battery. Independent of push mode — it runs while the service lives.
scope.launch { evictIdleReuseConnectionsLoop() }
}
/**
* A dataSync foreground start was skipped (still inside the runtime-cap window, [cause] null) or
* rejected by the platform ([cause] is the `ForegroundServiceStartNotAllowedException`). Either way
* degrade like the cap handler instead of crashing (#354): log PII-free and fall back to periodic
* sync. On an actual rejection, also (re)arm the cap window so the next restart skips the attempt.
*/
private fun degradeAfterBlockedForegroundStart(cause: Throwable?) {
if (cause == null) {
AppLog.i(TAG, "dataSync FGS cap still active; skipping foreground start, periodic sync covers mail")
} else {
AppLog.w(
TAG,
"dataSync FGS start rejected (runtime cap or background); staying on 15-minute periodic sync",
cause,
)
markCapReached()
}
degradeToPeriodicSync()
}
/**
* Periodically evicts IMAP connections the reuse cache kept warm once they go idle past the reuse
* idle timeout (issue #357 Part 2). A no-op when reuse is disabled or nothing is idle;
* [ImapClient.evictIdleReusedConnections] skips any connection currently in use, so a sweep never
* disturbs an in-flight sync or interactive fetch.
*/
private suspend fun evictIdleReuseConnectionsLoop() {
while (scope.isActive) {
delay(REUSE_EVICTION_SWEEP_MS)
runCatching { imapClient.evictIdleReusedConnections() }
.onFailure { AppLog.w(TAG, "reuse idle-eviction sweep failed", it) }
}
return START_STICKY
}
/**
@@ -152,24 +203,43 @@ class IdleService : Service() {
// here is effectively a no-op — done anyway so the fallback provably exists whenever push is
// paused, without disturbing the running period.
syncScheduler.schedulePeriodicSync()
// Mirror the IDLE teardown for the reuse cache (issue #357 Part 2): drop any warm
// interactive/sync connections so we hold no kept-alive IMAP sockets while conserving
// battery. They re-establish on the next sync/interactive op once battery recovers.
scope.launch { imapClient.closeReusedConnections() }
} else {
AppLog.i(TAG, "Battery recovered: resuming IMAP IDLE push")
}
}
/**
* Clean shutdown for the dataSync FGS runtime-cap timeout (issue #302): re-assert the periodic
* sync fallback, swap the persistent notification to the degraded text and DETACH it so it stays
* posted after we leave foreground state, then stop the service. Stopping foreground state is not
* optional here — a `dataSync` service that is still foreground when its timeout elapses is the
* exact condition the platform force-stops (and throws) on, so we must not keep running as an FGS.
* [stopSelf] then tears down [scope] in [onDestroy], closing the IDLE connections; mail arrives via
* the 15-minute periodic sync until push is started again (next app foreground / cap reset).
* The dataSync FGS runtime-cap timeout (issue #302): record the cap event so the next (re)start
* skips its now-illegal foreground start (#354), then degrade to the periodic-sync fallback. Kept
* fast/synchronous so a `dataSync` service that is still foreground when its timeout elapses — the
* exact condition the platform force-stops (and throws `ForegroundServiceDidNotStopInTimeException`)
* on — leaves foreground state within the grace window.
*/
private fun fallBackToPeriodicSync() {
AppLog.i(TAG, "dataSync FGS runtime cap reached: pausing IMAP IDLE; mail arrives via 15-minute periodic sync")
markCapReached()
degradeToPeriodicSync()
}
/**
* Shared degrade to the 15-minute periodic-sync fallback, used by the runtime-cap timeout
* ([fallBackToPeriodicSync]) and by [onStartCommand] when a dataSync foreground start is skipped or
* rejected (#354): re-assert the periodic sync, swap the persistent notification to the degraded
* text and DETACH it so it stays posted after we leave foreground state, then stop the service.
* Leaving foreground state is safe on the [onStartCommand] paths too (never-foregrounded there, so
* `stopForeground` is a no-op), and [stopSelf] must run promptly because that start arrived via
* `startForegroundService` — otherwise the platform raises the "did not call startForeground in
* time" ANR. [stopSelf] then tears down [scope] in [onDestroy], closing the IDLE connections; mail
* arrives via the 15-minute periodic sync until push is started again (next app foreground / cap
* reset).
*/
// Permission is checked via hasNotificationPermission() below; lint can't trace the indirect guard.
@SuppressLint("MissingPermission")
private fun fallBackToPeriodicSync() {
AppLog.i(TAG, "dataSync FGS runtime cap reached: pausing IMAP IDLE; mail arrives via 15-minute periodic sync")
private fun degradeToPeriodicSync() {
// Already scheduled at every app start (UPDATE, so a no-op here) — re-asserted so the fallback
// provably exists now that push is paused, mirroring the low-battery path in onPushModeChanged.
syncScheduler.schedulePeriodicSync()
@@ -187,6 +257,25 @@ class IdleService : Service() {
stopSelf()
}
/** Records the wall-independent time of the last dataSync cap event, arming [capWindowActive]. */
private fun markCapReached() {
capReachedElapsedMs = SystemClock.elapsedRealtime()
}
/**
* True while still within [CAP_WINDOW_MS] of the last cap event ([markCapReached]) — a burst of
* restarts in that window is certainly still capped, so [onStartCommand] skips the foreground start
* (and its now-guaranteed rejection) entirely. The window is anchored to the last real cap event
* and never refreshed by the skip itself, so it expires and lets a later restart re-probe; that
* probe is safe because a still-capped rejection is caught. Uses [SystemClock.elapsedRealtime] (not
* wall-clock) so it is immune to clock changes, and the companion field survives service
* re-creation within the process (which is where the restart storm happens).
*/
private fun capWindowActive(): Boolean {
val reachedAt = capReachedElapsedMs
return reachedAt != 0L && SystemClock.elapsedRealtime() - reachedAt < CAP_WINDOW_MS
}
private fun hasNotificationPermission(): Boolean =
ContextCompat.checkSelfPermission(this, Manifest.permission.POST_NOTIFICATIONS) ==
PackageManager.PERMISSION_GRANTED
@@ -200,6 +289,7 @@ class IdleService : Service() {
* idle()'s on-connect sync, so no mail is missed across renewals.
*/
private suspend fun watchAccount(account: Account) {
AppLog.i(TAG, "IDLE watch start ${accountLogRef(account.id)}")
var backoffMs = INITIAL_BACKOFF_MS
while (scope.isActive) {
try {
@@ -212,7 +302,7 @@ class IdleService : Service() {
} catch (e: CancellationException) {
throw e
} catch (e: Exception) {
Log.w(TAG, "IDLE for ${account.email} dropped; retrying in ${backoffMs}ms", e)
AppLog.w(TAG, "IDLE for ${accountLogRef(account.id)} dropped; retrying in ${backoffMs}ms", e)
delay(backoffMs)
backoffMs = (backoffMs * 2).coerceAtMost(MAX_BACKOFF_MS)
}
@@ -242,8 +332,24 @@ class IdleService : Service() {
const val INITIAL_BACKOFF_MS = 5_000L
const val MAX_BACKOFF_MS = 5 * 60_000L
// Cadence of the reuse-cache idle-eviction sweep (issue #357 Part 2). Tighter than the reuse
// idle timeout so an idle socket is closed shortly after it crosses it.
const val REUSE_EVICTION_SWEEP_MS = 2 * 60_000L
// Re-establish IDLE on this cadence — under RFC 2177's 29-minute ceiling and short enough
// to beat typical NAT/firewall idle-socket timeouts.
const val IDLE_RENEWAL_MS = 9 * 60_000L
// How long after a dataSync cap event onStartCommand skips the (still-illegal) foreground start
// outright (#354). The true rolling-24h budget reset is unknowable client-side, so this is a
// restart-storm damper, not a precise predictor: it matches the periodic-sync interval — the
// fallback already covering mail — so at most one foreground-start probe happens per cycle, and
// re-probing after it is safe because a still-capped rejection is caught, not fatal.
const val CAP_WINDOW_MS = 15 * 60_000L
// elapsedRealtime() of the last dataSync cap event; 0 = none this process. Companion-scoped so
// it survives IdleService re-creation within the process, where the restart storm happens.
@Volatile
private var capReachedElapsedMs = 0L
}
}
@@ -0,0 +1,39 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.reporting
import java.security.MessageDigest
/**
* A short, stable, non-reversible reference for an account — safe to write to logs and a [DebugReport].
*
* [Account.id][org.libremail.domain.model.Account.id] embeds the raw email address (e.g.
* `"outlook:user@domain.com"` or `"imap:user@domain.com"`), so it is PII and must **never** be logged
* directly. This returns the id's scheme prefix followed by a truncated SHA-256 of the whole id — e.g.
* `"outlook:a1b2c3"` — which is:
*
* - **stable**: the same id always maps to the same reference (a hash, not a random token), so lines
* for one account correlate across a session and across reports;
* - **non-reversible**: a truncated one-way hash cannot be turned back into the email;
* - **non-PII**: the scheme (`outlook`/`imap`) names the auth kind, not the user, and the hex hash
* contains no `@`, domain, or local part. If the part before the first `:` is not a bare token
* (e.g. an id that is itself an address), it is replaced with [GENERIC_SCHEME] so no address can
* leak through the prefix.
*/
fun accountLogRef(accountId: String): String {
val candidate = accountId.substringBefore(':', missingDelimiterValue = "")
val scheme = candidate.takeIf(::isBareScheme) ?: GENERIC_SCHEME
return "$scheme:${sha256Hex(accountId).take(REF_HASH_LENGTH)}"
}
/** Prefix used when an id has no scheme, or one that could itself carry PII (an `@`, a dot, …). */
private const val GENERIC_SCHEME = "acct"
/** Hex characters of the SHA-256 kept in the reference; short but collision-safe for a device's few accounts. */
private const val REF_HASH_LENGTH = 6
/** A bare scheme token is letters/digits only, so an address (with `@`/`.`) can never pass as a scheme. */
private fun isBareScheme(text: String): Boolean = text.isNotEmpty() && text.all(Char::isLetterOrDigit)
private fun sha256Hex(input: String): String = MessageDigest.getInstance("SHA-256")
.digest(input.toByteArray(Charsets.UTF_8))
.joinToString("") { "%02x".format(it) }
@@ -8,6 +8,12 @@ import android.util.Log
* be attached to a user-reviewed [DebugReport]. Call [install] once at startup. Never pass PII
* (email addresses, message content, credentials) to these methods — the buffer can end up in a
* report the user reviews and may submit.
*
* The throwable-carrying [d]/[w]/[e] overloads forward the throwable to Logcat as usual, but record a
* **PII-scrubbed** rendering of its stack trace into the buffer via [StackTraceScrubber]: exception
* class names and stack frames are kept, while exception *messages* — which can embed a server host or
* an account email (e.g. a `ConnectException` or an auth failure) — are stripped. To refer to an
* account in a log line, log [accountLogRef]`(account.id)` rather than the raw id, which is PII.
*/
object AppLog {
@Volatile
@@ -22,6 +28,11 @@ object AppLog {
buffer?.record('D', tag, message)
}
fun d(tag: String, message: String, throwable: Throwable?) {
Log.d(tag, message, throwable)
buffer?.record('D', tag, bufferLine(message, throwable))
}
fun i(tag: String, message: String) {
Log.i(tag, message)
buffer?.record('I', tag, message)
@@ -32,8 +43,23 @@ object AppLog {
buffer?.record('W', tag, message)
}
fun w(tag: String, message: String, throwable: Throwable?) {
Log.w(tag, message, throwable)
buffer?.record('W', tag, bufferLine(message, throwable))
}
fun e(tag: String, message: String, throwable: Throwable? = null) {
Log.e(tag, message, throwable)
buffer?.record('E', tag, message)
buffer?.record('E', tag, bufferLine(message, throwable))
}
/**
* The buffer text for a log call: the caller [message] alone, or — when a [throwable] is present —
* the message followed by the throwable's **scrubbed** stack trace (exception class names and stack
* frames kept; host/email-bearing exception messages stripped by [StackTraceScrubber]).
*/
private fun bufferLine(message: String, throwable: Throwable?): String {
if (throwable == null) return message
return "$message\n" + StackTraceScrubber.scrub(throwable.stackTraceToString())
}
}
@@ -0,0 +1,35 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.reporting
/**
* At-rest encryption seam for persisted [DebugReport]s (issue #369). When the opt-in cache encryption
* (`encryptCache`) is ON, [ReportStore] seals a report's storage JSON with this before writing it and
* unseals it on read; when OFF the JSON is stored plaintext exactly as before.
*
* Kept behind an interface so [ReportStore] stays a plain `java.io.File` component that JVM-tests
* without the Android Keystore. The on-device implementation
* (`org.libremail.data.security.KeystoreReportEncryption`) wraps the vetted Keystore crypto; [None] is
* the plaintext default used where encryption isn't wired and by tests that don't exercise it.
*/
interface ReportEncryption {
/**
* Whether reports must be encrypted at rest right now — a synchronous, crash-safe mirror of the
* `encryptCache` setting, never a live DataStore read (a crash-time [ReportStore.save] runs on the
* crashing thread). When false, [encrypt]/[decrypt] are never called.
*/
fun enabled(): Boolean
/** Seals [plaintext] to an opaque at-rest blob. Only invoked when [enabled] is true. */
fun encrypt(plaintext: String): String
/** Unseals a blob produced by [encrypt] back to the original plaintext. */
fun decrypt(encoded: String): String
/** The plaintext default: encryption disabled, so [encrypt]/[decrypt] are pass-throughs never used. */
object None : ReportEncryption {
override fun enabled(): Boolean = false
override fun encrypt(plaintext: String): String = plaintext
override fun decrypt(encoded: String): String = encoded
}
}
@@ -22,10 +22,19 @@ import java.io.File
* to [scope] (IO by default). Reactive consumers (Problem Reports, the startup prompt) observe
* [reports] and update when the scan lands; the empty window is momentary. Writes ([save] etc.)
* re-scan synchronously so a crash-time save is never lost to the pending initial scan.
*
* At-rest encryption (issue #369): when [encryption] reports it is [ReportEncryption.enabled], each
* report's storage JSON is sealed before it is written and tagged with [ENCRYPTED_PREFIX]; when it is
* off the JSON is stored plaintext exactly as before. Reads sniff the prefix, so plaintext reports from
* before the setting was turned on and sealed reports written after it coexist transparently. Writes
* FAIL CLOSED: if sealing throws while encryption is on, the report is dropped rather than written in
* plaintext, so the user's opt-in encryption is never silently defeated by leaving a plaintext report
* on disk. The default [ReportEncryption.None] keeps the store plaintext (its historical behaviour).
*/
class ReportStore(
private val directory: File,
scope: CoroutineScope = CoroutineScope(SupervisorJob() + Dispatchers.IO),
private val encryption: ReportEncryption = ReportEncryption.None,
) {
private val lock = Any()
private val _reports = MutableStateFlow<List<DebugReport>>(emptyList())
@@ -41,8 +50,9 @@ class ReportStore(
fun save(report: DebugReport) {
synchronized(lock) {
val serialized = serializeForDisk(report) ?: return
directory.mkdirs()
File(directory, fileName(report.id)).writeText(report.toStorageJson())
File(directory, fileName(report.id)).writeText(serialized)
_reports.value = scan()
}
}
@@ -58,7 +68,8 @@ class ReportStore(
synchronized(lock) {
val report = _reports.value.firstOrNull { it.id == id } ?: return
if (report.surfaced) return
File(directory, fileName(id)).writeText(report.copy(surfaced = true).toStorageJson())
val serialized = serializeForDisk(report.copy(surfaced = true)) ?: return
File(directory, fileName(id)).writeText(serialized)
_reports.value = scan()
}
}
@@ -85,13 +96,57 @@ class ReportStore(
val files = directory.listFiles { file -> file.isFile && file.name.endsWith(SUFFIX) }
?: return emptyList()
return files
.mapNotNull { file -> runCatching { DebugReport.fromStorageJson(file.readText()) }.getOrNull() }
.mapNotNull { file -> readReport(file) }
.sortedByDescending { it.createdAtMillis }
}
/**
* Renders [report] for on-disk storage. With encryption OFF this is the plaintext storage JSON,
* byte-for-byte as before. With it ON the JSON is sealed and tagged with [ENCRYPTED_PREFIX]. Returns
* `null` — so the caller writes nothing — when sealing fails while encryption is ON: the report is
* deliberately NOT written in plaintext (that would defeat the user's opt-in encryption, #369), so a
* failed seal drops the report rather than leaking it. A dropped crash report still lets the original
* crash propagate to the system handler.
*/
private fun serializeForDisk(report: DebugReport): String? {
val json = report.toStorageJson()
if (!encryption.enabled()) return json
return runCatching { ENCRYPTED_PREFIX + encryption.encrypt(json) }.getOrElse { e ->
AppLog.e(TAG, "Report encryption failed; not persisting to avoid a plaintext report on disk", e)
null
}
}
/**
* Reads one stored report, transparently unsealing files tagged with [ENCRYPTED_PREFIX]. A file that
* cannot be decrypted (e.g. the master key was cleared) is logged and skipped rather than crashing
* the list; an unparseable plaintext file is skipped silently, as before.
*/
private fun readReport(file: File): DebugReport? {
val raw = runCatching { file.readText() }.getOrNull() ?: return null
val json = if (raw.startsWith(ENCRYPTED_PREFIX)) {
runCatching { encryption.decrypt(raw.removePrefix(ENCRYPTED_PREFIX)) }.getOrElse { e ->
AppLog.w(TAG, "Skipping a stored report that could not be decrypted", e)
return null
}
} else {
raw
}
return runCatching { DebugReport.fromStorageJson(json) }.getOrNull()
}
private fun fileName(id: String) = "$id$SUFFIX"
private companion object {
const val SUFFIX = ".json"
const val TAG = "ReportStore"
/**
* Marks a file whose body is `Base64(iv || ciphertext)` rather than plaintext JSON. Contains
* characters outside Base64's alphabet (`.`/`:`) and never matches a plaintext report's leading
* `{`, so a read tells sealed from plaintext files unambiguously — the two coexist on disk after
* the encryption setting is toggled.
*/
const val ENCRYPTED_PREFIX = "libremail.report.enc.v1:"
}
}
@@ -5,7 +5,7 @@ import android.app.Activity
import android.content.Intent
import android.os.Bundle
import android.os.Process
import android.util.Log
import org.libremail.reporting.AppLog
/**
* Separate-process trampoline that performs an app relaunch from OUTSIDE the process being killed.
@@ -40,7 +40,10 @@ class RestartActivity : Activity() {
if (launchIntent != null) {
startActivity(launchIntent)
} else {
Log.w(TAG, "no launch intent for $packageName; cannot relaunch after restart")
// Logcat-only here: this runs in the ":restart" trampoline process, where
// LibreMailApplication.onCreate returns early and never calls AppLog.install, so the
// buffer is null and this breadcrumb never reaches a DebugReport.
AppLog.w(TAG, "no launch intent for $packageName; cannot relaunch after restart")
}
finish()
@@ -3,7 +3,6 @@ package org.libremail.ui.accountsetup
import android.content.ActivityNotFoundException
import android.content.Intent
import android.util.Log
import androidx.lifecycle.ViewModel
import androidx.lifecycle.viewModelScope
import dagger.hilt.android.lifecycle.HiltViewModel
@@ -15,6 +14,7 @@ import kotlinx.coroutines.launch
import org.libremail.auth.OutlookAuthManager
import org.libremail.domain.model.Account
import org.libremail.domain.repository.AccountRepository
import org.libremail.reporting.AppLog
import javax.inject.Inject
/** Stage of an account-setup attempt, shared by the Outlook and manual flows. */
@@ -68,12 +68,18 @@ class AccountSetupViewModel @Inject constructor(
Account.outlook(oauth.email).id
}.fold(
onSuccess = { accountId ->
// No email: the account id embeds it (see accountLogRef) and must never be logged.
AppLog.i(TAG, "Outlook account added")
_state.update { it.copy(status = SetupStatus.DONE, addedAccountId = accountId) }
},
onFailure = { e ->
// Stripped from release builds by the Log.d ProGuard rule (keeps any account
// address / token detail out of shipped logs); visible in debug for diagnosis.
Log.d(TAG, "Outlook sign-in failed after redirect", e)
// AppLog.d's Logcat mirror is stripped from release builds by the -assumenosideeffects
// Log.d ProGuard rule (keeps any account address / token detail out of shipped
// logcat), but that rule only elides the `Log.d(...)` call inside AppLog.d — the
// buffer.record(...) line right after it is untouched, so this breadcrumb still
// reaches a submitted report. The throwable's message may carry the account
// email/token; AppLog's StackTraceScrubber redacts it before it is recorded.
AppLog.d(TAG, "Outlook sign-in failed after redirect", e)
_state.update {
it.copy(status = SetupStatus.IDLE, error = e.message ?: "Microsoft sign-in failed")
}
@@ -5,7 +5,6 @@ import android.content.Context
import android.os.SystemClock
import android.security.keystore.KeyPermanentlyInvalidatedException
import android.security.keystore.UserNotAuthenticatedException
import android.util.Log
import androidx.annotation.VisibleForTesting
import androidx.lifecycle.ViewModel
import androidx.lifecycle.viewModelScope
@@ -32,6 +31,7 @@ import org.libremail.data.security.LockState
import org.libremail.data.security.PassphraseSession
import org.libremail.data.settings.SettingsRepository
import org.libremail.data.sync.SyncScheduler
import org.libremail.reporting.AppLog
import org.libremail.restart.ProcessRestarter
import java.util.concurrent.ExecutionException
import java.util.concurrent.TimeUnit
@@ -149,6 +149,8 @@ class AppLockViewModel @Inject constructor(
keyInvalidated = databaseKeyCipher.isInvalidated(),
)
}
// LockAction is a non-PII enum, so it's safe to record verbatim as a breadcrumb.
AppLog.i(TAG, "app-lock foreground decision: $action")
when (action) {
LockAction.PROCEED -> _uiState.value = AppLockUiState.Unlocked
@@ -191,6 +193,7 @@ class AppLockViewModel @Inject constructor(
viewModelScope.launch {
when (withContext(defaultDispatcher) { unlockOrArm() }) {
UnlockResult.OK -> {
AppLog.i(TAG, "auth seal unlocked; cache readable")
gate.onAuthenticated()
publish()
}
@@ -198,7 +201,7 @@ class AppLockViewModel @Inject constructor(
UnlockResult.UNRECOVERABLE -> {
// The passphrase is permanently unrecoverable (key invalidated or deleted by a
// screen-lock change). Wipe the cache safely at the next cold start and re-sync.
Log.w(TAG, "encrypted cache passphrase unrecoverable; clearing cache")
AppLog.w(TAG, "encrypted cache passphrase unrecoverable; clearing cache")
clearCacheAndRestart(disableAppLock = false)
}
@@ -240,7 +243,7 @@ class AppLockViewModel @Inject constructor(
return runCatching { databaseKeyStore.sealWithAuth() }.fold(
onSuccess = { UnlockResult.OK },
onFailure = { e ->
Log.w(TAG, "arming auth seal failed", e)
AppLog.w(TAG, "arming auth seal failed", e)
UnlockResult.RETRY
},
)
@@ -249,17 +252,17 @@ class AppLockViewModel @Inject constructor(
private suspend fun unwrapSealedPassphrase(): UnlockResult {
// A sealed passphrase exists but its key is gone entirely: it can never be unwrapped.
if (!databaseKeyCipher.hasKey()) {
Log.w(TAG, "auth-sealed passphrase present but key was deleted; cache unrecoverable")
AppLog.w(TAG, "auth-sealed passphrase present but key was deleted; cache unrecoverable")
return UnlockResult.UNRECOVERABLE
}
return try {
databaseKeyStore.unlockWithAuth()
UnlockResult.OK
} catch (e: KeyPermanentlyInvalidatedException) {
Log.w(TAG, "auth-bound key permanently invalidated", e)
AppLog.w(TAG, "auth-bound key permanently invalidated", e)
UnlockResult.UNRECOVERABLE
} catch (e: UserNotAuthenticatedException) {
Log.w(TAG, "auth window elapsed before unwrap; will retry", e)
AppLog.w(TAG, "auth window elapsed before unwrap; will retry", e)
UnlockResult.RETRY
} catch (e: CancellationException) {
throw e
@@ -268,12 +271,13 @@ class AppLockViewModel @Inject constructor(
// process right after a successful auth. Re-lock and let the user retry rather than wiping
// the cache on an ambiguous error (a genuinely lost key still surfaces as UNRECOVERABLE
// via hasKey()/KeyPermanentlyInvalidatedException above).
Log.w(TAG, "unexpected failure unwrapping auth-sealed passphrase; will retry", e)
AppLog.w(TAG, "unexpected failure unwrapping auth-sealed passphrase; will retry", e)
UnlockResult.RETRY
}
}
private suspend fun clearCacheAndRestart(disableAppLock: Boolean) {
AppLog.w(TAG, "clearing encrypted cache and restarting (disableAppLock=$disableAppLock)")
withContext(defaultDispatcher) {
// Record the wipe intent BEFORE flipping app-lock off, so a crash between the two writes
// leaves the wipe still pending (recoverable) rather than a disabled gate over a stale key.
@@ -302,12 +306,12 @@ class AppLockViewModel @Inject constructor(
try {
operation.result.get(SYNC_ENQUEUE_TIMEOUT_SECONDS, TimeUnit.SECONDS)
} catch (e: TimeoutException) {
Log.w(TAG, "re-sync enqueue not confirmed within timeout; restarting anyway", e)
AppLog.w(TAG, "re-sync enqueue not confirmed within timeout; restarting anyway", e)
} catch (e: ExecutionException) {
Log.w(TAG, "re-sync enqueue failed; restarting anyway", e)
AppLog.w(TAG, "re-sync enqueue failed; restarting anyway", e)
} catch (e: InterruptedException) {
Thread.currentThread().interrupt()
Log.w(TAG, "interrupted awaiting re-sync enqueue; restarting anyway", e)
AppLog.w(TAG, "interrupted awaiting re-sync enqueue; restarting anyway", e)
}
}
@@ -19,6 +19,7 @@ import org.libremail.domain.model.InlineImage
import org.libremail.domain.model.Message
import org.libremail.domain.model.ReplyMode
import org.libremail.domain.repository.MailRepository
import org.libremail.reporting.AppLog
import org.libremail.ui.navigation.Routes
import java.io.File
import javax.inject.Inject
@@ -74,6 +75,10 @@ class ReaderViewModel @Inject constructor(
}
}
viewModelScope.launch {
// Time the spinner: from launch to the state update that clears `loading` — what the user
// actually waits through. Covers openMessage (first-open body fetch) plus inline-image
// resolution, so a slow render can be split from a slow open in a debug report (issue #358).
val startNanos = System.nanoTime()
repository.openMessage(messageId).fold(
onSuccess = { message ->
// Resolve inline cid: images BEFORE the first render and publish them in the SAME
@@ -86,6 +91,11 @@ class ReaderViewModel @Inject constructor(
emptyMap()
}
_state.update { it.copy(loading = false, message = message, inlineImages = images) }
AppLog.d(
READER_TAG,
"reader ready took=${(System.nanoTime() - startNanos) / NANOS_PER_MS}ms " +
"html=${message.isHtml} inline=${images.size}",
)
},
onFailure = { e ->
_state.update {
@@ -95,6 +105,10 @@ class ReaderViewModel @Inject constructor(
e.message ?: "Could not load message",
)
}
AppLog.w(
READER_TAG,
"reader load failed took=${(System.nanoTime() - startNanos) / NANOS_PER_MS}ms",
)
},
)
}
@@ -155,4 +169,9 @@ class ReaderViewModel @Inject constructor(
_state.update { it.copy(deleted = true) }
}
}
private companion object {
const val READER_TAG = "Reader"
const val NANOS_PER_MS = 1_000_000L
}
}
@@ -0,0 +1,265 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.ui.security
import androidx.activity.compose.rememberLauncherForActivityResult
import androidx.activity.result.contract.ActivityResultContracts
import androidx.compose.foundation.layout.Arrangement
import androidx.compose.foundation.layout.Column
import androidx.compose.foundation.layout.Row
import androidx.compose.foundation.layout.Spacer
import androidx.compose.foundation.layout.fillMaxSize
import androidx.compose.foundation.layout.fillMaxWidth
import androidx.compose.foundation.layout.height
import androidx.compose.foundation.layout.padding
import androidx.compose.foundation.layout.width
import androidx.compose.foundation.rememberScrollState
import androidx.compose.foundation.text.selection.SelectionContainer
import androidx.compose.foundation.verticalScroll
import androidx.compose.material.icons.Icons
import androidx.compose.material.icons.automirrored.filled.ArrowBack
import androidx.compose.material.icons.filled.Lock
import androidx.compose.material3.Button
import androidx.compose.material3.CircularProgressIndicator
import androidx.compose.material3.ExperimentalMaterial3Api
import androidx.compose.material3.Icon
import androidx.compose.material3.IconButton
import androidx.compose.material3.MaterialTheme
import androidx.compose.material3.Scaffold
import androidx.compose.material3.SnackbarHost
import androidx.compose.material3.SnackbarHostState
import androidx.compose.material3.Surface
import androidx.compose.material3.Text
import androidx.compose.material3.TextButton
import androidx.compose.material3.TopAppBar
import androidx.compose.runtime.Composable
import androidx.compose.runtime.getValue
import androidx.compose.runtime.mutableStateOf
import androidx.compose.runtime.remember
import androidx.compose.runtime.rememberCoroutineScope
import androidx.compose.runtime.saveable.rememberSaveable
import androidx.compose.runtime.setValue
import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
import androidx.compose.ui.platform.LocalClipboard
import androidx.compose.ui.platform.LocalContext
import androidx.compose.ui.res.stringResource
import androidx.compose.ui.text.font.FontFamily
import androidx.compose.ui.text.style.TextAlign
import androidx.compose.ui.unit.dp
import androidx.hilt.lifecycle.viewmodel.compose.hiltViewModel
import androidx.lifecycle.compose.collectAsStateWithLifecycle
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.launch
import kotlinx.coroutines.withContext
import org.libremail.R
import org.libremail.ui.reporting.copyReportPayloadToClipboard
/**
* Fail-closed gate for the opt-in encrypted cache (issue #359). Wraps the whole app: while
* [CacheEncryptionGateViewModel] resolves whether the encrypted cache can be opened it shows a blank
* cover, and [content] — the real app — composes only once the gate reports
* [CacheEncryptionGateState.Ready]. If SQLCipher's native library will not load, the gate shows
* [CacheEncryptionErrorScreen] instead of ever opening the cache unencrypted or reaching the mailbox.
*
* Hosted INSIDE the app-lock gate (`AppLockGateHost`), so when app-lock is on the passphrase is already
* unlocked before the probe runs.
*/
@Composable
fun CacheEncryptionGate(viewModel: CacheEncryptionGateViewModel = hiltViewModel(), content: @Composable () -> Unit) {
val state by viewModel.state.collectAsStateWithLifecycle()
when (state) {
CacheEncryptionGateState.Checking -> GateCover()
CacheEncryptionGateState.Ready -> content()
CacheEncryptionGateState.Unavailable -> CacheEncryptionErrorFlow(viewModel)
}
}
/** Opaque cover shown while the gate resolves, so no DB-backed screen is visible before the decision. */
@Composable
private fun GateCover() {
Surface(modifier = Modifier.fillMaxSize(), color = MaterialTheme.colorScheme.background) {}
}
/** Error screen ↔ ephemeral report review, kept off the app's NavHost (the app never composed here). */
@Composable
private fun CacheEncryptionErrorFlow(viewModel: CacheEncryptionGateViewModel) {
var reviewing by rememberSaveable { mutableStateOf(false) }
if (reviewing) {
val payload by viewModel.reportPayload.collectAsStateWithLifecycle()
EphemeralReportReviewScreen(
payload = payload,
onBack = {
viewModel.dismissReport()
reviewing = false
},
)
} else {
CacheEncryptionErrorScreen(
onReportProblem = {
viewModel.prepareReport()
reviewing = true
},
)
}
}
/**
* Full-screen fail-closed notice shown when the encrypted cache cannot be opened. Presentational: the
* verbatim error message, a reassurance that nothing was lost, and a "Report a problem" action wired by
* the caller. Deliberately renders no mailbox content.
*/
@Composable
fun CacheEncryptionErrorScreen(onReportProblem: () -> Unit) {
Surface(modifier = Modifier.fillMaxSize(), color = MaterialTheme.colorScheme.background) {
Column(
modifier = Modifier
.fillMaxSize()
.verticalScroll(rememberScrollState())
.padding(32.dp),
horizontalAlignment = Alignment.CenterHorizontally,
verticalArrangement = Arrangement.Center,
) {
Icon(
imageVector = Icons.Filled.Lock,
contentDescription = null,
tint = MaterialTheme.colorScheme.error,
)
Spacer(Modifier.height(16.dp))
Text(
text = stringResource(R.string.cache_encryption_error_title),
style = MaterialTheme.typography.headlineSmall,
textAlign = TextAlign.Center,
)
Spacer(Modifier.height(8.dp))
// The exact maintainer-specified message; do not reword.
Text(
text = stringResource(R.string.cache_encryption_error_message),
style = MaterialTheme.typography.bodyMedium,
color = MaterialTheme.colorScheme.error,
textAlign = TextAlign.Center,
)
Spacer(Modifier.height(16.dp))
Text(
text = stringResource(R.string.cache_encryption_error_body),
style = MaterialTheme.typography.bodyMedium,
color = MaterialTheme.colorScheme.onSurfaceVariant,
textAlign = TextAlign.Center,
)
Spacer(Modifier.height(24.dp))
Button(onClick = onReportProblem) {
Text(stringResource(R.string.cache_encryption_report_action))
}
}
}
}
/**
* Ephemeral review of the PII-free diagnostic report generated from the error gate. The report is held
* only in memory (never written to disk); the copy states that plainly. The user can Copy it to the
* clipboard or Save it to a file they choose — the only ways it leaves this screen.
*/
@OptIn(ExperimentalMaterial3Api::class)
@Composable
private fun EphemeralReportReviewScreen(payload: String?, onBack: () -> Unit) {
val context = LocalContext.current
val clipboard = LocalClipboard.current
val scope = rememberCoroutineScope()
val snackbarHostState = remember { SnackbarHostState() }
val copiedMessage = stringResource(R.string.report_copied)
val savedMessage = stringResource(R.string.report_saved)
val saveLauncher = rememberLauncherForActivityResult(
ActivityResultContracts.CreateDocument("application/json"),
) { uri ->
val text = payload
if (uri != null && text != null) {
scope.launch {
withContext(Dispatchers.IO) {
runCatching {
context.contentResolver.openOutputStream(uri)?.use { it.write(text.toByteArray()) }
}
}
snackbarHostState.showSnackbar(savedMessage)
}
}
}
Scaffold(
topBar = {
TopAppBar(
title = { Text(stringResource(R.string.report_review_title)) },
navigationIcon = {
IconButton(onClick = onBack) {
Icon(
Icons.AutoMirrored.Filled.ArrowBack,
contentDescription = stringResource(R.string.action_back),
)
}
},
)
},
snackbarHost = { SnackbarHost(snackbarHostState) },
) { innerPadding ->
Column(
Modifier
.fillMaxSize()
.padding(innerPadding)
.verticalScroll(rememberScrollState())
.padding(16.dp),
) {
Text(
stringResource(R.string.cache_encryption_report_ephemeral_notice),
style = MaterialTheme.typography.bodyMedium,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
Spacer(Modifier.height(16.dp))
if (payload == null) {
Row(Modifier.fillMaxWidth(), horizontalArrangement = Arrangement.Center) {
CircularProgressIndicator()
}
Spacer(Modifier.height(8.dp))
Text(
stringResource(R.string.cache_encryption_report_generating),
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
} else {
SelectionContainer {
Surface(
color = MaterialTheme.colorScheme.surfaceVariant,
shape = MaterialTheme.shapes.small,
modifier = Modifier.fillMaxWidth(),
) {
Text(
text = payload,
style = MaterialTheme.typography.bodySmall,
fontFamily = FontFamily.Monospace,
modifier = Modifier.padding(12.dp),
)
}
}
Spacer(Modifier.height(16.dp))
Row(Modifier.fillMaxWidth()) {
TextButton(
onClick = {
scope.launch {
copyReportPayloadToClipboard(clipboard, payload)
snackbarHostState.showSnackbar(copiedMessage)
}
},
modifier = Modifier.weight(1f),
) {
Text(stringResource(R.string.report_copy))
}
Spacer(Modifier.width(8.dp))
TextButton(
onClick = { saveLauncher.launch("libremail-report.json") },
modifier = Modifier.weight(1f),
) {
Text(stringResource(R.string.report_save))
}
}
}
}
}
}
@@ -0,0 +1,111 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.ui.security
import androidx.annotation.VisibleForTesting
import androidx.lifecycle.ViewModel
import androidx.lifecycle.viewModelScope
import dagger.hilt.android.lifecycle.HiltViewModel
import kotlinx.coroutines.CoroutineDispatcher
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.flow.MutableStateFlow
import kotlinx.coroutines.flow.StateFlow
import kotlinx.coroutines.flow.asStateFlow
import kotlinx.coroutines.launch
import kotlinx.coroutines.withContext
import org.libremail.data.local.CacheEncryptionUnavailableException
import org.libremail.data.local.DatabaseProvisioner
import org.libremail.reporting.AppLog
import org.libremail.reporting.DiagnosticsCollector
import javax.inject.Inject
/** State of the fail-closed encrypted-cache gate that wraps the app (issue #359). */
sealed interface CacheEncryptionGateState {
/** Resolving whether the encrypted cache can be opened — render a blank cover, never the app. */
data object Checking : CacheEncryptionGateState
/** The cache is openable (encrypted-and-ready, or not encrypted): show the app content. */
data object Ready : CacheEncryptionGateState
/**
* SQLCipher's native library will not load, so the encrypted cache cannot be opened. FAIL CLOSED:
* show the error gate instead of the mailbox — never an unencrypted cache.
*/
data object Unavailable : CacheEncryptionGateState
}
/**
* Drives the fail-closed encrypted-cache gate. On first composition it proactively runs the shared
* startup sequence ([DatabaseProvisioner.prepareCache]) so the cache's open mode is resolved BEFORE any
* DB-backed screen composes. If that raises [CacheEncryptionUnavailableException] — SQLCipher's native
* library failed to load (issue #359) — the gate goes to [CacheEncryptionGateState.Unavailable] and the
* host shows the encryption error screen; otherwise it goes [CacheEncryptionGateState.Ready] and the app
* renders. The failure is never memoized by the provisioner, so a fresh process (e.g. after an app
* update that ships a loadable library) re-probes and recovers automatically.
*
* From the error screen the user can generate an **ephemeral** PII-free diagnostic report
* ([prepareReport]) reusing the app's existing [DiagnosticsCollector]. Because encryption is
* unavailable in this exact moment the report cannot be encrypted at rest, so it is deliberately NOT
* written to `ReportStore` — it lives only in memory for on-screen review and the user's explicit
* Copy/Save.
*
* This VM must be hosted only AFTER the app-lock gate unlocks (see `AppLockGateHost`), so when app-lock
* is on the auth-bound passphrase is already in [org.libremail.data.security.PassphraseSession] and
* `prepareCache()` does not park waiting for authentication.
*/
@HiltViewModel
class CacheEncryptionGateViewModel @Inject constructor(
private val provisioner: DatabaseProvisioner,
private val diagnostics: DiagnosticsCollector,
) : ViewModel() {
// Injectable so the report collection (which touches DataStore + the account store) is pushed off
// the main thread in production yet runs on the test scheduler in unit tests. prepareCache() already
// switches to its own IO dispatcher internally, so the probe does not need this.
@VisibleForTesting
internal var ioDispatcher: CoroutineDispatcher = Dispatchers.IO
private val _state = MutableStateFlow<CacheEncryptionGateState>(CacheEncryptionGateState.Checking)
val state: StateFlow<CacheEncryptionGateState> = _state.asStateFlow()
// The ephemeral report payload, or null before it has been generated / after dismissal. Never
// persisted — it exists only for on-screen review and the user's explicit Copy/Save.
private val _reportPayload = MutableStateFlow<String?>(null)
val reportPayload: StateFlow<String?> = _reportPayload.asStateFlow()
init {
probe()
}
private fun probe() {
viewModelScope.launch {
_state.value = try {
provisioner.prepareCache()
CacheEncryptionGateState.Ready
} catch (e: CacheEncryptionUnavailableException) {
AppLog.w(TAG, "encrypted cache unavailable; showing the fail-closed encryption gate", e)
CacheEncryptionGateState.Unavailable
}
}
}
/**
* Generate the ephemeral PII-free diagnostic report for on-screen review (idempotent while one is
* already prepared). Reuses [DiagnosticsCollector.collectManual] — the same PII-free assembly the
* normal "Report a problem" flow uses — but the result is held only in memory here, never saved.
*/
fun prepareReport() {
if (_reportPayload.value != null) return
viewModelScope.launch {
_reportPayload.value = withContext(ioDispatcher) { diagnostics.collectManual().toSubmissionPayload() }
}
}
/** Drop the in-memory report (on leaving the review screen). */
fun dismissReport() {
_reportPayload.value = null
}
private companion object {
const val TAG = "CacheEncryptionGate"
}
}
+12
View File
@@ -390,6 +390,18 @@
<string name="report_submit_unavailable">Online submission isn\'t available in this build. Use Copy or Save to share the report.</string>
<string name="report_copied">Copied to clipboard</string>
<string name="report_saved">Saved</string>
<!-- Fail-closed encrypted-cache gate (#359): shown when SQLCipher's native library will not load,
so the opt-in encrypted cache cannot be opened. The app fails closed here instead of falling
back to an unencrypted cache. -->
<string name="cache_encryption_error_title">Encrypted mail unavailable</string>
<!-- Verbatim maintainer-specified message; do not reword. -->
<string name="cache_encryption_error_message">Error - decryption could not proceed. Native decryption library load failure.</string>
<string name="cache_encryption_error_body">A required security component could not be loaded, so your encrypted mail can\'t be opened on this device right now. Nothing was changed or deleted — your mail is still on your server. LibreMail will try again the next time you open it.</string>
<string name="cache_encryption_report_action">Report a problem</string>
<!-- Ephemeral report: it is prepared in memory only, so the copy must say it is NOT saved unless
the user explicitly saves/exports it, and that it holds no personal information. -->
<string name="cache_encryption_report_ephemeral_notice">This report contains only diagnostic information — no email addresses, message content, or passwords. It is prepared just for you to review now and is NOT saved to this device unless you tap Save.</string>
<string name="cache_encryption_report_generating">Preparing report…</string>
<string name="crash_prompt_title">LibreMail closed unexpectedly</string>
<string name="crash_prompt_message">A problem report from the last crash is ready for you to review. Nothing is sent automatically.</string>
<string name="crash_prompt_review">Review</string>
@@ -14,6 +14,7 @@ import io.mockk.mockkConstructor
import io.mockk.mockkStatic
import io.mockk.runs
import io.mockk.unmockkAll
import io.mockk.verify
import kotlinx.coroutines.test.runTest
import net.openid.appauth.AuthState
import net.openid.appauth.AuthorizationException
@@ -30,6 +31,7 @@ import org.junit.Test
import org.libremail.BuildConfig
import kotlin.test.assertEquals
import kotlin.test.assertFailsWith
import kotlin.test.assertTrue
/**
* Drives the Outlook/Microsoft OAuth wrapper in a pure JVM test. The AppAuth types it builds
@@ -64,18 +66,23 @@ class OutlookAuthManagerTest {
}
@Test
fun `exchangeToken mints an Exchange token and reads the account email from the id_token`() = runTest {
fun `exchangeToken builds the result from the code-exchange token without a second refresh`() = runTest {
installStatics()
stubResponse(authResponseWithCode("auth-code"))
stubTokenRequests(
tokenResponse(access = "code-access", idToken = jwt("email" to "me@example.com"), refresh = "rt"),
tokenResponse(access = "outlook-access", idToken = null, refresh = "rt"),
tokenResponse(access = "code-access", idToken = jwt("email" to "me@example.com"), refresh = "durable-rt"),
)
val result = authManager().exchangeToken(mockk<Intent>())
assertEquals("me@example.com", result.email)
assertEquals("outlook-access", result.accessToken) // the Exchange-scoped token, not the code one
// The code exchange already named the Exchange Online resource, so its access token is used
// directly for IMAP verification — not a token from a second, redundant refresh round-trip.
assertEquals("code-access", result.accessToken)
// The durable AuthState — carrying the refresh token for later token refreshes — is persisted.
assertTrue(result.authStateJson.contains("durable-rt"))
// Exactly one token request (the code exchange); the redundant second refresh is gone.
verify(exactly = 1) { anyConstructed<AuthorizationService>().performTokenRequest(any(), any()) }
}
@Test
@@ -84,7 +91,6 @@ class OutlookAuthManagerTest {
stubResponse(authResponseWithCode("auth-code"))
stubTokenRequests(
tokenResponse("code-access", jwt("email" to "", "preferred_username" to "alt@example.com"), "rt"),
tokenResponse("outlook-access", null, "rt"),
)
assertEquals("alt@example.com", authManager().exchangeToken(mockk<Intent>()).email)
@@ -10,6 +10,7 @@ import io.mockk.every
import io.mockk.just
import io.mockk.mockk
import io.mockk.mockkObject
import io.mockk.mockkStatic
import io.mockk.unmockkAll
import io.mockk.verify
import kotlinx.coroutines.CompletableDeferred
@@ -28,7 +29,9 @@ import org.libremail.data.settings.SettingsRepository
import java.io.File
import java.util.concurrent.Executors
import kotlin.test.assertEquals
import kotlin.test.assertFailsWith
import kotlin.test.assertFalse
import kotlin.test.assertTrue
/**
* [DatabaseProvisioner] holds the one-time startup sequence that `DatabaseModule.provideDatabase` used
@@ -56,6 +59,11 @@ class DatabaseProvisionerTest {
.asCoroutineDispatcher()
mockkObject(DatabaseEncryption)
mockkObject(DatabaseFiles)
// `android.util.Log` is a no-op stub under plain JVM unit tests; the #359 degrade path breadcrumbs
// through AppLog.w, so statically mock Log (fully-qualified — a raw android.util.Log import is
// detekt-forbidden, epic #324) so it does not throw "not mocked".
mockkStatic(android.util.Log::class)
every { android.util.Log.w(any<String>(), any<String>(), any()) } returns 0
every { context.getDatabasePath(any()) } returns File("libremail.db")
every { DatabaseFiles.clear(any()) } just Runs
@@ -98,6 +106,62 @@ class DatabaseProvisionerTest {
verify(exactly = 1) { DatabaseEncryption.ensureNativeLibraryLoaded() }
}
@Test
fun `a native-library load failure fails closed without opening plaintext or touching the setting`() = runTest {
every { settingsRepository.settings } returns flowOf(AppSettings(encryptCache = true, appLock = false))
// Issue #359: the SQLCipher `.so` fails to load, throwing UnsatisfiedLinkError (a LinkageError) at
// the keyed open. The provisioner must FAIL CLOSED — raise a distinct signal, never silently
// degrade to an unencrypted cache (which would defeat the user's opt-in encryption).
val nativeLoadFailure = UnsatisfiedLinkError("dlopen failed: libsqlcipher.so is not 16 KB aligned")
every { DatabaseEncryption.ensureNativeLibraryLoaded() } throws nativeLoadFailure
val error = assertFailsWith<CacheEncryptionUnavailableException> { provisioner().prepareCache() }
// The distinct signal preserves the underlying native LinkageError in its cause chain (coroutine
// stack-trace recovery may re-wrap the exception across withContext, so assert the chain rather
// than exact instance identity), and NONE of the old fail-open side effects run: the setting is
// never flipped, nothing is wiped, and no seal is reset.
assertTrue(
generateSequence(error.cause) { it.cause }.any { it is LinkageError },
"the native-load LinkageError must be preserved as the cause",
)
coVerify(exactly = 0) { settingsRepository.setEncryptCache(any()) }
verify(exactly = 0) { DatabaseFiles.clear(any()) }
coVerify(exactly = 0) { keyStore.resetSealedPassphrase() }
}
@Test
fun `a native-library load failure never wipes an already-encrypted cache`() = runTest {
every { settingsRepository.settings } returns flowOf(AppSettings(encryptCache = true, appLock = false))
every { DatabaseEncryption.isEncrypted(any()) } returns true
every { DatabaseEncryption.ensureEncrypted(any(), any()) } throws
UnsatisfiedLinkError("dlopen failed: libsqlcipher.so is not 16 KB aligned")
assertFailsWith<CacheEncryptionUnavailableException> { provisioner().prepareCache() }
// The ciphertext the plaintext opener can't parse is deliberately PRESERVED (it may become
// readable again once the library loads on a later launch), the seals stay intact, and the
// setting is untouched — the opposite of the rejected degrade-and-wipe behaviour.
verify(exactly = 0) { DatabaseFiles.clear(any()) }
coVerify(exactly = 0) { keyStore.resetSealedPassphrase() }
coVerify(exactly = 0) { settingsRepository.setEncryptCache(any()) }
}
@Test
fun `a native-library load failure is not memoized and retries on the next open`() = runTest {
every { settingsRepository.settings } returns flowOf(AppSettings(encryptCache = true, appLock = false))
every { DatabaseEncryption.ensureNativeLibraryLoaded() } throws
UnsatisfiedLinkError("dlopen failed: libsqlcipher.so is not 16 KB aligned")
val provisioner = provisioner()
assertFailsWith<CacheEncryptionUnavailableException> { provisioner.prepareCache() }
assertFailsWith<CacheEncryptionUnavailableException> { provisioner.prepareCache() }
// A failure is NOT memoized (unlike a success): each open re-runs the whole sequence, so the
// migrator ran on BOTH attempts. That is what lets a later launch recover once the library loads.
coVerify(exactly = 2) { accountDataMigrator.migrateIfNeeded() }
}
@Test
fun `prepareCache suspends on the auth-bound passphrase until it resolves`() = runTest {
every { settingsRepository.settings } returns flowOf(AppSettings(encryptCache = true, appLock = true))
@@ -4,6 +4,7 @@ package org.libremail.data.repository
import android.content.ContentResolver
import android.content.Context
import android.net.Uri
import android.util.Log
import app.cash.turbine.test
import io.mockk.Runs
import io.mockk.coEvery
@@ -19,6 +20,7 @@ import jakarta.mail.Flags
import kotlinx.coroutines.flow.flowOf
import kotlinx.coroutines.test.runTest
import org.junit.After
import org.junit.Before
import org.junit.Test
import org.libremail.data.attachment.AttachmentUriGrants
import org.libremail.data.local.dao.AccountDao
@@ -97,6 +99,17 @@ class MailRepositoryImplCoverageTest {
attachmentUriGrants = mockk<AttachmentUriGrants>(relaxed = true),
)
// openMessage now breadcrumbs via AppLog (issue #358); android.util.Log is a no-op stub under plain
// JVM tests, so mock it class-wide so no test crashes on the unmocked method.
@Before
fun setUp() {
mockkStatic(Log::class)
every { Log.d(any(), any()) } returns 0
every { Log.i(any(), any()) } returns 0
every { Log.w(any<String>(), any<String>()) } returns 0
every { Log.e(any(), any(), any()) } returns 0
}
@After
fun tearDown() = unmockkAll()
@@ -2,6 +2,7 @@
package org.libremail.data.repository
import android.content.Context
import android.util.Log
import androidx.paging.PagingSource
import androidx.paging.PagingState
import androidx.paging.testing.asSnapshot
@@ -12,13 +13,17 @@ import io.mockk.coVerify
import io.mockk.every
import io.mockk.just
import io.mockk.mockk
import io.mockk.mockkStatic
import io.mockk.slot
import io.mockk.unmockkAll
import jakarta.mail.Flags
import kotlinx.coroutines.CompletableDeferred
import kotlinx.coroutines.flow.flowOf
import kotlinx.coroutines.runBlocking
import kotlinx.coroutines.test.runTest
import kotlinx.coroutines.withTimeout
import org.junit.After
import org.junit.Before
import org.junit.Test
import org.libremail.data.local.dao.AccountDao
import org.libremail.data.local.dao.AttachmentDao
@@ -47,6 +52,9 @@ import org.libremail.mail.FetchedFolder
import org.libremail.mail.ImapClient
import org.libremail.mail.MessageContent
import org.libremail.mail.ReplyContext
import org.libremail.reporting.AppLog
import org.libremail.reporting.RingLogBuffer
import org.libremail.reporting.accountLogRef
import java.io.File
import java.nio.file.Files
import kotlin.test.assertEquals
@@ -83,6 +91,20 @@ class MailRepositoryImplTest {
attachmentUriGrants = mockk(relaxed = true),
)
// openMessage now breadcrumbs via AppLog (issue #358); android.util.Log is a no-op stub under plain
// JVM tests, so mock it class-wide so no test crashes on the unmocked method.
@Before
fun setUp() {
mockkStatic(Log::class)
every { Log.d(any(), any()) } returns 0
every { Log.i(any(), any()) } returns 0
every { Log.w(any<String>(), any<String>()) } returns 0
every { Log.e(any(), any(), any()) } returns 0
}
@After
fun tearDown() = unmockkAll()
@Test
fun `pagedUnifiedFolderMessages maps the paged summaries to domain messages`() = runTest {
every { messageDao.pagingUnifiedFolderSummaries("INBOX") } returns FakeSummaryPagingSource(
@@ -214,6 +236,30 @@ class MailRepositoryImplTest {
coVerify { imapClient.fetchBodyMarkingSeen(any(), "Archive", "5") }
}
@Test
fun `openMessage breadcrumbs a PII-free open via the account ref and folder label`() = runTest {
val buffer = RingLogBuffer()
AppLog.install(buffer)
val id = "acct:INBOX:7"
coEvery { messageDao.getRouting(id) } returns messageRouting(id, "INBOX")
coEvery { messageDao.getById(id) } returns messageEntity(id, "INBOX")
coEvery { accountDao.getById("acct") } returns accountEntity()
coEvery { connectionFactory.imapParamsFor(any()) } returns imapParams()
coEvery { imapClient.fetchBodyMarkingSeen(any(), "INBOX", "7") } returns
MessageContent("Body text", isHtml = false)
coEvery { messageDao.updateBody(id, any(), any(), any()) } just Runs
coEvery { messageDao.setRead(id, true) } just Runs
repository.openMessage(id)
val breadcrumb = buffer.snapshot().map { it.message }.single { it.startsWith("openMessage ") }
// Account via the hashed accountLogRef (never the raw id), system folder by name, branch + timing.
assertTrue(breadcrumb.contains(accountLogRef("acct")), breadcrumb)
assertTrue(breadcrumb.contains("folder=INBOX"), breadcrumb)
assertTrue(breadcrumb.contains("fetchedBody=true"), breadcrumb)
assertTrue(breadcrumb.contains("took="), breadcrumb)
}
@Test
fun `openMessage derives a readable plain-text snippet from an HTML body`() = runTest {
val id = "acct:INBOX:20"
@@ -2,6 +2,11 @@
package org.libremail.data.security
import android.security.keystore.KeyGenParameterSpec
import android.security.keystore.StrongBoxUnavailableException
import io.mockk.every
import io.mockk.mockk
import io.mockk.mockkStatic
import io.mockk.unmockkStatic
import org.junit.Test
import java.security.GeneralSecurityException
import javax.crypto.AEADBadTagException
@@ -70,6 +75,27 @@ class AesGcmKeystoreCipherTest {
assertTrue(error.message!!.contains("test.alias"))
}
@Test
fun `key generation falls back to a TEE-backed key when StrongBox is unavailable`() {
// AppLog.i (breadcrumb on the fallback) forwards to android.util.Log, a no-op stub under plain JVM
// tests; mock it (fully-qualified — a raw android.util.Log import is detekt-forbidden, epic #324).
mockkStatic(android.util.Log::class)
every { android.util.Log.i(any<String>(), any<String>()) } returns 0
try {
val teeKey = newAesKey()
val cipher = StrongBoxFakeCipher(teeKey)
val key = cipher.createKey()
// StrongBox is attempted first; its StrongBoxUnavailableException triggers a single retry with
// strongBox = false, and that TEE-backed key is returned — so generation succeeds everywhere.
assertSame(teeKey, key)
assertEquals(listOf(true, false), cipher.attempts)
} finally {
unmockkStatic(android.util.Log::class)
}
}
@Test
fun `a non-AEAD failure propagates unwrapped so key invalidation still surfaces`() {
// Only AEADBadTagException is remapped; every other cipher failure — including the
@@ -110,7 +136,36 @@ class AesGcmKeystoreCipherTest {
return onDecrypt(key, encoded)
}
override fun keySpec(): KeyGenParameterSpec = error("keySpec is not exercised in the JVM base test")
override fun keySpec(strongBox: Boolean): KeyGenParameterSpec =
error("keySpec is not exercised in the JVM base test")
}
/**
* A JVM-only cipher that exercises the REAL [getOrCreateKey]/StrongBox-fallback control flow (unlike
* [FakeCipher], which stubs [getOrCreateKey] out). [existingKey] returns null so a key is generated,
* and the [generateKey] seam simulates the device: a StrongBox attempt fails, the TEE attempt yields
* [teeKey]. [attempts] records the `strongBox` value of each generation attempt in order.
*/
private class StrongBoxFakeCipher(private val teeKey: SecretKey) :
AesGcmKeystoreCipher(alias = "test.alias", generateKeyOnDecrypt = true) {
val attempts = mutableListOf<Boolean>()
override fun existingKey(): SecretKey? = null
override fun generateKey(strongBox: Boolean): SecretKey {
attempts += strongBox
// Objenesis-instantiated (no stubbed-constructor call) so it is throwable under the android.jar
// stub; it is still a StrongBoxUnavailableException, so the production catch clause matches.
if (strongBox) throw mockk<StrongBoxUnavailableException>(relaxed = true)
return teeKey
}
override fun keySpec(strongBox: Boolean): KeyGenParameterSpec =
error("keySpec is bypassed because generateKey is overridden")
/** Invokes the protected [getOrCreateKey] so the StrongBox fallback runs without touching Base64. */
fun createKey(): SecretKey = getOrCreateKey()
}
private companion object {
@@ -0,0 +1,81 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.data.security
import io.mockk.every
import io.mockk.mockk
import io.mockk.mockkStatic
import io.mockk.unmockkAll
import kotlinx.coroutines.flow.flowOf
import kotlinx.coroutines.test.runTest
import org.junit.After
import org.junit.Before
import org.junit.Test
import org.libremail.data.settings.AppSettings
import org.libremail.data.settings.SettingsRepository
import kotlin.test.assertEquals
import kotlin.test.assertFalse
import kotlin.test.assertTrue
/**
* [KeystoreReportEncryption] delegates the crypto to [KeystoreCrypto] (device-bound, tested in
* `KeystoreCryptoTest`) and answers [org.libremail.reporting.ReportEncryption.enabled] from an
* in-memory mirror of the `encryptCache` setting. These pin the delegation and that the mirror tracks
* the observed setting — including the crash-safe default of "off" before the first value lands.
*/
class KeystoreReportEncryptionTest {
private val crypto = mockk<KeystoreCrypto>()
private val settingsRepository = mockk<SettingsRepository>()
private val encryption = KeystoreReportEncryption(crypto, settingsRepository)
@Before
fun setUp() {
// observeEncryptCacheSetting breadcrumbs the on/off state through AppLog -> android.util.Log,
// a no-op stub that throws "not mocked" under plain JVM tests. Fully-qualified (a raw
// android.util.Log import is detekt-forbidden, epic #324).
mockkStatic(android.util.Log::class)
every { android.util.Log.i(any<String>(), any<String>()) } returns 0
}
@After
fun tearDown() = unmockkAll()
@Test
fun `is disabled before the setting has been observed`() {
assertFalse(encryption.enabled())
}
@Test
fun `encrypt delegates to the keystore crypto`() {
every { crypto.encrypt("report-json") } returns "cipher"
assertEquals("cipher", encryption.encrypt("report-json"))
}
@Test
fun `decrypt delegates to the keystore crypto`() {
every { crypto.decrypt("cipher") } returns "report-json"
assertEquals("report-json", encryption.decrypt("cipher"))
}
@Test
fun `enabled mirrors the latest encryptCache setting when on`() = runTest {
every { settingsRepository.settings } returns
flowOf(AppSettings(encryptCache = false), AppSettings(encryptCache = true))
encryption.observeEncryptCacheSetting()
assertTrue(encryption.enabled())
}
@Test
fun `enabled mirrors the latest encryptCache setting when off`() = runTest {
every { settingsRepository.settings } returns
flowOf(AppSettings(encryptCache = true), AppSettings(encryptCache = false))
encryption.observeEncryptCacheSetting()
assertFalse(encryption.enabled())
}
}
@@ -1,31 +1,62 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.data.sync
import android.util.Log
import androidx.work.ListenableWorker.Result
import dagger.Lazy
import io.mockk.coEvery
import io.mockk.coVerify
import io.mockk.every
import io.mockk.mockk
import io.mockk.mockkStatic
import io.mockk.unmockkAll
import io.mockk.verify
import kotlinx.coroutines.test.runTest
import org.junit.After
import org.junit.Before
import org.junit.Test
import org.libremail.data.security.EncryptedCacheGuard
import org.libremail.reporting.AppLog
import org.libremail.reporting.RingLogBuffer
import kotlin.test.assertEquals
import kotlin.test.assertFalse
import kotlin.test.assertTrue
/**
* [BackfillWorker] must defer while the encrypted cache is locked rather than park this WorkManager
* thread opening the DB. It gates on [EncryptedCacheGuard], resolving the (`Lazy`) [MailBackfiller]
* only once unlocked — the same pre-auth invariant `SyncWorker`/`SendWorker` enforce.
* only once unlocked — the same pre-auth invariant `SyncWorker`/`SendWorker` enforce. Also covers issue
* #329's AppLog breadcrumbs on the deferred/success/retry outcomes.
*/
class BackfillWorkerTest {
private val backfiller = mockk<MailBackfiller>()
private val lazyBackfiller = mockk<Lazy<MailBackfiller>> { every { get() } returns backfiller }
private val cacheGuard = mockk<EncryptedCacheGuard>()
private val logBuffer = RingLogBuffer()
private fun worker() = BackfillWorker(mockk(relaxed = true), mockk(relaxed = true), lazyBackfiller, cacheGuard)
@Before
fun setUp() {
// `android.util.Log` is a no-op stub under plain JVM unit tests, so it is statically mocked here,
// mirroring org.libremail.reporting.AppLogTest — doWork() now breadcrumbs through AppLog.
mockkStatic(Log::class)
every { Log.d(any(), any()) } returns 0
every { Log.d(any(), any(), any()) } returns 0
every { Log.i(any(), any()) } returns 0
every { Log.w(any<String>(), any<String>()) } returns 0
every { Log.w(any<String>(), any<String>(), any()) } returns 0
every { Log.e(any(), any(), any()) } returns 0
AppLog.install(logBuffer)
}
@After
fun tearDown() {
unmockkAll()
DebugFetchGate.reset() // the gate is a process-global object; don't leak a pause to other tests
}
@Test
fun `retries without resolving the backfiller when the cache is locked`() = runTest {
coEvery { cacheGuard.isCacheLocked() } returns true
@@ -54,4 +85,84 @@ class BackfillWorkerTest {
assertEquals(Result.retry(), worker().doWork())
}
// --- issue #329: AppLog breadcrumbs ---------------------------------------------------------
@Test
fun `logs a deferred breadcrumb when the cache is locked, without touching the backfiller`() = runTest {
coEvery { cacheGuard.isCacheLocked() } returns true
worker().doWork()
val entry = logBuffer.snapshot().single()
assertEquals('I', entry.level)
assertEquals("backfill deferred: cache locked", entry.message)
}
@Test
fun `logs a success breadcrumb once every chained slice completes`() = runTest {
coEvery { cacheGuard.isCacheLocked() } returns false
coEvery { backfiller.runBackfill(any()) } returnsMany listOf(true, false)
worker().doWork()
val entry = logBuffer.snapshot().single()
assertEquals('I', entry.level)
assertEquals("backfill worker: success", entry.message)
}
@Test
fun `logs a scrubbed retry breadcrumb when backfilling throws`() = runTest {
coEvery { cacheGuard.isCacheLocked() } returns false
coEvery { backfiller.runBackfill(any()) } throws
IllegalStateException("auth failed for a@example.org")
worker().doWork()
val entry = logBuffer.snapshot().single()
assertEquals('W', entry.level)
assertTrue(entry.message.startsWith("backfill worker: retry"), entry.message)
assertTrue(entry.message.contains("IllegalStateException"), entry.message)
assertFalse(entry.message.contains("a@example.org"), entry.message)
}
// --- issue #393: debug-only fetch gate ------------------------------------------------------
@Test
fun `defers without resolving the backfiller when the fetch gate pauses backfill`() = runTest {
// Cache unlocked, so the ONLY reason to defer is the gate. (BuildConfig.DEBUG is true under
// testDebugUnitTest, so the gate branch is live.)
coEvery { cacheGuard.isCacheLocked() } returns false
DebugFetchGate.pause(setOf(FetchScope.BACKFILL))
assertEquals(Result.retry(), worker().doWork())
// Same invariant as the cache-lock deferral: never resolve the DB-backed collaborator.
verify(exactly = 0) { lazyBackfiller.get() }
coVerify(exactly = 0) { backfiller.runBackfill(any()) }
}
@Test
fun `pausing only prefetch does NOT defer the backfill worker`() = runTest {
// The worker gate honours BACKFILL only — a PREFETCH pause must leave history paging running.
coEvery { cacheGuard.isCacheLocked() } returns false
coEvery { backfiller.runBackfill(any()) } returns false
DebugFetchGate.pause(setOf(FetchScope.PREFETCH))
assertEquals(Result.success(), worker().doWork())
coVerify(exactly = 1) { backfiller.runBackfill(any()) }
}
@Test
fun `logs a deferred breadcrumb when the fetch gate pauses backfill`() = runTest {
coEvery { cacheGuard.isCacheLocked() } returns false
DebugFetchGate.pause(setOf(FetchScope.BACKFILL))
worker().doWork()
val entry = logBuffer.snapshot().single()
assertEquals('I', entry.level)
assertEquals("backfill deferred: fetch-gate paused", entry.message)
}
}
@@ -0,0 +1,164 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.data.sync
import org.junit.After
import org.junit.Before
import org.junit.Test
import java.util.concurrent.CountDownLatch
import java.util.concurrent.Executors
import java.util.concurrent.TimeUnit
import kotlin.test.assertEquals
import kotlin.test.assertFalse
import kotlin.test.assertTrue
/**
* Unit tests for the debug-only fetch gate (issue #393): its default not-paused state, pause/resume/query
* per [FetchScope], the `all` alias + scope-string parsing the receiver relies on, the ordered-broadcast
* read-back string, and thread-safety of the in-memory holder. [DebugFetchGate] is a process-global
* object, so each test resets it to isolate from the others.
*/
class DebugFetchGateTest {
@Before
@After
fun resetGate() = DebugFetchGate.reset()
@Test
fun `defaults to nothing paused for every scope`() {
FetchScope.entries.forEach { scope ->
assertFalse(DebugFetchGate.isPaused(scope), "$scope must default to not-paused")
}
assertTrue(DebugFetchGate.pausedScopes().isEmpty())
assertEquals("paused=[]", DebugFetchGate.pausedResult())
}
@Test
fun `pausing one scope leaves the other live`() {
DebugFetchGate.pause(setOf(FetchScope.BACKFILL))
assertTrue(DebugFetchGate.isPaused(FetchScope.BACKFILL))
assertFalse(DebugFetchGate.isPaused(FetchScope.PREFETCH))
assertEquals(setOf(FetchScope.BACKFILL), DebugFetchGate.pausedScopes())
assertEquals("paused=[backfill]", DebugFetchGate.pausedResult())
}
@Test
fun `pausing both scopes reports both, in declaration order`() {
DebugFetchGate.pause(setOf(FetchScope.PREFETCH, FetchScope.BACKFILL))
assertTrue(DebugFetchGate.isPaused(FetchScope.BACKFILL))
assertTrue(DebugFetchGate.isPaused(FetchScope.PREFETCH))
// Declaration order (BACKFILL before PREFETCH), regardless of the set's insertion order.
assertEquals("paused=[backfill,prefetch]", DebugFetchGate.pausedResult())
}
@Test
fun `pause is additive across calls`() {
DebugFetchGate.pause(setOf(FetchScope.BACKFILL))
DebugFetchGate.pause(setOf(FetchScope.PREFETCH))
assertEquals(setOf(FetchScope.BACKFILL, FetchScope.PREFETCH), DebugFetchGate.pausedScopes())
}
@Test
fun `resume clears only the named scope`() {
DebugFetchGate.pause(setOf(FetchScope.BACKFILL, FetchScope.PREFETCH))
DebugFetchGate.resume(setOf(FetchScope.BACKFILL))
assertFalse(DebugFetchGate.isPaused(FetchScope.BACKFILL))
assertTrue(DebugFetchGate.isPaused(FetchScope.PREFETCH))
assertEquals("paused=[prefetch]", DebugFetchGate.pausedResult())
}
@Test
fun `resume of an unpaused scope is a no-op`() {
DebugFetchGate.resume(setOf(FetchScope.BACKFILL))
assertEquals("paused=[]", DebugFetchGate.pausedResult())
}
@Test
fun `empty pause and resume are no-ops`() {
DebugFetchGate.pause(emptySet())
assertEquals("paused=[]", DebugFetchGate.pausedResult())
DebugFetchGate.pause(setOf(FetchScope.PREFETCH))
DebugFetchGate.resume(emptySet())
assertEquals("paused=[prefetch]", DebugFetchGate.pausedResult())
}
@Test
fun `reset clears every pause`() {
DebugFetchGate.pause(setOf(FetchScope.BACKFILL, FetchScope.PREFETCH))
DebugFetchGate.reset()
assertTrue(DebugFetchGate.pausedScopes().isEmpty())
}
// --- FetchScope.parse (the scope-string contract the receiver depends on) -------------------------
@Test
fun `parse maps a comma list of known scopes`() {
assertEquals(setOf(FetchScope.BACKFILL, FetchScope.PREFETCH), FetchScope.parse("backfill,prefetch"))
assertEquals(setOf(FetchScope.BACKFILL), FetchScope.parse("backfill"))
}
@Test
fun `parse expands the all alias to every scope`() {
assertEquals(FetchScope.entries.toSet(), FetchScope.parse("all"))
// The alias wins even mixed with other tokens.
assertEquals(FetchScope.entries.toSet(), FetchScope.parse("backfill,all"))
}
@Test
fun `parse is case- and whitespace-insensitive`() {
assertEquals(setOf(FetchScope.BACKFILL, FetchScope.PREFETCH), FetchScope.parse(" BACKFILL , Prefetch "))
}
@Test
fun `parse ignores unknown and blank tokens`() {
assertEquals(setOf(FetchScope.PREFETCH), FetchScope.parse("prefetch,bogus,,"))
assertTrue(FetchScope.parse("nope").isEmpty())
}
@Test
fun `parse of null or blank is the empty set`() {
assertTrue(FetchScope.parse(null).isEmpty())
assertTrue(FetchScope.parse("").isEmpty())
assertTrue(FetchScope.parse(" ").isEmpty())
}
// --- thread-safety --------------------------------------------------------------------------------
@Test
fun `concurrent pause and resume never corrupts the holder`() {
val threads = 8
val iterations = 2_000
val pool = Executors.newFixedThreadPool(threads)
val start = CountDownLatch(1)
try {
val futures = (0 until threads).map { index ->
pool.submit {
start.await()
val scope = FetchScope.entries[index % FetchScope.entries.size]
repeat(iterations) {
DebugFetchGate.pause(setOf(scope))
DebugFetchGate.isPaused(scope)
DebugFetchGate.pausedResult()
DebugFetchGate.resume(setOf(scope))
}
}
}
start.countDown()
futures.forEach { it.get(30, TimeUnit.SECONDS) }
} finally {
pool.shutdownNow()
}
// No exception above (a data race on a plain HashSet would throw), and the holder is consistent:
// pausedScopes() only ever reports declared scopes.
assertTrue(DebugFetchGate.pausedScopes().all { it in FetchScope.entries })
}
}
@@ -2,12 +2,15 @@
package org.libremail.data.sync
import android.content.Context
import android.util.Log
import com.icegreen.greenmail.util.GreenMail
import com.icegreen.greenmail.util.ServerSetupTest
import io.mockk.coEvery
import io.mockk.coVerify
import io.mockk.every
import io.mockk.mockk
import io.mockk.mockkStatic
import io.mockk.unmockkAll
import jakarta.mail.Folder
import jakarta.mail.Message
import jakarta.mail.Session
@@ -38,6 +41,8 @@ import org.libremail.mail.FetchedMessage
import org.libremail.mail.ImapClient
import org.libremail.power.BatteryStatus
import org.libremail.power.BatteryStatusProvider
import org.libremail.reporting.AppLog
import org.libremail.reporting.RingLogBuffer
import java.util.Date
import java.util.Properties
import kotlin.test.assertEquals
@@ -55,7 +60,9 @@ import kotlin.test.assertTrue
class MailBackfillerTest {
private lateinit var greenMail: GreenMail
private val client = ImapClient()
// Reuse off: this suite pins the connect-per-operation backfill behaviour it was written against.
private val client = ImapClient(reuseConnections = false)
private val accountEntity = AccountEntity(
id = "acct",
@@ -74,15 +81,34 @@ class MailBackfillerTest {
// though `cached` would silently absorb it. Isolates the "no message fetched twice" guarantee.
private var totalOffered = 0
// issue #329: AppLog breadcrumbs — a fresh buffer per test (a new instance per @Test, per JUnit4).
private val logBuffer = RingLogBuffer()
@Before
fun setUp() {
greenMail = GreenMail(ServerSetupTest.SMTP_IMAP)
greenMail.start()
greenMail.setUser("alice@example.org", "secret")
// `android.util.Log` is a no-op stub under plain JVM unit tests, so it is statically mocked here
// for the whole class, mirroring org.libremail.reporting.AppLogTest — every test now exercises
// real backfill code, which breadcrumbs through AppLog.
mockkStatic(Log::class)
every { Log.d(any(), any()) } returns 0
every { Log.d(any(), any(), any()) } returns 0
every { Log.i(any(), any()) } returns 0
every { Log.w(any<String>(), any<String>()) } returns 0
every { Log.w(any<String>(), any<String>(), any()) } returns 0
every { Log.e(any(), any(), any()) } returns 0
AppLog.install(logBuffer)
}
@After
fun tearDown() = greenMail.stop()
fun tearDown() {
greenMail.stop()
unmockkAll()
DebugFetchGate.reset() // the gate is a process-global object; don't leak a pause to other tests
}
private fun params() = ImapConnectionParams(
host = "127.0.0.1",
@@ -209,6 +235,23 @@ class MailBackfillerTest {
coVerify(exactly = 0) { requireNotNull(lastMailRepository).prefetchMessage(any()) }
}
// --- issue #393: debug-only fetch gate ------------------------------------------------------
@Test
fun `a paused prefetch gate skips body prefetch but still pages history headers`() = runTest {
appendMessages(60)
seedForegroundWindow()
DebugFetchGate.pause(setOf(FetchScope.PREFETCH))
// ALWAYS + healthy battery: without the gate this WOULD prefetch, so the gate is the only cause.
backfiller(AccountSettings("acct"), fetchPolicy = FetchPolicy.ALWAYS).runBackfill()
// The gate stops only the body/attachment prefetch; the history headers still page in.
assertEquals(60, distinctCachedUids().size, "header paging itself is not gated")
coVerify(exactly = 0) { requireNotNull(lastMailRepository).prefetchMessage(any()) }
assertTrue(logBuffer.snapshot().any { it.message == "prefetch skipped: fetch-gate paused" })
}
@Test
fun `charging at low percent keeps the backfill content prefetch running`() = runTest {
appendMessages(60)
@@ -353,6 +396,68 @@ class MailBackfillerTest {
}
}
// --- issue #329: AppLog breadcrumbs ---------------------------------------------------------
@Test
fun `runBackfill logs a start breadcrumb, a per-folder breadcrumb, and a done breadcrumb`() = runTest {
appendMessages(TOTAL)
seedForegroundWindow()
backfiller(AccountSettings("acct")).runBackfill(maxBatches = 3)
val snapshot = logBuffer.snapshot()
val start = snapshot.first { it.message.startsWith("backfill slice:") }
val perFolder = snapshot.first { it.message.startsWith("backfill acct:") }
val done = snapshot.first { it.message.startsWith("backfill slice done:") }
assertEquals('I', start.level)
assertEquals("backfill slice: maxBatches=3", start.message)
assertEquals('D', perFolder.level)
assertTrue(perFolder.message.contains("folder=INBOX"), perFolder.message)
assertTrue(perFolder.message.contains("pages="), perFolder.message)
assertTrue(perFolder.message.contains("complete="), perFolder.message)
assertEquals('I', done.level)
assertTrue(done.message.contains("moreWork="), done.message)
snapshot.forEach { assertNoPii(it.message) }
}
@Test
fun `an already-complete folder logs no redundant per-folder breadcrumb on the next slice`() = runTest {
appendMessages(60)
seedForegroundWindow()
val backfiller = backfiller(AccountSettings("acct"))
var guard = 0
while (backfiller.runBackfill() && guard++ < 10) { /* drive to completion */ }
logBuffer.clear()
// The folder is already complete; this slice does no work for it.
backfiller.runBackfill()
assertTrue(
logBuffer.snapshot().none { it.message.startsWith("backfill acct:") },
"an already-complete folder must not spam a breadcrumb on every subsequent slice",
)
}
@Test
fun `no slice of a full backfill ever logs the account's email, host, or a message address`() = runTest {
appendMessages(TOTAL)
seedForegroundWindow()
val backfiller = backfiller(AccountSettings("acct"))
var guard = 0
while (backfiller.runBackfill() && guard++ < 10) { /* drive to completion */ }
val snapshot = logBuffer.snapshot()
assertTrue(snapshot.isNotEmpty(), "the run must have logged something to be a meaningful check")
snapshot.forEach { assertNoPii(it.message) }
}
/** No test fixture's email address or host may ever reach a log line — the hard PII rule. */
private fun assertNoPii(message: String) {
assertFalse(message.contains("@example.org"), message) // covers the account and every sender
assertFalse(message.contains("127.0.0.1"), message) // the GreenMail host
}
private fun fetchedMessage(uid: String) = FetchedMessage(
uid = uid,
sender = "Sender",
@@ -1,9 +1,12 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.data.sync
import android.util.Log
import io.mockk.coEvery
import io.mockk.every
import io.mockk.mockk
import io.mockk.mockkStatic
import io.mockk.unmockkAll
import kotlinx.coroutines.CompletableDeferred
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.delay
@@ -13,6 +16,8 @@ import kotlinx.coroutines.launch
import kotlinx.coroutines.runBlocking
import kotlinx.coroutines.sync.withLock
import kotlinx.coroutines.yield
import org.junit.After
import org.junit.Before
import org.junit.Test
import org.libremail.data.local.dao.AccountDao
import org.libremail.data.local.dao.MessageDao
@@ -28,6 +33,8 @@ import org.libremail.mail.FetchedMessage
import org.libremail.mail.ImapClient
import org.libremail.power.BatteryStatus
import org.libremail.power.BatteryStatusProvider
import org.libremail.reporting.AppLog
import org.libremail.reporting.RingLogBuffer
import java.util.Collections
import java.util.concurrent.atomic.AtomicInteger
import kotlin.test.assertEquals
@@ -52,6 +59,25 @@ class MailMaintenanceGateTest {
smtp = ServerConfigEmbedded("127.0.0.1", 465, "NONE"),
)
// issue #329: MailBackfiller/MailPruner now breadcrumb through AppLog, whose Logcat forwarding
// (`android.util.Log`) is a no-op stub under plain JVM unit tests — mock it statically, mirroring
// org.libremail.reporting.AppLogTest. The breadcrumbs themselves are asserted in
// MailBackfillerTest/MailPrunerTest; this class only needs to not crash.
@Before
fun setUp() {
mockkStatic(Log::class)
every { Log.d(any(), any()) } returns 0
every { Log.d(any(), any(), any()) } returns 0
every { Log.i(any(), any()) } returns 0
every { Log.w(any<String>(), any<String>()) } returns 0
every { Log.w(any<String>(), any<String>(), any()) } returns 0
every { Log.e(any(), any(), any()) } returns 0
AppLog.install(RingLogBuffer())
}
@After
fun tearDown() = unmockkAll()
/**
* The gate's contract: everyone who acquires the *same* gate takes the *same* exclusive lock, so no
* two critical sections overlap. Guards against a regression that hands out a fresh lock per access
@@ -2,15 +2,20 @@
package org.libremail.data.sync
import android.content.Context
import android.util.Log
import io.mockk.Runs
import io.mockk.coEvery
import io.mockk.coVerify
import io.mockk.every
import io.mockk.just
import io.mockk.mockk
import io.mockk.mockkStatic
import io.mockk.slot
import io.mockk.unmockkAll
import kotlinx.coroutines.flow.flowOf
import kotlinx.coroutines.test.runTest
import org.junit.After
import org.junit.Before
import org.junit.Test
import org.libremail.data.local.dao.AccountDao
import org.libremail.data.local.dao.MessageDao
@@ -20,8 +25,11 @@ import org.libremail.data.settings.AccountSettingsRepository
import org.libremail.data.settings.AppSettings
import org.libremail.data.settings.SettingsRepository
import org.libremail.domain.model.AccountSettings
import org.libremail.reporting.AppLog
import org.libremail.reporting.RingLogBuffer
import java.io.File
import kotlin.test.assertEquals
import kotlin.test.assertFalse
import kotlin.test.assertTrue
/**
@@ -40,6 +48,27 @@ class MailPrunerTest {
smtp = ServerConfigEmbedded("smtp.example.org", 465, "SSL_TLS"),
)
// issue #329: AppLog breadcrumbs — a fresh buffer per test (a new instance per @Test, per JUnit4).
private val logBuffer = RingLogBuffer()
@Before
fun setUp() {
// `android.util.Log` is a no-op stub under plain JVM unit tests, so it is statically mocked here
// for the whole class, mirroring org.libremail.reporting.AppLogTest — every test now exercises
// real prune code, which breadcrumbs through AppLog.
mockkStatic(Log::class)
every { Log.d(any(), any()) } returns 0
every { Log.d(any(), any(), any()) } returns 0
every { Log.i(any(), any()) } returns 0
every { Log.w(any<String>(), any<String>()) } returns 0
every { Log.w(any<String>(), any<String>(), any()) } returns 0
every { Log.e(any(), any(), any()) } returns 0
AppLog.install(logBuffer)
}
@After
fun tearDown() = unmockkAll()
private fun pruner(global: AppSettings, accountSettings: AccountSettings, messageDao: MessageDao): MailPruner {
val accountDao = mockk<AccountDao>()
coEvery { accountDao.getAll() } returns listOf(account)
@@ -136,4 +165,61 @@ class MailPrunerTest {
coVerify(exactly = 0) { messageDao.syncedIdsOlderThan(any(), any()) }
coVerify(exactly = 0) { messageDao.syncedIdsBeyondCountInFolder(any(), any(), any()) }
}
// --- issue #329: AppLog breadcrumbs ---------------------------------------------------------
@Test
fun `prune logs a done breadcrumb with the removed count`() = runTest {
val messageDao = mockk<MessageDao>(relaxed = true)
coEvery { messageDao.syncedFolders("acct") } returns listOf("INBOX", "Archive")
coEvery { messageDao.syncedIdsBeyondCountInFolder("acct", "INBOX", 2) } returns
listOf("acct:INBOX:1", "acct:INBOX:2")
coEvery { messageDao.syncedIdsBeyondCountInFolder("acct", "Archive", 2) } returns listOf("acct:Archive:9")
coEvery { messageDao.deleteByIds(any()) } just Runs
pruner(
global = AppSettings(retentionCount = 2, retentionMonths = 0),
accountSettings = AccountSettings("acct"),
messageDao = messageDao,
).prune()
val entry = logBuffer.snapshot().single()
assertEquals('I', entry.level)
assertEquals("prune done: removed=3", entry.message)
}
@Test
fun `prune logs a done breadcrumb even when nothing is removed`() = runTest {
val messageDao = mockk<MessageDao>(relaxed = true)
pruner(
global = AppSettings(retentionCount = 0, retentionMonths = 0),
accountSettings = AccountSettings("acct"),
messageDao = messageDao,
).prune()
val entry = logBuffer.snapshot().single()
assertEquals('I', entry.level)
assertEquals("prune done: removed=0", entry.message)
}
@Test
fun `prune never logs the account email or host, even while removing rows by age and count`() = runTest {
val messageDao = mockk<MessageDao>(relaxed = true)
coEvery { messageDao.syncedFolders("acct") } returns listOf("INBOX")
coEvery { messageDao.syncedIdsOlderThan(any(), any()) } returns listOf("acct:INBOX:1")
coEvery { messageDao.syncedIdsBeyondCountInFolder("acct", "INBOX", 2) } returns listOf("acct:INBOX:2")
coEvery { messageDao.deleteByIds(any()) } just Runs
pruner(
global = AppSettings(retentionCount = 2, retentionMonths = 6),
accountSettings = AccountSettings("acct"),
messageDao = messageDao,
).prune()
logBuffer.snapshot().forEach { entry ->
assertFalse(entry.message.contains("a@example.org"), entry.message)
assertFalse(entry.message.contains("example.org"), entry.message)
}
}
}
@@ -2,14 +2,19 @@
package org.libremail.data.sync
import android.content.Context
import android.util.Log
import io.mockk.coEvery
import io.mockk.every
import io.mockk.mockk
import io.mockk.mockkStatic
import io.mockk.unmockkAll
import kotlinx.coroutines.CompletableDeferred
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.flow.flowOf
import kotlinx.coroutines.launch
import kotlinx.coroutines.runBlocking
import org.junit.After
import org.junit.Before
import org.junit.Test
import org.libremail.data.local.dao.AccountDao
import org.libremail.data.local.dao.BackfillProgressDao
@@ -28,6 +33,8 @@ import org.libremail.mail.FetchedMessage
import org.libremail.mail.ImapClient
import org.libremail.power.BatteryStatus
import org.libremail.power.BatteryStatusProvider
import org.libremail.reporting.AppLog
import org.libremail.reporting.RingLogBuffer
import java.io.File
import java.util.concurrent.atomic.AtomicInteger
import kotlin.test.assertEquals
@@ -74,6 +81,25 @@ class MailSyncConcurrencyTest {
useXoauth2 = false,
)
// issue #329: MailSyncer/MailBackfiller/MailPruner now breadcrumb through AppLog, whose Logcat
// forwarding (`android.util.Log`) is a no-op stub under plain JVM unit tests — mock it statically,
// mirroring org.libremail.reporting.AppLogTest. The breadcrumbs themselves are asserted in
// MailSyncerTest/MailBackfillerTest/MailPrunerTest; this class only needs to not crash.
@Before
fun setUp() {
mockkStatic(Log::class)
every { Log.d(any(), any()) } returns 0
every { Log.d(any(), any(), any()) } returns 0
every { Log.i(any(), any()) } returns 0
every { Log.w(any<String>(), any<String>()) } returns 0
every { Log.w(any<String>(), any<String>(), any()) } returns 0
every { Log.e(any(), any(), any()) } returns 0
AppLog.install(RingLogBuffer())
}
@After
fun tearDown() = unmockkAll()
// --- sync ↔ backfill -----------------------------------------------------------------------
/**
@@ -5,12 +5,17 @@ import android.content.Context
import android.net.ConnectivityManager
import android.net.Network
import android.net.NetworkCapabilities
import android.util.Log
import io.mockk.coEvery
import io.mockk.coVerify
import io.mockk.every
import io.mockk.mockk
import io.mockk.mockkStatic
import io.mockk.unmockkAll
import kotlinx.coroutines.flow.flowOf
import kotlinx.coroutines.test.runTest
import org.junit.After
import org.junit.Before
import org.junit.Test
import org.libremail.data.local.dao.AccountDao
import org.libremail.data.local.dao.MessageDao
@@ -28,12 +33,24 @@ import org.libremail.mail.ImapClient
import org.libremail.notifications.MailNotifier
import org.libremail.power.BatteryStatus
import org.libremail.power.BatteryStatusProvider
import org.libremail.reporting.AppLog
import org.libremail.reporting.RingLogBuffer
import java.io.IOException
import kotlin.test.assertEquals
import kotlin.test.assertFalse
import kotlin.test.assertTrue
/**
* Also covers issue #329's net-new [AppLog] breadcrumbs: every test exercises real sync code, so
* `android.util.Log` (a no-op stub under plain JVM unit tests) is statically mocked here for the whole
* class, mirroring [org.libremail.reporting.AppLogTest]. [logBuffer] is fresh per test (a new
* [MailSyncerTest] instance per `@Test`, per JUnit4) and installed before every test so the dedicated
* breadcrumb tests below can assert against it directly.
*/
class MailSyncerTest {
private val logBuffer = RingLogBuffer()
private val account = AccountEntity(
id = "acct",
email = "a@example.org",
@@ -43,6 +60,24 @@ class MailSyncerTest {
smtp = ServerConfigEmbedded("smtp.example.org", 465, "SSL_TLS"),
)
@Before
fun setUp() {
mockkStatic(Log::class)
every { Log.d(any(), any()) } returns 0
every { Log.d(any(), any(), any()) } returns 0
every { Log.i(any(), any()) } returns 0
every { Log.w(any<String>(), any<String>()) } returns 0
every { Log.w(any<String>(), any<String>(), any()) } returns 0
every { Log.e(any(), any(), any()) } returns 0
AppLog.install(logBuffer)
}
@After
fun tearDown() {
unmockkAll()
DebugFetchGate.reset() // the gate is a process-global object; don't leak a pause to other tests
}
/** The IMAP client of the most recently built [syncer], for verifying the fetch window size. */
private lateinit var lastImapClient: ImapClient
@@ -196,6 +231,22 @@ class MailSyncerTest {
coVerify(exactly = 0) { repo.prefetchMessage(any()) }
}
// --- issue #393: debug-only fetch gate ------------------------------------------------------
@Test
fun `a paused prefetch gate skips prefetch while the header sync still runs`() = runTest {
val repo = mockk<MailRepository>(relaxed = true)
DebugFetchGate.pause(setOf(FetchScope.PREFETCH))
// ALWAYS + healthy battery: without the gate this WOULD prefetch, so the gate is the only cause.
val result = syncer(FetchPolicy.ALWAYS, repo).syncFolder("acct", "INBOX")
assertEquals(0, result.getOrNull()) // header sync still ran and succeeded
coVerify { lastImapClient.fetchRecent(any(), "INBOX", any()) } // headers still fetched
coVerify(exactly = 0) { repo.prefetchMessage(any()) }
assertTrue(logBuffer.snapshot().any { it.message == "prefetch skipped: fetch-gate paused" })
}
@Test
fun `at exactly the 20 percent threshold prefetch is paused`() = runTest {
val repo = mockk<MailRepository>(relaxed = true)
@@ -382,4 +433,61 @@ class MailSyncerTest {
every { it.getSystemService(ConnectivityManager::class.java) } returns manager
}
}
// --- issue #329: AppLog breadcrumbs ---------------------------------------------------------
@Test
fun `syncFolder records a debug breadcrumb with the account ref, folder, and fetched count`() = runTest {
val repo = mockk<MailRepository>(relaxed = true)
val fetched = listOf(
FetchedMessage("1", "Ada", "ada@example.org", "Secret subject", 1_000L, isRead = true, isFlagged = false),
)
syncer(FetchPolicy.ON_DEMAND, repo, fetched = fetched).syncFolder("acct", "INBOX")
val entry = logBuffer.snapshot().single()
assertEquals('D', entry.level)
assertTrue(entry.message.contains("folder=INBOX"), entry.message)
assertTrue(entry.message.contains("fetched=1"), entry.message)
assertNoPii(entry.message)
}
@Test
fun `syncAll logs a start breadcrumb and a done breadcrumb with the fetched total`() = runTest {
val syncer = syncAllSyncer(accounts = listOf(accountEntity("one"), accountEntity("two")))
syncer.syncAll()
val snapshot = logBuffer.snapshot()
val start = snapshot.first { it.message.startsWith("sync all:") }
val done = snapshot.first { it.message.startsWith("sync all done:") }
assertEquals('I', start.level)
assertEquals("sync all: 2 accounts", start.message)
assertEquals('I', done.level)
assertEquals("sync all done: fetched=2", done.message)
snapshot.forEach { assertNoPii(it.message) }
}
@Test
fun `syncAll logs a warning breadcrumb with the scrubbed failure when every account fails`() = runTest {
val syncer = syncAllSyncer(
accounts = listOf(accountEntity("bad1"), accountEntity("bad2")),
failingIds = setOf("bad1", "bad2"),
)
syncer.syncAll()
val entry = logBuffer.snapshot().first { it.message.startsWith("sync all failed") }
assertEquals('W', entry.level)
// The throwable's class survives the scrub; its free-text message does not (issue #325).
assertTrue(entry.message.contains("IOException"), entry.message)
assertFalse(entry.message.contains("no network"), entry.message)
logBuffer.snapshot().forEach { assertNoPii(it.message) }
}
/** No test fixture's email address or host may ever reach a log line — the hard PII rule. */
private fun assertNoPii(message: String) {
assertFalse(message.contains("@example.org"), message)
assertFalse(message.contains("example.org"), message)
}
}
@@ -1,32 +1,60 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.data.sync
import android.util.Log
import androidx.work.ListenableWorker.Result
import dagger.Lazy
import io.mockk.coEvery
import io.mockk.coVerify
import io.mockk.every
import io.mockk.mockk
import io.mockk.mockkStatic
import io.mockk.unmockkAll
import io.mockk.verify
import kotlinx.coroutines.test.runTest
import org.junit.After
import org.junit.Before
import org.junit.Test
import org.libremail.data.security.EncryptedCacheGuard
import org.libremail.reporting.AppLog
import org.libremail.reporting.RingLogBuffer
import kotlin.test.assertEquals
import kotlin.test.assertFalse
import kotlin.test.assertTrue
/**
* [PruneWorker] must not touch the database while the encrypted cache is locked: opening it would park
* this WorkManager thread on an unsatisfiable passphrase await (and wedge the shared serial executor).
* It gates on [EncryptedCacheGuard] and only resolves the (`Lazy`) [MailPruner] once unlocked — the
* invariant every pre-auth DB entry point shares with `SyncWorker`/`SendWorker`.
* invariant every pre-auth DB entry point shares with `SyncWorker`/`SendWorker`. Also covers issue
* #329's AppLog breadcrumbs on the deferred/success/retry outcomes.
*/
class PruneWorkerTest {
private val pruner = mockk<MailPruner>()
private val lazyPruner = mockk<Lazy<MailPruner>> { every { get() } returns pruner }
private val cacheGuard = mockk<EncryptedCacheGuard>()
private val logBuffer = RingLogBuffer()
private fun worker() = PruneWorker(mockk(relaxed = true), mockk(relaxed = true), lazyPruner, cacheGuard)
@Before
fun setUp() {
// `android.util.Log` is a no-op stub under plain JVM unit tests, so it is statically mocked here,
// mirroring org.libremail.reporting.AppLogTest — doWork() now breadcrumbs through AppLog.
mockkStatic(Log::class)
every { Log.d(any(), any()) } returns 0
every { Log.d(any(), any(), any()) } returns 0
every { Log.i(any(), any()) } returns 0
every { Log.w(any<String>(), any<String>()) } returns 0
every { Log.w(any<String>(), any<String>(), any()) } returns 0
every { Log.e(any(), any(), any()) } returns 0
AppLog.install(logBuffer)
}
@After
fun tearDown() = unmockkAll()
@Test
fun `retries without resolving the pruner when the cache is locked`() = runTest {
coEvery { cacheGuard.isCacheLocked() } returns true
@@ -55,4 +83,43 @@ class PruneWorkerTest {
assertEquals(Result.retry(), worker().doWork())
}
// --- issue #329: AppLog breadcrumbs ---------------------------------------------------------
@Test
fun `logs a deferred breadcrumb when the cache is locked, without touching the pruner`() = runTest {
coEvery { cacheGuard.isCacheLocked() } returns true
worker().doWork()
val entry = logBuffer.snapshot().single()
assertEquals('I', entry.level)
assertEquals("prune deferred: cache locked", entry.message)
}
@Test
fun `logs a success breadcrumb when pruning succeeds`() = runTest {
coEvery { cacheGuard.isCacheLocked() } returns false
coEvery { pruner.prune(any()) } returns 5
worker().doWork()
val entry = logBuffer.snapshot().single()
assertEquals('I', entry.level)
assertEquals("prune worker: success", entry.message)
}
@Test
fun `logs a scrubbed retry breadcrumb when pruning throws`() = runTest {
coEvery { cacheGuard.isCacheLocked() } returns false
coEvery { pruner.prune(any()) } throws IllegalStateException("auth failed for a@example.org")
worker().doWork()
val entry = logBuffer.snapshot().single()
assertEquals('W', entry.level)
assertTrue(entry.message.startsWith("prune worker: retry"), entry.message)
assertTrue(entry.message.contains("IllegalStateException"), entry.message)
assertFalse(entry.message.contains("a@example.org"), entry.message)
}
}
@@ -2,8 +2,9 @@
package org.libremail.data.sync
import android.content.Context
import android.util.Log
import androidx.work.ListenableWorker.Result
import com.icegreen.greenmail.util.GreenMail
import com.icegreen.greenmail.util.ServerSetupTest
import dagger.Lazy
import io.mockk.coEvery
import io.mockk.coVerify
@@ -25,12 +26,16 @@ import org.libremail.data.local.entity.OutboxEntity
import org.libremail.data.local.entity.ServerConfigEmbedded
import org.libremail.data.local.toOutgoingAttachmentsJson
import org.libremail.data.security.EncryptedCacheGuard
import org.libremail.domain.model.MailSecurity
import org.libremail.domain.model.OutgoingAttachment
import org.libremail.domain.model.SmtpParams
import org.libremail.mail.GraphSendException
import org.libremail.mail.GraphSender
import org.libremail.mail.SendableAttachment
import org.libremail.mail.SmtpSender
import org.libremail.reporting.AppLog
import org.libremail.reporting.RingLogBuffer
import org.libremail.reporting.accountLogRef
import java.io.File
import kotlin.test.assertEquals
import kotlin.test.assertFalse
@@ -42,6 +47,11 @@ import kotlin.test.assertTrue
* the outbox: sending each queued message over SMTP or Microsoft Graph, deleting it on success,
* flagging failures for a retry, and — crucially for the "may have sent" case — never auto-retrying
* a Graph request that might already have delivered.
*
* It also logs breadcrumbs via [AppLog] (outbox-drain size, per-message send result, the Graph→SMTP
* fallback) — asserted here against a real [RingLogBuffer] rather than a mocked `Log`, per #328. Those
* assertions double as the regression cover for #297: `account.email` must never reach a log line,
* only the non-PII [accountLogRef].
*/
class SendWorkerTest {
@@ -56,6 +66,7 @@ class SendWorkerTest {
private val lazyConnection = mockk<Lazy<MailConnectionFactory>> { every { get() } returns connectionFactory }
private val lazyGrants = mockk<Lazy<AttachmentUriGrants>> { every { get() } returns attachmentUriGrants }
private val cacheGuard = mockk<EncryptedCacheGuard>()
private val logBuffer = RingLogBuffer()
private lateinit var cacheDir: File
private lateinit var appContext: Context
@@ -66,6 +77,16 @@ class SendWorkerTest {
appContext = mockk(relaxed = true)
every { appContext.cacheDir } returns cacheDir
coEvery { cacheGuard.isCacheLocked() } returns false
// AppLog forwards every call to Logcat; stub the Android stub (by fully-qualified name, so
// this file — like the production code it exercises — never imports android.util.Log; only
// AppLog.kt may) so a JVM unit test doesn't crash on the unmocked method, and install a real
// buffer so the breadcrumb + #297 no-PII assertions below read actual recorded lines rather
// than a `Log` verification.
mockkStatic(android.util.Log::class)
every { android.util.Log.i(any(), any()) } returns 0
every { android.util.Log.w(any<String>(), any<String>()) } returns 0
every { android.util.Log.w(any<String>(), any<String>(), any()) } returns 0
AppLog.install(logBuffer)
}
@After
@@ -128,6 +149,17 @@ class SendWorkerTest {
coVerify(exactly = 0) { smtpSender.send(any(), any(), any(), any()) }
}
@Test
fun `doWork logs the outbox-drain breadcrumb with the queued count`() = runTest {
coEvery { outboxDao.getAll() } returns listOf(entity())
coEvery { accountDao.getById("acct") } returns account("acct", "PASSWORD_IMAP")
coEvery { connectionFactory.smtpParamsFor(any()) } returns mockk<SmtpParams>()
worker().doWork()
assertTrue(logBuffer.snapshot().map { it.message }.contains("outbox drain: 1 queued"))
}
@Test
fun `drops a queued message whose account was removed`() = runTest {
val row = entity()
@@ -142,7 +174,7 @@ class SendWorkerTest {
}
@Test
fun `sends a password account message over SMTP then clears it`() = runTest {
fun `sends a password account message over SMTP then clears it, logging a PII-free breadcrumb`() = runTest {
coEvery { outboxDao.getAll() } returns listOf(entity())
coEvery { accountDao.getById("acct") } returns account("acct", "PASSWORD_IMAP")
coEvery { connectionFactory.smtpParamsFor(any()) } returns mockk<SmtpParams>()
@@ -152,10 +184,13 @@ class SendWorkerTest {
coVerify { smtpSender.send(any(), "acct@example.org", any(), any()) }
coVerify { outboxDao.delete("m1") }
coVerify { attachmentUriGrants.releaseUnreferenced(any()) }
val messages = logBuffer.snapshot().map { it.message }
assertTrue(messages.contains("sent ${accountLogRef("acct")} via SMTP"), "messages=$messages")
messages.forEach { assertFalse(it.contains("acct@example.org"), it) }
}
@Test
fun `an SMTP failure flags the row and retries`() = runTest {
fun `an SMTP failure flags the row, retries, and logs a PII-free failure breadcrumb`() = runTest {
coEvery { outboxDao.getAll() } returns listOf(entity())
coEvery { accountDao.getById("acct") } returns account("acct", "PASSWORD_IMAP")
coEvery { connectionFactory.smtpParamsFor(any()) } returns mockk<SmtpParams>()
@@ -165,10 +200,13 @@ class SendWorkerTest {
coVerify { outboxDao.setError("m1", "smtp down") }
coVerify(exactly = 0) { outboxDao.delete(any()) }
val messages = logBuffer.snapshot().map { it.message }
assertTrue(messages.contains("send failed for ${accountLogRef("acct")}; will retry"), "messages=$messages")
messages.forEach { assertFalse(it.contains("acct@example.org"), it) }
}
@Test
fun `sends an Outlook message over Graph`() = runTest {
fun `sends an Outlook message over Graph, logging a PII-free breadcrumb`() = runTest {
coEvery { outboxDao.getAll() } returns listOf(entity())
coEvery { accountDao.getById("acct") } returns account("acct", "OAUTH_OUTLOOK")
coEvery { connectionFactory.graphTokenFor(any()) } returns "graph-token"
@@ -178,6 +216,9 @@ class SendWorkerTest {
coVerify { graphSender.send("graph-token", any(), any()) }
coVerify { outboxDao.delete("m1") }
coVerify(exactly = 0) { smtpSender.send(any(), any(), any(), any()) }
val messages = logBuffer.snapshot().map { it.message }
assertTrue(messages.contains("sent ${accountLogRef("acct")} via Graph"), "messages=$messages")
messages.forEach { assertFalse(it.contains("acct@example.org"), it) }
}
@Test
@@ -212,9 +253,7 @@ class SendWorkerTest {
}
@Test
fun `a Graph transport error falls back to SMTP`() = runTest {
mockkStatic(Log::class)
every { Log.w(any(), any<String>(), any<Throwable>()) } returns 0
fun `a Graph transport error falls back to SMTP and logs a PII-free fallback breadcrumb`() = runTest {
coEvery { outboxDao.getAll() } returns listOf(entity())
coEvery { accountDao.getById("acct") } returns account("acct", "OAUTH_OUTLOOK")
coEvery { connectionFactory.graphTokenFor(any()) } throws RuntimeException("token refresh failed")
@@ -224,6 +263,75 @@ class SendWorkerTest {
coVerify { smtpSender.send(any(), "acct@example.org", any(), any()) }
coVerify { outboxDao.delete("m1") }
val messages = logBuffer.snapshot().map { it.message }
assertTrue(
messages.any { it.startsWith("Graph send failed for ${accountLogRef("acct")}; falling back to SMTP") },
"messages=$messages",
)
messages.forEach { assertFalse(it.contains("acct@example.org"), it) }
}
@Test
fun `regression #297 - no send breadcrumb ever contains the account email`() = runTest {
coEvery { outboxDao.getAll() } returns listOf(entity())
coEvery { accountDao.getById("acct") } returns account("acct", "OAUTH_OUTLOOK")
coEvery { connectionFactory.graphTokenFor(any()) } throws RuntimeException("token refresh failed")
coEvery { connectionFactory.smtpParamsFor(any()) } returns mockk<SmtpParams>()
worker().doWork()
val snapshot = logBuffer.snapshot()
// This path logs the outbox-drain, the Graph->SMTP fallback (with a scrubbed throwable), and
// the send-result breadcrumbs — the richest set of log lines for one message. None may carry
// the account's raw email, only the non-reversible accountLogRef.
assertTrue(snapshot.isNotEmpty(), "expected breadcrumbs to be recorded")
snapshot.forEach { entry ->
assertFalse(entry.message.contains("acct@example.org"), entry.message)
assertFalse(entry.message.contains("@example.org"), entry.message)
}
}
@Test
fun `sends over a real SMTP server end-to-end and logs a PII-free breadcrumb`() = runTest {
// The "connectivity/send" E2E surface for this ticket: a real SmtpSender talking to a real
// (in-process) SMTP server, rather than the mocked smtpSender used by the tests above.
val greenMail = GreenMail(ServerSetupTest.SMTP)
greenMail.start()
greenMail.setUser("acct@example.org", "smtp-secret")
try {
coEvery { outboxDao.getAll() } returns listOf(entity())
coEvery { accountDao.getById("acct") } returns account("acct", "PASSWORD_IMAP")
coEvery { connectionFactory.smtpParamsFor(any()) } returns SmtpParams(
host = "127.0.0.1",
port = greenMail.smtp.port,
security = MailSecurity.NONE,
username = "acct@example.org",
secret = "smtp-secret",
useXoauth2 = false,
)
val realSmtpWorker = SendWorker(
appContext,
mockk(relaxed = true),
lazyOutbox,
lazyAccount,
SmtpSender(),
graphSender,
lazyConnection,
cacheGuard,
lazyGrants,
)
assertEquals(Result.success(), realSmtpWorker.doWork())
greenMail.waitForIncomingEmail(1)
assertEquals(1, greenMail.receivedMessages.size)
coVerify { outboxDao.delete("m1") }
val messages = logBuffer.snapshot().map { it.message }
assertTrue(messages.contains("sent ${accountLogRef("acct")} via SMTP"), "messages=$messages")
messages.forEach { assertFalse(it.contains("acct@example.org"), it) }
} finally {
greenMail.stop()
}
}
@Test
@@ -0,0 +1,57 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.data.sync
import org.junit.Test
import kotlin.test.assertEquals
/**
* [logSafeFolderLabel] is the PII gate between a raw IMAP folder path and an [org.libremail.reporting.AppLog]
* breadcrumb (issue #329's folder-name caveat): a known system folder logs by name, anything else logs
* as the fixed placeholder — never the real name, regardless of nesting or delimiter.
*/
class SyncLoggingTest {
@Test
fun `a top-level system folder logs verbatim`() {
assertEquals("INBOX", logSafeFolderLabel("INBOX"))
assertEquals("Sent", logSafeFolderLabel("Sent"))
assertEquals("Drafts", logSafeFolderLabel("Drafts"))
assertEquals("Trash", logSafeFolderLabel("Trash"))
assertEquals("Archive", logSafeFolderLabel("Archive"))
}
@Test
fun `matching is case-insensitive but preserves the server's own casing`() {
assertEquals("JUNK", logSafeFolderLabel("JUNK"))
assertEquals("sent items", logSafeFolderLabel("sent items"))
}
@Test
fun `a system folder nested under a Gmail-style slash path logs only its leaf name`() {
assertEquals("Sent Mail", logSafeFolderLabel("[Gmail]/Sent Mail"))
assertEquals("All Mail", logSafeFolderLabel("[Gmail]/All Mail"))
}
@Test
fun `a system folder nested under a Dovecot-style dot path logs only its leaf name`() {
assertEquals("Trash", logSafeFolderLabel("INBOX.Trash"))
}
@Test
fun `a user-created top-level folder never logs its real name`() {
assertEquals("<folder>", logSafeFolderLabel("Projects"))
assertEquals("<folder>", logSafeFolderLabel("Invoice from Acme Corp"))
}
@Test
fun `a user-created nested folder logs the placeholder, never its name or its parent`() {
assertEquals("<folder>", logSafeFolderLabel("Work/Reports"))
assertEquals("<folder>", logSafeFolderLabel("Clients/Acme Corp"))
}
@Test
fun `a folder that merely shares a system folder's leaf name still logs only that leaf`() {
// Even a custom parent is never exposed — only the recognized leaf is ever logged, by design.
assertEquals("Archive", logSafeFolderLabel("Clients/Archive"))
}
}
@@ -1,31 +1,59 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.data.sync
import android.util.Log
import androidx.work.ListenableWorker
import dagger.Lazy
import io.mockk.coEvery
import io.mockk.coVerify
import io.mockk.every
import io.mockk.mockk
import io.mockk.mockkStatic
import io.mockk.unmockkAll
import io.mockk.verify
import kotlinx.coroutines.test.runTest
import org.junit.After
import org.junit.Before
import org.junit.Test
import org.libremail.data.security.EncryptedCacheGuard
import org.libremail.reporting.AppLog
import org.libremail.reporting.RingLogBuffer
import kotlin.test.assertEquals
import kotlin.test.assertFalse
import kotlin.test.assertTrue
/**
* [SyncWorker] already gates on [EncryptedCacheGuard]; this locks that invariant in as regression
* cover for the class of bug fixed in PruneWorker/BackfillWorker — while the cache is locked it must
* defer without resolving the (`Lazy`) [MailSyncer] (whose construction opens the Room DB).
* defer without resolving the (`Lazy`) [MailSyncer] (whose construction opens the Room DB). Also covers
* issue #329's AppLog breadcrumbs on the deferred/success/retry outcomes.
*/
class SyncWorkerTest {
private val mailSyncer = mockk<MailSyncer>()
private val lazySyncer = mockk<Lazy<MailSyncer>> { every { get() } returns mailSyncer }
private val cacheGuard = mockk<EncryptedCacheGuard>()
private val logBuffer = RingLogBuffer()
private fun worker() = SyncWorker(mockk(relaxed = true), mockk(relaxed = true), lazySyncer, cacheGuard)
@Before
fun setUp() {
// `android.util.Log` is a no-op stub under plain JVM unit tests, so it is statically mocked here,
// mirroring org.libremail.reporting.AppLogTest — doWork() now breadcrumbs through AppLog.
mockkStatic(Log::class)
every { Log.d(any(), any()) } returns 0
every { Log.d(any(), any(), any()) } returns 0
every { Log.i(any(), any()) } returns 0
every { Log.w(any<String>(), any<String>()) } returns 0
every { Log.w(any<String>(), any<String>(), any()) } returns 0
every { Log.e(any(), any(), any()) } returns 0
AppLog.install(logBuffer)
}
@After
fun tearDown() = unmockkAll()
@Test
fun `retries without resolving the syncer when the cache is locked`() = runTest {
coEvery { cacheGuard.isCacheLocked() } returns true
@@ -53,4 +81,43 @@ class SyncWorkerTest {
assertEquals(ListenableWorker.Result.retry(), worker().doWork())
}
// --- issue #329: AppLog breadcrumbs ---------------------------------------------------------
@Test
fun `logs a deferred breadcrumb when the cache is locked, without touching the syncer`() = runTest {
coEvery { cacheGuard.isCacheLocked() } returns true
worker().doWork()
val entry = logBuffer.snapshot().single()
assertEquals('I', entry.level)
assertEquals("sync deferred: cache locked", entry.message)
}
@Test
fun `logs a success breadcrumb when syncing succeeds`() = runTest {
coEvery { cacheGuard.isCacheLocked() } returns false
coEvery { mailSyncer.syncAll() } returns Result.success(3)
worker().doWork()
val entry = logBuffer.snapshot().single()
assertEquals('I', entry.level)
assertEquals("sync worker: success", entry.message)
}
@Test
fun `logs a scrubbed retry breadcrumb when syncing fails`() = runTest {
coEvery { cacheGuard.isCacheLocked() } returns false
coEvery { mailSyncer.syncAll() } returns Result.failure(IllegalStateException("auth failed for a@example.org"))
worker().doWork()
val entry = logBuffer.snapshot().single()
assertEquals('W', entry.level)
assertTrue(entry.message.startsWith("sync worker: retry"), entry.message)
assertTrue(entry.message.contains("IllegalStateException"), entry.message)
assertFalse(entry.message.contains("a@example.org"), entry.message)
}
}
@@ -36,6 +36,9 @@ class CountingImapProxy(private val backendHost: String, private val backendPort
/** Client → server pump threads, tracked so tests can wait for the parsed command stream to settle. */
private val clientPumps = Collections.synchronizedList(mutableListOf<Thread>())
/** Accepted client-side sockets, so a test can force-drop them to simulate a server/NAT disconnect. */
private val acceptedSockets = Collections.synchronizedList(mutableListOf<Socket>())
@Volatile private var running = true
/** The local port to point [ImapClient] at; it forwards to the backend. */
@@ -69,6 +72,17 @@ class CountingImapProxy(private val backendHost: String, private val backendPort
}
}
/**
* Force-closes every currently-accepted client socket, simulating a server idle-timeout / NAT
* rebind / network drop of the kept-alive reused connection. The client's next use of that socket
* then fails with an I/O error, which the connection-reuse cache should transparently reconnect
* from. [connectionCount] keeps counting, so a subsequent reconnect makes it rise.
*/
fun dropAcceptedConnections() {
val snapshot = synchronized(acceptedSockets) { acceptedSockets.toList().also { acceptedSockets.clear() } }
snapshot.forEach { runCatching { it.close() } }
}
override fun close() {
running = false
runCatching { server.close() }
@@ -88,6 +102,7 @@ class CountingImapProxy(private val backendHost: String, private val backendPort
runCatching { client.close() }
continue
}
acceptedSockets.add(client)
val upstream = Thread({ pumpCountingCommands(client, backend) }, "imap-proxy-up").apply { isDaemon = true }
val downstream = Thread({ pump(backend, client) }, "imap-proxy-down").apply { isDaemon = true }
clientPumps.add(upstream)
@@ -3,6 +3,9 @@ package org.libremail.mail
import com.icegreen.greenmail.util.GreenMail
import com.icegreen.greenmail.util.ServerSetupTest
import io.mockk.every
import io.mockk.mockkStatic
import io.mockk.unmockkAll
import jakarta.mail.Folder
import jakarta.mail.Message
import jakarta.mail.Session
@@ -27,17 +30,30 @@ import kotlin.test.assertTrue
class ImapClientBackfillTest {
private lateinit var greenMail: GreenMail
private val client = ImapClient()
// Reuse off: this suite pins the connect-per-operation paging behaviour it was written against.
private val client = ImapClient(reuseConnections = false)
@Before
fun setUp() {
greenMail = GreenMail(ServerSetupTest.SMTP_IMAP)
greenMail.start()
greenMail.setUser("alice@example.org", "secret")
// Every IMAP op now breadcrumbs through AppLog (per-op connect/work timings, issue #358), and
// android.util.Log is a no-op stub under plain JVM tests. Mock it class-wide — fully qualified so
// this file still never imports android.util.Log — so no test crashes on the unmocked method.
mockkStatic(android.util.Log::class)
every { android.util.Log.d(any(), any()) } returns 0
every { android.util.Log.i(any(), any()) } returns 0
every { android.util.Log.w(any<String>(), any<String>()) } returns 0
}
@After
fun tearDown() = greenMail.stop()
fun tearDown() {
greenMail.stop()
unmockkAll()
}
private fun params() = ImapConnectionParams(
host = "127.0.0.1",
@@ -1,7 +1,6 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.mail
import android.util.Log
import com.icegreen.greenmail.util.GreenMail
import com.icegreen.greenmail.util.GreenMailUtil
import com.icegreen.greenmail.util.ServerSetupTest
@@ -31,6 +30,8 @@ import org.junit.Before
import org.junit.Test
import org.libremail.domain.model.ImapConnectionParams
import org.libremail.domain.model.MailSecurity
import org.libremail.reporting.AppLog
import org.libremail.reporting.RingLogBuffer
import java.util.Properties
import kotlin.test.assertContentEquals
import kotlin.test.assertEquals
@@ -41,13 +42,24 @@ import kotlin.test.assertTrue
class ImapClientTest {
private lateinit var greenMail: GreenMail
private val client = ImapClient()
// Reuse off: this suite pins the connect-per-operation IMAP behaviour it was written against. The
// connection-reuse path (production default) has its own coverage in ImapFolderOpenLatencyTest.
private val client = ImapClient(reuseConnections = false)
@Before
fun setUp() {
greenMail = GreenMail(ServerSetupTest.SMTP_IMAP)
greenMail.start()
greenMail.setUser("alice@example.org", "secret")
// Every IMAP op now breadcrumbs through AppLog (per-op connect/work timings, issue #358), and
// android.util.Log is a no-op stub under plain JVM tests. Mock it class-wide — fully qualified so
// this file still never imports android.util.Log — so no test crashes on the unmocked method.
mockkStatic(android.util.Log::class)
every { android.util.Log.d(any(), any()) } returns 0
every { android.util.Log.i(any(), any()) } returns 0
every { android.util.Log.w(any<String>(), any<String>()) } returns 0
}
@After
@@ -101,6 +113,35 @@ class ImapClientTest {
assertTrue(client.fetchRecent(params(), "INBOX", limit = 50).first().isRead, "should be marked read")
}
@Test
fun `fetchBodyMarkingSeen breadcrumbs connect and phase timings without leaking PII`() = runTest {
GreenMailUtil.sendTextEmailTest("alice@example.org", "bob@example.org", "Secret subject", "Body.")
greenMail.waitForIncomingEmail(1)
val uid = client.fetchRecent(params(), "INBOX", limit = 50).first().uid
val buffer = RingLogBuffer()
AppLog.install(buffer)
client.fetchBodyMarkingSeen(params(), "INBOX", uid)
val messages = buffer.snapshot().map { it.message }
// withStore labels the op and splits connect (CONNECT+TLS+LOGIN) from work — timings only.
assertTrue(
messages.any { it.startsWith("body-fetch connect=") && it.contains("work=") && it.contains("live=") },
"messages=$messages",
)
// fetchBodyMarkingSeen adds the select/body/flag phase split plus PII-free size counts.
assertTrue(
messages.any { it.startsWith("body-fetch select=") && it.contains("rfc822=") && it.contains("att=") },
"messages=$messages",
)
// Breadcrumbs are numbers and fixed labels only — never the subject, sender, or recipient.
messages.forEach { message ->
assertFalse(message.contains("Secret subject"), message)
assertFalse(message.contains("bob@example.org"), message)
assertFalse(message.contains("alice@example.org"), message)
}
}
@Test
fun `search returns only messages matching the query`() = runTest {
GreenMailUtil.sendTextEmailTest("alice@example.org", "bob@example.org", "Vacation plans", "Beach trip")
@@ -359,8 +400,14 @@ class ImapClientTest {
@Test
fun `idle syncs once on connect, again on newly delivered mail, and stops on cancel`() = runBlocking {
mockkStatic(Log::class) // idle() logs connect/push at debug level
every { Log.d(any(), any()) } returns 0
// idle() logs connect/push breadcrumbs via AppLog, which forwards to Logcat; stub the Android
// stub (by fully-qualified name, so this file never imports android.util.Log — only AppLog.kt
// may) so the JVM test doesn't crash on the unmocked method, and install a real buffer so the
// breadcrumbs can be asserted directly instead of via a Log verification.
mockkStatic(android.util.Log::class)
every { android.util.Log.d(any(), any()) } returns 0
val buffer = RingLogBuffer()
AppLog.install(buffer)
val activity = Channel<Unit>(Channel.UNLIMITED)
val job = launch(Dispatchers.IO) { client.idle(params()) { activity.send(Unit) } }
try {
@@ -373,6 +420,17 @@ class ImapClientTest {
job.cancelAndJoin() // cancelling closes the connection and unblocks idle()
}
assertTrue(job.isCompleted, "the idle loop must terminate on cancellation")
val messages = buffer.snapshot().map { it.message }
assertTrue(messages.contains("IDLE connected"), "messages=$messages")
assertTrue(messages.any { it == "IDLE push: 1 new message(s)" }, "messages=$messages")
// idle() only holds host/username via ImapConnectionParams and must stay account-agnostic —
// attribution belongs to the IdleService caller (accountLogRef) — regression guard for #297.
val connectionParams = params()
messages.forEach { message ->
assertFalse(message.contains(connectionParams.host), message)
assertFalse(message.contains(connectionParams.username), message)
}
}
/** Creates [folderName] (holding messages) if it does not already exist. */
@@ -1,7 +1,10 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.mail
import io.mockk.every
import io.mockk.mockk
import io.mockk.mockkStatic
import io.mockk.unmockkAll
import io.mockk.verify
import jakarta.mail.Folder
import jakarta.mail.FolderClosedException
@@ -9,6 +12,8 @@ import jakarta.mail.MessagingException
import jakarta.mail.Store
import jakarta.mail.StoreClosedException
import kotlinx.coroutines.test.runTest
import org.junit.After
import org.junit.Before
import org.junit.Test
import org.libremail.domain.model.ImapConnectionParams
import org.libremail.domain.model.MailSecurity
@@ -17,10 +22,11 @@ import kotlin.test.assertEquals
import kotlin.test.assertFailsWith
/**
* The connection-reuse cache (issue #125 spike): one authenticated [Store] per account behind a
* mutex, established lazily and kept open. These tests pin the reuse guarantee and the lazy
* catch-and-retry-once stale handling — a dropped connection is rebuilt and the op retried, a second
* failure clears the slot, and a genuine protocol error is propagated without ever reconnecting.
* The connection-reuse cache (issue #357 Part 2, wiring the #125 spike): one authenticated [Store] per
* account behind a mutex, established lazily and kept open. These tests pin the reuse guarantee, the
* lazy catch-and-retry-once stale handling — a dropped connection is rebuilt and the op retried, a
* second failure clears the slot, and a genuine protocol error is propagated without ever reconnecting
* — plus idle eviction (a connection unused past the timeout is closed) and teardown.
*/
class ImapConnectionCacheTest {
@@ -35,18 +41,41 @@ class ImapConnectionCacheTest {
private var connects = 0
/** A cache whose connect step counts calls and returns [supply] (a fresh relaxed [Store] by default). */
private fun cache(supply: () -> Store = { mockk(relaxed = true) }) = ImapConnectionCache {
connects++
supply()
/** Injected monotonic clock (nanos) so idle eviction is deterministic; advanced by the test. */
private var nowNanos = 0L
@Before
fun setUp() {
// The cache breadcrumbs through AppLog, which forwards to android.util.Log — a throwing no-op
// stub under plain JVM tests. Mock it class-wide (fully qualified, so this file never imports
// android.util.Log) so no test crashes on the unmocked method.
mockkStatic(android.util.Log::class)
every { android.util.Log.d(any(), any()) } returns 0
every { android.util.Log.d(any(), any(), any()) } returns 0
every { android.util.Log.w(any<String>(), any<String>()) } returns 0
every { android.util.Log.w(any<String>(), any<String>(), any()) } returns 0
}
@After
fun tearDown() = unmockkAll()
/** A cache whose connect step counts calls and returns [supply] (a fresh relaxed [Store] by default). */
private fun cache(idleTimeoutMillis: Long = IDLE_TIMEOUT_MS, supply: () -> Store = { mockk(relaxed = true) }) =
ImapConnectionCache(
connect = {
connects++
supply()
},
idleTimeoutMillis = idleTimeoutMillis,
nowNanos = { nowNanos },
)
@Test
fun `establishes one connection and reuses it across calls`() = runTest {
val cache = cache()
assertEquals("a", cache.withStore(params) { "a" })
assertEquals("b", cache.withStore(params) { "b" })
assertEquals("a", cache.withStore(params, op = "test") { "a" })
assertEquals("b", cache.withStore(params, op = "test") { "b" })
assertEquals(1, connects, "the second op reuses the first connection")
}
@@ -59,7 +88,7 @@ class ImapConnectionCacheTest {
}
var attempts = 0
val result = cache.withStore(params) {
val result = cache.withStore(params, op = "test") {
attempts++
if (attempts == 1) throw IOException("dropped") else "recovered"
}
@@ -73,10 +102,10 @@ class ImapConnectionCacheTest {
fun `a second failure after reconnect clears the slot so the next call reconnects`() = runTest {
val cache = cache()
assertFailsWith<IOException> { cache.withStore(params) { throw IOException("still down") } }
assertFailsWith<IOException> { cache.withStore(params, op = "test") { throw IOException("still down") } }
assertEquals(2, connects, "initial connect plus one rebuild")
cache.withStore(params) { "ok" }
cache.withStore(params, op = "test") { "ok" }
assertEquals(3, connects, "the cleared slot forces a fresh connect")
}
@@ -85,7 +114,7 @@ class ImapConnectionCacheTest {
val cache = cache()
assertFailsWith<IllegalStateException> {
cache.withStore(params) { throw IllegalStateException("bad login") }
cache.withStore(params, op = "test") { throw IllegalStateException("bad login") }
}
assertEquals(1, connects, "a non-drop error must not trigger a reconnect")
@@ -104,22 +133,44 @@ class ImapConnectionCacheTest {
val cache = cache()
assertFailsWith<MessagingException> {
cache.withStore(params) { throw MessagingException("server said no") }
cache.withStore(params, op = "test") { throw MessagingException("server said no") }
}
assertEquals(1, connects)
}
@Test
fun `evictIdle closes a connection idle past the timeout but keeps a fresh one`() = runTest {
val store = mockk<Store>(relaxed = true)
val cache = cache(idleTimeoutMillis = IDLE_TIMEOUT_MS) { store }
nowNanos = 0L
cache.withStore(params, op = "test") { "a" } // establish; lastUsed = 0
// Still within the timeout: not yet idle -> kept.
nowNanos = (IDLE_TIMEOUT_MS - 1) * NANOS_PER_MS
cache.evictIdle()
verify(exactly = 0) { store.close() }
// Idle past the timeout -> evicted, and the next op reconnects.
nowNanos = IDLE_TIMEOUT_MS * NANOS_PER_MS
cache.evictIdle()
verify(exactly = 1) { store.close() }
cache.withStore(params, op = "test") { "b" }
assertEquals(2, connects, "after idle eviction the next op reconnects")
}
@Test
fun `closeAll tears down and forgets every cached connection`() = runTest {
val store = mockk<Store>(relaxed = true)
val cache = cache { store }
cache.withStore(params) { "a" }
cache.withStore(params, op = "test") { "a" }
cache.closeAll()
verify { store.close() }
cache.withStore(params) { "b" }
cache.withStore(params, op = "test") { "b" }
assertEquals(2, connects, "after closeAll the next op reconnects")
}
@@ -129,7 +180,7 @@ class ImapConnectionCacheTest {
val cache = cache()
var attempts = 0
val result = cache.withStore(params) {
val result = cache.withStore(params, op = "test") {
attempts++
if (attempts == 1) throw error else "ok"
}
@@ -137,4 +188,9 @@ class ImapConnectionCacheTest {
assertEquals("ok", result)
assertEquals(2, connects, "${error.javaClass.simpleName} should have been retried on a fresh socket")
}
private companion object {
const val IDLE_TIMEOUT_MS = 60_000L
const val NANOS_PER_MS = 1_000_000L
}
}
@@ -4,6 +4,9 @@ package org.libremail.mail
import com.icegreen.greenmail.util.GreenMail
import com.icegreen.greenmail.util.GreenMailUtil
import com.icegreen.greenmail.util.ServerSetupTest
import io.mockk.every
import io.mockk.mockkStatic
import io.mockk.unmockkAll
import kotlinx.coroutines.runBlocking
import kotlinx.coroutines.test.runTest
import org.junit.After
@@ -12,6 +15,7 @@ import org.junit.Test
import org.libremail.domain.model.ImapConnectionParams
import org.libremail.domain.model.MailSecurity
import kotlin.test.assertEquals
import kotlin.test.assertFailsWith
import kotlin.test.assertTrue
/**
@@ -19,28 +23,36 @@ import kotlin.test.assertTrue
* deterministically and without a real network, by routing [ImapClient] through a [CountingImapProxy]
* that counts the TCP connections and IMAP commands it establishes.
*
* The finding these tests pin down: [ImapClient] wraps every operation in its own short-lived
* [jakarta.mail.Store] (`withStore`), so **each folder-open pays a fresh CONNECT + LOGIN + SELECT +
* FETCH + LOGOUT** — nothing is reused between operations. On a real network the CONNECT + TLS + LOGIN
* group is several RTTs of user-perceived latency that a pooled/kept-alive connection would pay only
* once. See `docs/perf/issue-125-imap-folder-open.md`.
* Two contrasting behaviours are pinned. With reuse **off**, [ImapClient] wraps every operation in its
* own short-lived [jakarta.mail.Store] (`withStore`), so **each folder-open pays a fresh CONNECT +
* LOGIN + SELECT + FETCH + LOGOUT** — nothing is reused. With reuse **on** (the production default,
* issue #357 Part 2), the same real IMAP operations collapse onto **one** kept-alive connection / one
* LOGIN, with the necessary per-folder EXAMINE unchanged — and a dropped socket is transparently
* reconnected, an application error is not mistaken for a drop, and an idle socket is evicted. On a
* real network the CONNECT + TLS + LOGIN group is several RTTs of user-perceived latency (and, on
* Gmail, the connect *volume* that trips server-side throttling) that reuse pays only once. See
* `docs/perf/issue-125-imap-folder-open.md`.
*
* These assertions encode the *current* (no-reuse) behaviour. They are also the harness to validate a
* future connection-reuse fix: when the client reuses one authenticated connection across folder
* switches, the connection/auth counts here drop below the operation count — flip the expectations to
* assert reuse and the tests confirm the win against a real IMAP server.
* The reuse-on assertions are also the regression guard: they fail if reuse ever silently regresses to
* connect-per-operation.
*/
class ImapFolderOpenLatencyTest {
private lateinit var greenMail: GreenMail
private lateinit var proxy: CountingImapProxy
/** Flag OFF (production default): a fresh connect + LOGOUT per operation. */
private val client = ImapClient()
/** Reuse OFF: a fresh connect + LOGOUT per operation — the baseline these counts contrast against. */
private val client = ImapClient(reuseConnections = false)
/** Flag ON (the spike prototype, issue #125): one kept-alive connection reused across operations. */
/** Reuse ON (production default, issue #357 Part 2): one kept-alive connection reused across ops. */
private val reuseClient = ImapClient(reuseConnections = true)
/**
* Reuse ON with a zero idle timeout, so `evictIdleReusedConnections()` closes the kept-alive socket
* immediately — lets the idle-eviction test assert the teardown deterministically without a clock.
*/
private val evictClient = ImapClient(reuseConnections = true, reuseIdleTimeoutMillis = 0L)
@Before
fun setUp() {
greenMail = GreenMail(ServerSetupTest.SMTP_IMAP)
@@ -48,13 +60,28 @@ class ImapFolderOpenLatencyTest {
greenMail.setUser("alice@example.org", "secret")
// All IMAP traffic goes through the proxy so we can count it; the proxy forwards to GreenMail.
proxy = CountingImapProxy(backendHost = "127.0.0.1", backendPort = greenMail.imap.port)
// Every IMAP op now breadcrumbs through AppLog (per-op connect/work timings, issue #358), and
// android.util.Log is a no-op stub under plain JVM tests. Mock it class-wide — fully qualified so
// this file still never imports android.util.Log — so no test crashes on the unmocked method.
mockkStatic(android.util.Log::class)
every { android.util.Log.d(any(), any()) } returns 0
every { android.util.Log.d(any(), any(), any()) } returns 0 // reuse-stale reconnect logs with a throwable
every { android.util.Log.i(any(), any()) } returns 0
every { android.util.Log.w(any<String>(), any<String>()) } returns 0
every { android.util.Log.w(any<String>(), any<String>(), any()) } returns 0
}
@After
fun tearDown() {
runBlocking { reuseClient.closeReusedConnections() } // release any kept-alive socket before the server stops
// Release any kept-alive socket before the server stops (a no-op for a client that never reused).
runBlocking {
reuseClient.closeReusedConnections()
evictClient.closeReusedConnections()
}
proxy.close()
greenMail.stop()
unmockkAll()
}
/** Points [ImapClient] at the counting proxy rather than directly at GreenMail. */
@@ -144,7 +171,7 @@ class ImapFolderOpenLatencyTest {
assertEquals(2, proxy.authCommandCount(), "list + read each pay a full LOGIN")
}
// --- Flag ON: the spike prototype reuses one connection across operations (issue #125). ---
// --- Reuse ON (production default): one connection is reused across operations (issue #357 Part 2). ---
// These are the deterministic proof that reuse works: the SAME real-IMAP operations that cost N
// connections / N LOGINs above collapse to ONE connection / ONE LOGIN here, with the necessary
// per-open EXAMINE unchanged. Localhost is ~0 RTT so this proves the STRUCTURE, not wall-clock.
@@ -184,6 +211,58 @@ class ImapFolderOpenLatencyTest {
assertEquals(1, proxy.authCommandCount(), "reuse: one LOGIN covers both the list and the read")
}
// --- Hardening: transparent stale recovery, narrow drop detection, and idle eviction. ---
@Test
fun `with reuse on, a dropped connection is transparently reconnected on the next op`() = runTest {
seedInbox(1)
reuseClient.fetchRecent(params(), "INBOX", limit = 50) // establish the kept-alive connection
assertEquals(1, proxy.connectionCount, "one connection is established and kept alive")
// The server (or NAT / a network change) silently drops the idle socket.
proxy.dropAcceptedConnections()
// The next op must NOT surface an error: the cache detects the dead socket, reconnects once,
// and completes the operation, returning its real result.
val messages = reuseClient.fetchRecent(params(), "INBOX", limit = 50)
assertEquals(1, messages.size, "the op still returns its result after a transparent reconnect")
assertEquals(2, proxy.connectionCount, "a dropped reused socket is transparently reconnected (1 -> 2)")
}
@Test
fun `with reuse on, an application error reuses the live connection rather than reconnecting`() = runTest {
seedInbox(1)
reuseClient.fetchRecent(params(), "INBOX", limit = 50) // establish the kept-alive connection
assertEquals(1, proxy.connectionCount)
// A non-connection error (the UID doesn't exist) must propagate as-is, NOT be mistaken for a
// dropped socket — so the live connection is neither torn down nor needlessly reconnected, and a
// mutation would never be silently re-issued over a working socket.
assertFailsWith<Exception> { reuseClient.fetchBodyMarkingSeen(params(), "INBOX", "999999") }
assertEquals(1, proxy.connectionCount, "an application error keeps reusing the one live connection")
}
@Test
fun `with reuse on, idle eviction closes the socket and the next op reconnects`() = runTest {
seedInbox(1)
evictClient.fetchRecent(params(), "INBOX", limit = 50) // establish the kept-alive connection
assertEquals(1, proxy.connectionCount)
// Idle timeout is zero for evictClient, so the sweep closes the just-used socket now.
evictClient.evictIdleReusedConnections()
proxy.awaitClientStreamsSettled() // let the LOGOUT + close flush through the proxy
assertEquals(1, proxy.commandCount("LOGOUT"), "idle eviction tears the socket down with a LOGOUT")
evictClient.fetchRecent(params(), "INBOX", limit = 50) // must reconnect, the socket is gone
assertEquals(2, proxy.connectionCount, "the next op after idle eviction reconnects (1 -> 2)")
}
private companion object {
const val OPENS = 3
}
@@ -0,0 +1,104 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.push
import android.app.Service
import org.junit.Assert.assertEquals
import org.junit.Assert.assertFalse
import org.junit.Assert.assertNull
import org.junit.Assert.assertSame
import org.junit.Assert.assertTrue
import org.junit.Test
/**
* JVM coverage of [IdleForegroundStarter] — the dataSync foreground-start decision
* `IdleService.onStartCommand` delegates to (#354). Extracted out of the Android `Service` precisely
* so this branchy logic — skip while capped, catch the runtime-cap
* `ForegroundServiceStartNotAllowedException` (surfaced via its [IllegalStateException] supertype) and
* degrade, otherwise proceed — is unit-testable with no emulator and no Robolectric (this repo has
* neither). It reads only the `Service.START_NOT_STICKY` constant; no `Service` is instantiated.
*/
class IdleForegroundStarterTest {
@Test
fun `a successful foreground start proceeds to onStarted and returns START_NOT_STICKY`() {
var started = false
var degradeCalled = false
val result = IdleForegroundStarter.startForegroundOrDegrade(
capActive = false,
enterForeground = { /* startForeground succeeds */ },
onStarted = { started = true },
onDegraded = { degradeCalled = true },
)
assertEquals(Service.START_NOT_STICKY, result)
assertTrue("onStarted must run after a successful foreground start", started)
assertFalse("degrade must not run when the start succeeds", degradeCalled)
}
@Test
fun `a runtime-cap rejection is caught, routed to degrade with the cause, and never propagates`() {
// The exact platform rejection is a ForegroundServiceStartNotAllowedException (API 31+); here we
// throw its IllegalStateException supertype, which is what the seam catches (and what the stub
// android.jar lets us construct on the JVM).
val rejection = IllegalStateException("Time limit already exhausted for foreground service type dataSync")
var started = false
var degradedWith: Throwable? = null
// No exception escapes — that is the whole point of the fix (was an uncaught crash → restart loop).
val result = IdleForegroundStarter.startForegroundOrDegrade(
capActive = false,
enterForeground = { throw rejection },
onStarted = { started = true },
onDegraded = { cause -> degradedWith = cause },
)
assertEquals(Service.START_NOT_STICKY, result)
assertFalse("a rejected foreground start must not begin IDLE watching", started)
assertSame("the rejection cause must reach the degrade path", rejection, degradedWith)
}
@Test
fun `an active cap window skips the foreground start entirely and degrades with no cause`() {
var enterForegroundAttempted = false
var started = false
var degradeCalled = false
var degradedWith: Throwable? = null
val result = IdleForegroundStarter.startForegroundOrDegrade(
capActive = true,
enterForeground = { enterForegroundAttempted = true },
onStarted = { started = true },
onDegraded = { cause ->
degradeCalled = true
degradedWith = cause
},
)
assertEquals(Service.START_NOT_STICKY, result)
assertFalse("must not attempt a dataSync FGS start while still inside the cap window", enterForegroundAttempted)
assertFalse(started)
assertTrue("must still degrade to periodic sync while capped", degradeCalled)
assertNull("the cap-window skip carries no throwable cause", degradedWith)
}
@Test
fun `a non-IllegalStateException from the foreground start propagates unchanged`() {
// Only the runtime-cap/background ISE is a safe-degrade condition; anything else (e.g. a
// SecurityException) is a genuine bug we must not swallow.
val boom = SecurityException("not an FGS runtime-cap rejection")
var degradeCalled = false
val thrown = runCatching {
IdleForegroundStarter.startForegroundOrDegrade(
capActive = false,
enterForeground = { throw boom },
onStarted = {},
onDegraded = { degradeCalled = true },
)
}.exceptionOrNull()
assertSame("an unrelated failure must propagate, not degrade", boom, thrown)
assertFalse("degrade must not run for a non-ISE failure", degradeCalled)
}
}
@@ -0,0 +1,72 @@
// SPDX-License-Identifier: GPL-3.0-or-later
package org.libremail.reporting
import org.junit.Test
import kotlin.test.assertEquals
import kotlin.test.assertFalse
import kotlin.test.assertNotEquals
import kotlin.test.assertTrue
/**
* [accountLogRef] turns a PII-bearing `Account.id` (which embeds the raw email) into a short, stable,
* non-reversible reference that is safe to write to logs and a [DebugReport].
*/
class AccountLogRefTest {
@Test
fun `is deterministic and stable for the same id`() {
val id = "outlook:user@example.com"
assertEquals(accountLogRef(id), accountLogRef(id))
}
@Test
fun `differs across different ids`() {
assertNotEquals(
accountLogRef("outlook:user@example.com"),
accountLogRef("outlook:other@example.com"),
)
// Same address under a different scheme is a different account, so it must map to a different ref.
assertNotEquals(
accountLogRef("outlook:user@example.com"),
accountLogRef("imap:user@example.com"),
)
}
@Test
fun `contains neither the email nor its domain and is not the raw id`() {
val id = "imap:Alice.Smith@example.com"
val ref = accountLogRef(id)
assertNotEquals(id, ref)
assertFalse(ref.contains("@"), ref)
assertFalse(ref.contains("Alice.Smith"), ref)
assertFalse(ref.contains("example.com"), ref)
assertFalse(ref.contains("example"), ref)
}
@Test
fun `keeps the non-PII scheme prefix so refs stay readable`() {
assertTrue(accountLogRef("outlook:user@example.com").startsWith("outlook:"), "outlook prefix")
assertTrue(accountLogRef("imap:user@example.com").startsWith("imap:"), "imap prefix")
}
@Test
fun `falls back to a generic prefix for an id with no scheme, never leaking the address`() {
val ref = accountLogRef("user@example.com")
assertTrue(ref.startsWith("acct:"), ref)
assertFalse(ref.contains("@"), ref)
assertFalse(ref.contains("example"), ref)
}
@Test
fun `never treats an address-shaped prefix as the scheme`() {
// If the part before the first ':' is itself an address, it must not surface as the prefix.
val ref = accountLogRef("user@example.com:143")
assertTrue(ref.startsWith("acct:"), ref)
assertFalse(ref.contains("@"), ref)
assertFalse(ref.contains("example"), ref)
}
}

Some files were not shown because too many files have changed in this diff Show More