Root cause: after #302's runtime-cap fallBackToPeriodicSync() stops the
dataSync foreground service, IdleService was restarted (START_STICKY
null-intent redelivery + explicit startForegroundService) and onStartCommand
unconditionally called startForeground(DATA_SYNC) while the rolling-24h budget
was still exhausted. The platform rejected the start with
ForegroundServiceStartNotAllowedException; it was uncaught, the process
crashed, and START_STICKY restarted straight back into the same rejection -- a
crash loop until the 24h window freed budget (#354).
Fix (IdleService.kt):
- onStartCommand now returns START_NOT_STICKY. Push is app-managed
(LibreMailApplication.ensurePushStarted deterministically restarts it), so the
sticky null-intent auto-restart was redundant and fired exactly when a dataSync
FGS start is illegal.
- Guard the foreground start via a new JVM-testable IdleForegroundStarter seam:
a ForegroundServiceStartNotAllowedException (caught via its IllegalStateException
supertype, so no minSdk-29 class load) degrades like the cap handler --
schedulePeriodicSync(), keep the degraded POLLING notification, stopSelf()
promptly (avoids the "did not call startForeground in time" ANR) -- instead of
propagating.
- Record the cap event (elapsedRealtime); while still inside the cap window,
onStartCommand skips the now-guaranteed-illegal foreground start entirely.
- onTimeout stop path kept fast so ForegroundServiceDidNotStopInTimeException
stays mitigated.
PII-free AppLog.w/i on the degrade paths.
Tests:
- Unit (IdleForegroundStarterTest): onStartCommand returns START_NOT_STICKY; a
rejected start is caught and routed to degrade without propagating; the cap
window skips the attempt; a non-ISE propagates.
- Instrumented (IdleServiceForegroundStartInstrumentedTest): the degrade path on
a real Context -- rejection caught, periodic-sync fallback scheduled, degraded
"instant delivery paused" notification built, watching skipped.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adds AppLog breadcrumbs to the message-open path so a debug report can show
where the reader's spinner time goes:
- ImapClient.withStore: per-op connect vs. work timing plus a live
connect-per-op connection gauge (issue #125's provider-ceiling context).
- fetchBodyMarkingSeen: select/body/flag phase timings plus PII-free size
counts (RFC822 size, body chars, attachment count).
- MailRepositoryImpl.openMessage: end-to-end open latency plus the
cached-vs-fetched branch, keyed by accountLogRef and logSafeFolderLabel.
- ReaderViewModel: spinner-to-ready latency, split success vs. failure.
All breadcrumbs are PII-free: accounts are logged via the existing
accountLogRef hash, folders via the existing logSafeFolderLabel allowlist,
and everything else is sizes/durations/booleans only.
Fixes the 4 unit-test classes that exercise this code without mocking
android.util.Log (a throwing stub under plain JVM tests): mockkStatic(Log)
is now installed in MailRepositoryImplCoverageTest, ImapClientBackfillTest,
ImapFolderOpenLatencyTest, and ReaderViewModelActionsTest, following the
existing MailBackfillerTest/ImapClientTest conventions. detekt.yml gains two
more ForbiddenImport excludes for the newly Log-importing test files.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Move the traffic-control (runner-priority orchestration) job verbatim out of
.github/workflows/ci.yml into a new standalone workflow,
.github/workflows/traffic-control.yml, so the heavy CI jobs no longer depend
on it. The job's YAML (name, runs-on, timeout-minutes, permissions, env,
steps) and its documentation comment move unchanged; the decision core
.github/scripts/traffic_control.py is untouched and still unit-tested by the
traffic-control-tests job in ci.yml.
In ci.yml: removed the traffic-control job, dropped needs: traffic-control
from the five heavy jobs (static-analysis, debug-build, unit-tests, e2e,
e2e-preview) and from traffic-control-tests (its only needs, which would
otherwise dangle at a now-deleted job), and updated the now-stale header and
ci-passed comments to point at the extracted workflow.
The new workflow will be disabled pending a rebuild as a published GitHub
Action.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
#350 made ci-trigger.yml dispatch ci.yml with the built-in GITHUB_TOKEN, on the
claim that a workflow_dispatch is anti-recursion-exempt so no PAT is needed. In
practice a GITHUB_TOKEN-triggered run is held in `action_required` awaiting manual
approval and never runs un-attended, so auto-updated PRs' CI never ran (stalled
#285). The original #349 design was right: dispatch with a PAT so the run executes
as the authorized owner with no approval gate.
- ci-trigger.yml: the trigger step's GH_TOKEN is now
`${{ secrets.AUTOUPDATE_TOKEN || github.token }}` (was `${{ github.token }}`).
AUTOUPDATE_TOKEN (the PAT) is REQUIRED for the scheduler; the `|| github.token`
fallback stays fail-open but only starts CI if repo settings don't gate
GITHUB_TOKEN-triggered runs.
- autoupdate.yml: branch update stays on GITHUB_TOKEN (must NOT retrigger CI --
that would re-introduce the cascade). Clarified that AUTOUPDATE_TOKEN is still
required by the repo (by ci-trigger.yml) so the secret isn't deleted.
- Corrected the now-wrong "no PAT needed / workflow_dispatch anti-recursion-exempt"
comments in ci-trigger.yml and the traffic_control.py docstrings.
updates = GITHUB_TOKEN, triggering = PAT.
Validation: all three workflow YAMLs parse clean; traffic-control unit tests still
pass (59 tests) -- the change is workflow-env only, script logic unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Resolve DatabaseModule conflict from #320: main replaced the explicit .addMigrations(...) chain with .addMigrations(*ALL_MIGRATIONS) plus an introspectable ALL_MIGRATIONS list guarded by databaseModuleRegistersEveryDeclaredMigration (registered == declared). Add MIGRATION_19_20 to ALL_MIGRATIONS so the unified-inbox covering-index migration (cache schema v19->v20) is both registered on the Room builder and satisfies that safety-net test. Schema 20.json, the v20 @Database version, and DatabaseEncryptionTest's schema-version assertion (20) are unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
End the merge cascade and give the traffic-controller ownership of CI *triggering*.
- autoupdate.yml updates PR branches with the built-in GITHUB_TOKEN instead of a PAT,
so an update push no longer auto-retriggers CI (GitHub's anti-recursion rule) — the
cascade (every merge re-runs every PR, cancel-in-progress thrashing them) is gone.
- New scheduler ci-trigger.yml -> traffic_control.py --mode trigger (re-)triggers CI
for the highest-priority PR(s) whose head SHA has absent/stale checks, a few at a
time (inflight cap), in the existing P0-P9 / broken-draft priority order — a
poor-man's merge queue reusing the priority core. It runs after autoupdate finishes
(workflow_run, race-free) plus a cron backstop plus manual dispatch.
- Triggering uses workflow_dispatch, which is EXEMPT from anti-recursion, so the
built-in GITHUB_TOKEN (actions: write) starts the run — NO PAT / secret change needed.
- ci.yml gains a workflow_dispatch trigger (pr/head_sha/reason inputs) and a per-PR
concurrency group unifying pull_request and dispatch runs; its on: pull_request path
is kept so brand-new PRs, human pushes, and fork PRs always get CI (fail-open).
Pure select_triggers / classify_sha_runs decision core added to traffic_control.py with
24 new unit tests (priority order, oldest-first fairness, inflight cap, fork skip, P0
bypass+preempt, head-SHA needy classification, and a liveness/anti-starvation simulation).
Closes#349
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add a detekt style>ForbiddenImport rule that forbids `import android.util.Log`
so all logging flows through org.libremail.reporting.AppLog, which mirrors each
line into the debug-report RingLogBuffer. A raw android.util.Log import writes to
Logcat only and never reaches a user-reviewed DebugReport (epic #324, strangler
final step).
Excludes the AppLog facade itself (the one sanctioned wrapper) and the unit tests
that mockkStatic(Log) to verify forwarding — AppLog forwards to Log, a throwing
stub under plain JVM unit tests, so those tests must mock it; they do not bypass
the facade.
Closes#331
Part of #324
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adds a fast traffic-control-tests job (ubuntu, actions/checkout +
actions/setup-python, no emulator/Gradle) that runs the 37 pure-stdlib
unit tests for .github/scripts/traffic_control.py on every PR, and
wires it into ci-passed's needs so a regression blocks merge instead
of only being caught locally.
Closes#346
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Restore (and extend to drafts) the old bash's broken-reclaim behaviour that the
initial Python refactor had dropped. runs_to_cancel now cancels an OTHER PR's
active/queued runs when EITHER:
(a) THIS PR is P0 and that PR is strictly-lower (reclaim every lower runner); OR
(b) that PR is broken/draft (effective priority 10) and THIS PR is strictly-higher
(effective priority < 10) — a wasted run any ready PR may reclaim.
P1-P9 still never bump a *normal* (non-broken/draft) lower run; a broken/draft PR
(P10) preempts nothing (nothing is strictly-lower than the bottom, and the
equal-or-higher invariant means a P10 never cancels another P10). Self / main-push /
equal-or-higher invariants unchanged.
Updates the module docstring + ci.yml comments (the "only P0 preempts" wording
becomes: P0 preempts everything strictly-lower; additionally, any strictly-higher PR
preempts a broken/draft run) and the job step/permission/needs comments. Adds unit
tests: P3 reclaims a broken P10 run and a draft P10 run; P3 does not bump a normal P5
run; a P10 self preempts nothing; plus an end-to-end P5-reclaims-draft-then-waits
scenario. 37 unit tests pass; ci.yml parses clean; --dry-run shows a P3 cancelling a
draft (and broken) run while still yielding to a higher P1.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>