Files
LibreMail/.github/workflows/ci.yml
T
JMR-devandClaude Opus 4.8 05d06eb45b ci: traffic-controller owns CI triggering (GITHUB_TOKEN updates, priority-ordered dispatch)
End the merge cascade and give the traffic-controller ownership of CI *triggering*.

- autoupdate.yml updates PR branches with the built-in GITHUB_TOKEN instead of a PAT,
  so an update push no longer auto-retriggers CI (GitHub's anti-recursion rule) — the
  cascade (every merge re-runs every PR, cancel-in-progress thrashing them) is gone.
- New scheduler ci-trigger.yml -> traffic_control.py --mode trigger (re-)triggers CI
  for the highest-priority PR(s) whose head SHA has absent/stale checks, a few at a
  time (inflight cap), in the existing P0-P9 / broken-draft priority order — a
  poor-man's merge queue reusing the priority core. It runs after autoupdate finishes
  (workflow_run, race-free) plus a cron backstop plus manual dispatch.
- Triggering uses workflow_dispatch, which is EXEMPT from anti-recursion, so the
  built-in GITHUB_TOKEN (actions: write) starts the run — NO PAT / secret change needed.
- ci.yml gains a workflow_dispatch trigger (pr/head_sha/reason inputs) and a per-PR
  concurrency group unifying pull_request and dispatch runs; its on: pull_request path
  is kept so brand-new PRs, human pushes, and fork PRs always get CI (fail-open).

Pure select_triggers / classify_sha_runs decision core added to traffic_control.py with
24 new unit tests (priority order, oldest-first fairness, inflight cap, fork skip, P0
bypass+preempt, head-SHA needy classification, and a liveness/anti-starvation simulation).

Closes #349

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 01:27:40 -05:00

581 lines
29 KiB
YAML
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# SPDX-License-Identifier: GPL-3.0-or-later
name: CI
on:
pull_request:
branches: [main]
# The traffic-controller SCHEDULER (ci-trigger.yml, issue #349) (re-)triggers CI for a
# specific PR via this workflow_dispatch after a GITHUB_TOKEN auto-update has left the PR's
# head SHA with absent/stale checks. Dispatched on the PR's head BRANCH, so the run's checks
# land on the PR head SHA and satisfy branch protection. `on: pull_request` above is KEPT so
# brand-new PRs, human pushes, and fork PRs still get CI directly — this is the fail-open
# guarantee: CI is always triggerable even if the scheduler is broken or absent.
workflow_dispatch:
inputs:
pr:
description: "PR number this run is for (set by the traffic-controller scheduler)."
required: false
type: string
head_sha:
description: "Expected head SHA (informational, for traceability in the run log)."
required: false
type: string
reason:
description: "Why this run was dispatched (informational)."
required: false
type: string
# A new trigger for a PR cancels that PR's own in-flight run (a newer head SHA supersedes).
# The group is keyed to the PR NUMBER so a `pull_request` run and a scheduler
# `workflow_dispatch` run for the SAME PR share one concurrency group (either supersedes a
# stale run of the other); it falls back to the ref when no PR number is in context.
concurrency:
group: ci-pr-${{ github.event.pull_request.number || inputs.pr || github.ref }}
cancel-in-progress: true
permissions:
contents: read
env:
# SDK packages this project builds against (compileSdk 37 / build-tools 37.0.0).
# Quote the package ids when passed to sdkmanager — the ';' is a shell separator.
ANDROID_PLATFORM: "platforms;android-37.0"
ANDROID_BUILD_TOOLS: "build-tools;37.0.0"
jobs:
# ── Priority-based runner orchestration ─────────────────────────────
# Runs FIRST (the heavy jobs below all `needs: traffic-control`). It reads THIS
# PR's P0–P9 label, `broken` label, and draft state to order runner access. The
# decision logic lives in .github/scripts/traffic_control.py — a pure, unit-tested
# core (see .github/scripts/test_traffic_control.py) plus a thin gh-I/O shell; this
# step just checks out the repo and runs it.
#
# Effective priority: a `broken` OR `draft` PR => 10 (BOTTOM, below P9), overriding
# any P0–P9; else the lowest-numbered P0–P9 label present (P0 = highest); else P5.
#
# • P0 = EMERGENCY ONLY (app broken in production / emergency security update).
# P0 PREEMPTS: it cancels the in-progress / queued CI runs of ALL strictly-
# LOWER-priority OTHER open PRs to grab their runners immediately. A preempted
# PR simply re-runs on its next push / autoupdate rebase. P0 is the ONLY
# priority that preempts a *normal* lower run — P1–P9 never bump those (a
# higher PR may still reclaim a broken/draft lower run — see below).
#
# • P1–P9 = YIELD WITHOUT BUMPING a *normal* lower run. They do NOT cancel a
# normal lower-priority run already going — a higher-priority PR does not evict
# it, it just takes the next free slot (it MAY still reclaim a broken/draft
# lower run — see below). Mechanism: a bounded hold-back. This job defers (up to
# HOLD_BACK_BUDGET_SECONDS, kept well under timeout-minutes) while any strictly-
# higher-priority OTHER open PR still has an active/queued CI run, and — within
# its OWN priority level — while any peer is ordered ahead of it (an in-flight
# run keeps its place; then oldest createdAt first). It proceeds the moment it
# is at the front, or when the budget elapses (a PR never blocks itself).
#
# • `broken` / `draft` = BOTTOM (effective P10). Always yields, never preempts —
# and because its run is wasted (a broken PR can't merge; a draft isn't merge-
# ready), ANY higher-priority PR (not just P0) MAY cancel that run to reclaim
# its runner (still the strictly-lower rule: broken/draft is the bottom, so any
# ready PR outranks it). A maintainer marks a stuck/failing PR `broken` to drop
# it below everything so others aren't blocked behind it AND may reclaim its
# runner; a draft behaves the same until it is marked ready for review.
#
# Hard safety invariants, enforced in the script:
# • never cancels a run on main / a push event (the gh query filters
# --event pull_request and drops headBranch == main);
# • never cancels THIS PR's own run (skips self by PR number + run id);
# • never cancels an equal-or-higher-priority PR (only strictly-lower, prio > self);
# • P1–P9 never bump a *normal* lower run (they only reclaim broken/draft) —
# otherwise they just wait (bounded), then proceed.
#
# Honest limitation: GitHub Actions has no native priority queue and assigns
# runners roughly FIFO, so the hold-back is a BEST-EFFORT head-start, not a hard
# guarantee — under sustained contention the bounded wait can expire before a
# higher-priority PR drains. The waiting job also occupies a (cheap, short-lived)
# runner meanwhile, which is exactly why the wait is kept bounded.
#
# It is deliberately NOT a merge-gate check: it is absent from `ci-passed`'s
# needs, every API call is guarded, the script always exits 0, and the step is
# `continue-on-error` — so a hiccup (API error, missing permission, fork PR) can
# never fail or block CI. The heavy jobs only *order* after it via `needs`; if it
# were ever skipped/failed they'd be skipped, which `ci-passed` treats as a gate
# failure (fail-safe: blocks merge, never spuriously passes).
traffic-control:
name: Traffic control (runner priority)
runs-on: ubuntu-latest
timeout-minutes: 6 # hard backstop; the P1–P9 hold-back budget stays well under this
permissions:
contents: read # check out .github/scripts/traffic_control.py
actions: write # cancel lower-priority runs (P0 emergencies + broken/draft reclaim)
pull-requests: read # read PR P0–P9 labels + draft state
env:
GH_TOKEN: ${{ github.token }}
GH_REPO: ${{ github.repository }}
# On a `pull_request` run this is the PR number and the script does its full in-run
# runner-priority orchestration. On a scheduler `workflow_dispatch` run (issue #349)
# the event is not `pull_request`, so the script no-ops here (`--mode orchestrate`
# only acts on pull_request events) — priority was ALREADY applied at trigger time by
# ci-trigger.yml, so re-doing the in-run hold-back would just waste runner time. The
# `|| inputs.pr` keeps the number in the log for a dispatched run.
SELF_PR: ${{ github.event.pull_request.number || inputs.pr }}
# P1–P9 bounded hold-back knobs, read by traffic_control.py. BUDGET must stay
# comfortably below timeout-minutes so the poll loop always exits 0 before the
# hard job timeout fires — a timed-out job would skip the heavy jobs and fail
# `ci-passed`.
HOLD_BACK_BUDGET_SECONDS: "180"
HOLD_BACK_POLL_SECONDS: "15"
steps:
- name: Check out source
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
# gh is auto-configured from GH_TOKEN / GH_REPO; python3 is preinstalled on the
# runner. The script guards every API call and always exits 0 (belt-and-braces
# with continue-on-error), so it can never fail or block CI.
- name: Apply runner priority (P0/broken/draft preempt; P1–P9 hold back)
continue-on-error: true
run: python3 .github/scripts/traffic_control.py
# Fast, pure-stdlib-Python unit tests for the traffic-control decision core
# (.github/scripts/traffic_control.py / test_traffic_control.py — see the
# `traffic-control` job above). No emulator, no Gradle: this runs in seconds,
# independently of the Android jobs below, so a regression in the runner-priority
# logic fails fast and blocks merge via `ci-passed`.
traffic-control-tests:
name: Traffic-control unit tests
needs: traffic-control # order after runner-priority orchestration (P0/broken/draft preempt; P1–P9 hold-back)
runs-on: ubuntu-latest
steps:
- name: Check out source
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
- name: Set up Python
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6.3.0
with:
python-version: "3.x"
- name: Run traffic-controller unit tests
run: python -m unittest discover -s .github/scripts -p 'test_*.py' -v
debug-build:
name: Debug build
needs: traffic-control # order after runner-priority orchestration (P0/broken/draft preempt; P1–P9 hold-back)
# x86_64: Linux-arm64 runners can't set up this SDK — android-actions/setup-android's sdkmanager
# fails (exit 1) on the android-37.0 preview platform, and the emulator package has no arm64-Linux
# build. Build/unit-test results are host-arch-independent anyway (R8/AGP/JVM); real arm64
# device-ABI coverage would need arm64 emulators, which require macOS hosts.
runs-on: ubuntu-latest
steps:
- name: Check out source
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
- name: Set up JDK 21
uses: actions/setup-java@1bcf9fb12cf4aa7d266a90ae39939e61372fe520 # v5.4.0
with:
distribution: temurin
java-version: "21"
- name: Set up Android SDK
uses: android-actions/setup-android@40fd30fb8d7440372e1316f5d1809ec01dcd3699 # v4.0.1
- name: Install SDK platform and build-tools
run: sdkmanager "$ANDROID_PLATFORM" "$ANDROID_BUILD_TOOLS"
- name: Set up Gradle
uses: gradle/actions/setup-gradle@3f131e8634966bd73d06cc69884922b02e6faf92 # v6.2.0
- name: Assemble debug APK
run: ./gradlew assembleDebug --stacktrace
- name: Upload debug APK
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: debug-apk
path: app/build/outputs/apk/debug/*.apk
if-no-files-found: error
unit-tests:
name: Unit tests
needs: traffic-control # order after runner-priority orchestration (P0/broken/draft preempt; P1–P9 hold-back)
runs-on: ubuntu-latest
steps:
- name: Check out source
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
- name: Set up JDK 21
uses: actions/setup-java@1bcf9fb12cf4aa7d266a90ae39939e61372fe520 # v5.4.0
with:
distribution: temurin
java-version: "21"
- name: Set up Android SDK
uses: android-actions/setup-android@40fd30fb8d7440372e1316f5d1809ec01dcd3699 # v4.0.1
- name: Install SDK platform and build-tools
run: sdkmanager "$ANDROID_PLATFORM" "$ANDROID_BUILD_TOOLS"
- name: Set up Gradle
uses: gradle/actions/setup-gradle@3f131e8634966bd73d06cc69884922b02e6faf92 # v6.2.0
- name: Run unit tests
run: ./gradlew testDebugUnitTest --stacktrace
- name: Upload unit test report
if: ${{ !cancelled() }}
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: unit-test-report
path: app/build/reports/tests/testDebugUnitTest/
if-no-files-found: warn
# JaCoCo XML + HTML coverage for the JVM unit tests (issue #192), scoped to the JVM-testable
# surface (issues #290/#292). The report is generated and uploaded first, then a no-regression
# gate (issue #251) fails the job if overall LINE coverage drops below the floor pinned in
# app/build.gradle.kts. Kept in this unit-test job so it is part of the `CI passed` gate.
- name: Generate JaCoCo coverage report
if: ${{ !cancelled() }}
run: ./gradlew :app:jacocoTestReport --stacktrace
- name: Upload coverage report
if: ${{ !cancelled() }}
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: jacoco-coverage-report
path: app/build/reports/jacoco/jacocoTestReport/
if-no-files-found: warn
# No-regression coverage gate (issue #251): fails CI if scoped LINE coverage regresses below the
# floor pinned in app/build.gradle.kts. Runs after the upload so the HTML/XML report is always
# archived for triage even when this step goes red.
- name: Verify JaCoCo coverage (no-regression floor)
if: ${{ !cancelled() }}
run: ./gradlew :app:jacocoTestCoverageVerification --stacktrace
static-analysis:
name: Static analysis
needs: traffic-control # order after runner-priority orchestration (P0/broken/draft preempt; P1–P9 hold-back)
runs-on: ubuntu-latest
steps:
- name: Check out source
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
- name: Set up JDK 21
uses: actions/setup-java@1bcf9fb12cf4aa7d266a90ae39939e61372fe520 # v5.4.0
with:
distribution: temurin
java-version: "21"
# AGP configuration needs the SDK even for ktlint/detekt (they run on the :app module).
- name: Set up Android SDK
uses: android-actions/setup-android@40fd30fb8d7440372e1316f5d1809ec01dcd3699 # v4.0.1
- name: Install SDK platform and build-tools
run: sdkmanager "$ANDROID_PLATFORM" "$ANDROID_BUILD_TOOLS"
- name: Set up Gradle
uses: gradle/actions/setup-gradle@3f131e8634966bd73d06cc69884922b02e6faf92 # v6.2.0
# --continue so a ktlint failure still lets detekt report (and vice versa).
- name: Run ktlint and detekt
run: ./gradlew :app:ktlintCheck :app:detekt --continue --stacktrace
- name: Upload analysis reports
if: ${{ !cancelled() }}
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: static-analysis-reports
path: |
app/build/reports/ktlint/
app/build/reports/detekt/
if-no-files-found: warn
e2e:
name: E2E
needs: traffic-control # order after runner-priority orchestration (P0/broken/draft preempt; P1–P9 hold-back)
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
# Every Android API level across the rolling ~7-year support window: minSdk (29 / Android 10,
# 2019) through the latest stable. Each level boots its own emulator and runs the full
# instrumented + Compose UI (E2E) suite; all of them fan in to the "CI passed" gate. When a
# new Android ships, add it and drop the oldest level that has aged out of ~7 years. API 37
# (preview) is NOT in this matrix because emulator-runner can't provision its nonstandard
# android-37.0 / google_apis_ps16k image (it would wedge the gate) — it's covered separately
# by the custom-provisioned `e2e-preview` job below. Keep in sync with
# testOptions.managedDevices in app/build.gradle.kts.
api-level: [29, 30, 31, 32, 33, 34, 35, 36]
steps:
- name: Check out source
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
- name: Set up JDK 21
uses: actions/setup-java@1bcf9fb12cf4aa7d266a90ae39939e61372fe520 # v5.4.0
with:
distribution: temurin
java-version: "21"
- name: Set up Android SDK
uses: android-actions/setup-android@40fd30fb8d7440372e1316f5d1809ec01dcd3699 # v4.0.1
- name: Install SDK platform and build-tools
run: sdkmanager "$ANDROID_PLATFORM" "$ANDROID_BUILD_TOOLS"
- name: Set up Gradle
uses: gradle/actions/setup-gradle@3f131e8634966bd73d06cc69884922b02e6faf92 # v6.2.0
# The hardware-accelerated emulator needs KVM, which is gated behind a udev rule.
- name: Enable KVM
run: |
echo 'KERNEL=="kvm", GROUP="kvm", MODE="0666", OPTIONS+="static_node=kvm"' | sudo tee /etc/udev/rules.d/99-kvm4all.rules
sudo udevadm control --reload-rules
sudo udevadm trigger --name-match=kvm
- name: Cache AVD snapshot
uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9 # v6.1.0
id: avd-cache
with:
path: |
~/.android/avd/*
~/.android/adb*
key: avd-${{ matrix.api-level }}-google_apis-x86_64
# On a cache miss, cold-boot the emulator once so its snapshot can be cached,
# making subsequent runs start from a warm snapshot.
- name: Create AVD and generate snapshot for caching
if: steps.avd-cache.outputs.cache-hit != 'true'
uses: reactivecircus/android-emulator-runner@e89f39f1abbbd05b1113a29cf4db69e7540cae5a # v2.37.0
with:
api-level: ${{ matrix.api-level }}
target: google_apis
arch: x86_64
force-avd-creation: false
emulator-options: -no-window -gpu swiftshader_indirect -noaudio -no-boot-anim -camera-back none
disable-animations: false
script: echo "Generated AVD snapshot for caching."
# reactivecircus/android-emulator-runner runs an un-guarded, fatal `adb shell input keyevent 82`
# after boot. On snapshot resume that can race system_server (sys.boot_completed=1 before the
# `input` binder service is republished), aborting the job before Gradle runs with
# "No service published for: input" — an ~2%, API-29-only infra flake, not a test failure. Make
# the step non-fatal and retry once: two independent boots drop the race to ~0.04%. The definitive
# fix (adopt the e2e-preview job's manual-boot + `keyevent 82 || true`) is tracked separately.
- name: Run E2E tests
id: e2e
continue-on-error: true
uses: reactivecircus/android-emulator-runner@e89f39f1abbbd05b1113a29cf4db69e7540cae5a # v2.37.0
with:
api-level: ${{ matrix.api-level }}
target: google_apis
arch: x86_64
force-avd-creation: false
emulator-options: -no-snapshot-save -no-window -gpu swiftshader_indirect -noaudio -no-boot-anim -camera-back none
disable-animations: true
script: ./gradlew connectedDebugAndroidTest --stacktrace
- name: Run E2E tests (retry after emulator boot race)
if: steps.e2e.outcome == 'failure'
uses: reactivecircus/android-emulator-runner@e89f39f1abbbd05b1113a29cf4db69e7540cae5a # v2.37.0
with:
api-level: ${{ matrix.api-level }}
target: google_apis
arch: x86_64
force-avd-creation: false
emulator-options: -no-snapshot-save -no-window -gpu swiftshader_indirect -noaudio -no-boot-anim -camera-back none
disable-animations: true
script: ./gradlew connectedDebugAndroidTest --stacktrace
- name: Upload E2E test report
if: ${{ !cancelled() }}
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: e2e-test-report-api${{ matrix.api-level }}
path: app/build/reports/androidTests/connected/
if-no-files-found: warn
# API 37 (Android 17, preview) E2E. Its only system image is the nonstandard
# android-37.0 / google_apis_ps16k (16 KB page size), which reactivecircus/android-emulator-runner
# can't provision (it builds android-37 / google_apis, neither of which exists), so this job
# CUSTOM-PROVISIONS the emulator with sdkmanager/avdmanager/emulator directly. It is REQUIRED:
# part of the "CI passed" gate's needs (the preview emulator has proven stable in practice), so a
# genuine failure blocks merges. When a stable, emulator-runner-friendly API 37 image ships, fold
# 37 into the main `e2e` matrix and delete this job.
e2e-preview:
name: E2E (API 37 preview)
needs: traffic-control # order after runner-priority orchestration (P0/broken/draft preempt; P1–P9 hold-back)
runs-on: ubuntu-latest
timeout-minutes: 35
env:
API37_IMAGE: "system-images;android-37.0;google_apis_ps16k;x86_64"
steps:
- name: Check out source
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
- name: Set up JDK 21
uses: actions/setup-java@1bcf9fb12cf4aa7d266a90ae39939e61372fe520 # v5.4.0
with:
distribution: temurin
java-version: "21"
- name: Set up Android SDK
uses: android-actions/setup-android@40fd30fb8d7440372e1316f5d1809ec01dcd3699 # v4.0.1
- name: Set up Gradle
uses: gradle/actions/setup-gradle@3f131e8634966bd73d06cc69884922b02e6faf92 # v6.2.0
# The hardware-accelerated emulator needs KVM, which is gated behind a udev rule.
- name: Enable KVM
run: |
echo 'KERNEL=="kvm", GROUP="kvm", MODE="0666", OPTIONS+="static_node=kvm"' | sudo tee /etc/udev/rules.d/99-kvm4all.rules
sudo udevadm control --reload-rules
sudo udevadm trigger --name-match=kvm
# Cache the ~1 GB preview system image so only the first run pays the download.
- name: Cache API 37 system image
uses: actions/cache@55cc8345863c7cc4c66a329aec7e433d2d1c52a9 # v6.1.0
with:
# GitHub-hosted ubuntu runners install the SDK at /usr/local/lib/android/sdk; caching the
# image dir (with its package metadata) lets sdkmanager treat it as installed and skip the
# re-download on a cache hit.
path: /usr/local/lib/android/sdk/system-images/android-37.0
key: sysimg-android-37.0-google_apis_ps16k-x86_64
- name: Install SDK packages + preview system image
run: sdkmanager "$ANDROID_PLATFORM" "$ANDROID_BUILD_TOOLS" "platform-tools" "emulator" "$API37_IMAGE"
- name: Create API 37 AVD
run: |
# avdmanager and the emulator disagree on the default AVD dir when ANDROID_SDK_HOME is set
# on the runner (avdmanager writes $ANDROID_SDK_HOME/.android/avd; the emulator looks in
# $ANDROID_SDK_HOME/avd), which made the boot step report "Unknown AVD name [api37]". Pin
# ANDROID_AVD_HOME so both agree, and carry it to the boot step via $GITHUB_ENV.
export ANDROID_AVD_HOME="$HOME/.android/avd"
echo "ANDROID_AVD_HOME=$ANDROID_AVD_HOME" >> "$GITHUB_ENV"
mkdir -p "$ANDROID_AVD_HOME"
echo "no" | avdmanager create avd -n api37 -k "$API37_IMAGE" -d pixel_2 --force
echo "AVDs visible to the emulator:"; "$ANDROID_SDK_ROOT/emulator/emulator" -list-avds
- name: Boot emulator and run E2E
run: |
set -euo pipefail
EMU_LOG="${RUNNER_TEMP:-/tmp}/emulator.log"
LOGCAT_LOG="${RUNNER_TEMP:-/tmp}/logcat.txt"
DIAG_LOG="${RUNNER_TEMP:-/tmp}/boot-diagnostics.txt"
GPU_MODE="swiftshader_indirect"
# On a boot timeout, capture the full system state (accel/KVM/GPU/mem/disk/AVD config +
# emulator.log tail) into $DIAG_LOG for the artifact upload, then print a CONCISE summary
# (accel/KVM status + last 50 lines of emulator.log) to the step log so the cause is
# visible in the run output without downloading artifacts. Every probe is guarded (|| true)
# so a missing tool can't abort the retry under `set -e`.
dump_diagnostics() {
local attempt="$1" accel kvm
accel=$("$ANDROID_SDK_ROOT/emulator/emulator" -accel-check 2>&1) || true
kvm=$(ls -l /dev/kvm 2>&1) || true
{
echo "===== API 37 boot diagnostics (attempt $attempt) ====="
echo "--- adb devices ---"; adb devices 2>&1 || true
echo "--- emulator -accel-check ---"; echo "$accel"
echo "--- /dev/kvm ---"; echo "$kvm"
echo "--- GPU mode ---"; echo "$GPU_MODE"
echo "--- free memory ---"; free -h 2>&1 || true
echo "--- free disk ---"; df -h 2>&1 || true
echo "--- AVD config.ini ---"; cat "${ANDROID_AVD_HOME:-$HOME/.android/avd}/api37.avd/config.ini" 2>&1 || true
echo "--- emulator.log (tail 200) ---"; tail -200 "$EMU_LOG" 2>&1 || true
} >> "$DIAG_LOG" 2>&1 || true
echo "----- BOOT FAILURE SUMMARY (attempt $attempt) -----"
echo "accel-check: $accel"
echo "/dev/kvm: $kvm"
echo "--- emulator.log (tail 50) ---"; tail -50 "$EMU_LOG" 2>&1 || true
}
boot_emulator() {
echo "::group::Start API 37 emulator (attempt $1)"
# Capture the emulator's own output — without this a boot failure is invisible.
# -verbose -debug init,avd_config,kernel turns boot logging on by default so a boot flake
# is diagnosable from $EMU_LOG; DIAGNOSTICS ONLY — no boot-affecting flag is changed.
"$ANDROID_SDK_ROOT/emulator/emulator" -avd api37 \
-no-window -no-audio -no-boot-anim -no-snapshot -accel on \
-gpu "$GPU_MODE" -camera-back none -camera-front none \
-verbose -debug init,avd_config,kernel > "$EMU_LOG" 2>&1 &
# Stream logcat from the moment the device registers (wait-for-device blocks until then)
# into a file that survives to the artifact upload. Appended (with a header) per attempt.
echo "===== logcat (attempt $1) =====" >> "$LOGCAT_LOG"
adb wait-for-device logcat -v time >> "$LOGCAT_LOG" 2>&1 &
logcat_pid=$!
# ONE bounded wait covering both device registration and full boot, so a stuck emulator
# fails fast instead of hanging the whole job until the 35-min cap (the original bug).
if timeout 300 adb wait-for-device shell \
'while [ "$(getprop sys.boot_completed | tr -d "\r")" != "1" ]; do sleep 2; done'; then
echo "::endgroup::"; return 0
fi
echo "::endgroup::"
echo "::warning::API 37 emulator did not boot within 300s (attempt $1)"
dump_diagnostics "$1"
kill "$logcat_pid" 2>/dev/null || true
adb emu kill 2>/dev/null || true
sleep 5
return 1
}
booted=0
for attempt in 1 2; do boot_emulator "$attempt" && { booted=1; break; }; done
[ "$booted" = "1" ] || { echo "::error::API 37 preview emulator failed to boot after 2 attempts"; exit 1; }
adb shell input keyevent 82 || true
./gradlew connectedDebugAndroidTest --stacktrace
- name: Dump emulator log on failure
if: failure()
run: |
echo "--- emulator.log ---"; tail -200 "${RUNNER_TEMP:-/tmp}/emulator.log" 2>/dev/null || echo "(none)"
echo "--- logcat ---"; adb logcat -d 2>/dev/null | tail -120 || echo "(device unavailable)"
# Always upload the emulator log + captured logcat + boot-diagnostics dump so a boot flake
# (which can time out or be cancelled) is diagnosable from artifacts without a re-run.
- name: Upload emulator boot diagnostics
if: always()
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: e2e-api37-boot-diagnostics
path: |
${{ runner.temp }}/emulator.log
${{ runner.temp }}/logcat.txt
${{ runner.temp }}/boot-diagnostics.txt
if-no-files-found: warn
- name: Shut down emulator
if: always()
run: adb emu kill || true
- name: Upload E2E (API 37) report
if: ${{ !cancelled() }}
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: e2e-test-report-api37-preview
path: app/build/reports/androidTests/connected/
if-no-files-found: warn
# Single aggregating gate so branch protection can require ALL CI jobs with one stable status
# check. It depends on every job — including each api-level of the E2E matrix — so adding/removing
# a matrix level needs no change to branch protection (the per-"(api-level)" check names would
# otherwise have to be re-listed each time).
ci-passed:
name: CI passed
if: always()
# `traffic-control` is intentionally NOT listed here — it is a best-effort
# optimizer, not a merge requirement. But because the heavy jobs `needs:` it,
# a (should-never-happen) traffic-control failure would mark them 'skipped';
# treating 'skipped' as a gate failure below keeps that fail-safe (blocks the
# merge rather than letting it through untested).
needs: [traffic-control-tests, static-analysis, debug-build, unit-tests, e2e, e2e-preview]
runs-on: ubuntu-latest
steps:
- name: Verify every required job succeeded
if: ${{ contains(needs.*.result, 'failure') || contains(needs.*.result, 'cancelled') || contains(needs.*.result, 'skipped') }}
run: |
echo "Required CI jobs did not all succeed:"
echo " traffic-control-tests: ${{ needs.traffic-control-tests.result }}"
echo " static-analysis: ${{ needs.static-analysis.result }}"
echo " debug-build: ${{ needs.debug-build.result }}"
echo " unit-tests: ${{ needs.unit-tests.result }}"
echo " e2e: ${{ needs.e2e.result }}"
echo " e2e-preview: ${{ needs.e2e-preview.result }}"
exit 1