The JVM unit test suite can deadlock in Room/WorkManager and hang forever #125

Closed
opened 2026-08-26 03:48:36 +00:00 by JMR-dev · 5 comments
JMR-dev commented 2026-08-26 03:48:36 +00:00 (Migrated from github.com)

:app:testDebugUnitTest can deadlock and hang forever. Caught locally with
jstack during unrelated work on #118; reproduced on 1 of 2 runs. The suite has
no per-test timeout, so nothing breaks the hang — one local run sat in it for
47 minutes before being killed. On CI it would burn the job's 60-minute cap
and report as a timeout with no cause.

It is a real lock-order inversion, named by the JVM

Found one Java-level deadlock:
=============================
"SDK 36 Main Thread @coroutine#1758":
  waiting to lock monitor 0x…3490 (object 0x…ed94cfe0, a java.lang.Object),
  which is held by "DefaultDispatcher-worker-3 @coroutine#1693"

"DefaultDispatcher-worker-3 @coroutine#1693":
  waiting to lock monitor 0x…3610 (object 0x…ed9501d0, a java.lang.Object),
  which is held by "SDK 36 Main Thread @coroutine#1758"

Two library executors, each holding the lock the other wants:

thread holds waits on
SDK 36 Main Thread ed9501d0 — androidx.room.TransactionExecutor ed94cfe0 — androidx.work.impl.utils.SerialExecutorImpl.execute:49
DefaultDispatcher-worker-3 ed94cfe0 — inside SerialExecutorImpl$Task.run:96 ed9501d0 — androidx.room.TransactionExecutor.execute:34

Both arrive through Room's FlowUtil$createFlow →
DBUtil.performSuspending → withContext, i.e. WorkManager's Room-backed
WorkInfo flow being collected while WorkManager's own serial executor is
draining a task.

Where it enters our code

ConversionViewModelCleanupTest."start over on a finished conversion deletes the staged file"(:56)
  -> convertedViewModel(:114)
    -> ConversionViewModel.convert(:372)
      -> ConversionViewModel.observe(:383)          // getWorkInfoByIdFlow(...).collect
        -> JobSnapshotsKt$jobSnapshots$2(:27)

Nothing here is obviously wrong: the test drives a real ConversionViewModel
against a real WorkManagerTestInitHelper and collects the WorkInfo flow, which
is exactly what the production code does. The inversion is between two androidx
libraries, so this is most likely not ours to fix directly — but it is ours
to stop it costing an hour.

Why this is filed rather than fixed

Two separable pieces, and the second is the valuable one:

  1. Diagnose the inversion — needs a reliable reproduction; 1-in-2 locally is
    workable but the JVM's own deadlock report already names both monitors, so
    most of the diagnosis is done. Worth checking whether a newer
    work-runtime/room pairing has fixed it before investing (both float on
    minor+patch here, so the pair can change without a commit).
  2. Bound it. The JVM suite has no test timeout at all. A global timeout would
    turn "hangs until someone notices" into a normal failure with a stack — the
    same argument as #122, which is a hang in the instrumented suite for an
    unrelated reason. Two different hangs, one missing guardrail.

Do not respond by deleting or @Ignore-ing the test. It is asserting real
cleanup behaviour and it passes; the hang is a scheduling accident, not a fault
in what it checks.

Evidence

Full jstack dump attached below as a gist-able file; the deadlock section and
both stacks are quoted above verbatim. Captured 2026-08-25 on main at
b49295d, JDK 25, Robolectric sdk=36.

It is not caused by any PR in flight — the run that hung was gating a
two-.sh-plus-one-.md diff, the same commit passed CI's Unit tests job, and the
local retry passed 444/0.

`:app:testDebugUnitTest` can deadlock and hang **forever**. Caught locally with `jstack` during unrelated work on #118; reproduced on 1 of 2 runs. The suite has no per-test timeout, so nothing breaks the hang — one local run sat in it for **47 minutes** before being killed. On CI it would burn the job's 60-minute cap and report as a timeout with no cause. ## It is a real lock-order inversion, named by the JVM ``` Found one Java-level deadlock: ============================= "SDK 36 Main Thread @coroutine#1758": waiting to lock monitor 0x…3490 (object 0x…ed94cfe0, a java.lang.Object), which is held by "DefaultDispatcher-worker-3 @coroutine#1693" "DefaultDispatcher-worker-3 @coroutine#1693": waiting to lock monitor 0x…3610 (object 0x…ed9501d0, a java.lang.Object), which is held by "SDK 36 Main Thread @coroutine#1758" ``` Two library executors, each holding the lock the other wants: | thread | holds | waits on | |---|---|---| | `SDK 36 Main Thread` | `ed9501d0` — `androidx.room.TransactionExecutor` | `ed94cfe0` — `androidx.work.impl.utils.SerialExecutorImpl.execute:49` | | `DefaultDispatcher-worker-3` | `ed94cfe0` — inside `SerialExecutorImpl$Task.run:96` | `ed9501d0` — `androidx.room.TransactionExecutor.execute:34` | Both arrive through Room's `FlowUtil$createFlow` → `DBUtil.performSuspending` → `withContext`, i.e. WorkManager's Room-backed `WorkInfo` flow being collected while WorkManager's own serial executor is draining a task. ## Where it enters our code ``` ConversionViewModelCleanupTest."start over on a finished conversion deletes the staged file"(:56) -> convertedViewModel(:114) -> ConversionViewModel.convert(:372) -> ConversionViewModel.observe(:383) // getWorkInfoByIdFlow(...).collect -> JobSnapshotsKt$jobSnapshots$2(:27) ``` Nothing here is obviously wrong: the test drives a real `ConversionViewModel` against a real `WorkManagerTestInitHelper` and collects the WorkInfo flow, which is exactly what the production code does. The inversion is between two androidx libraries, so this is most likely **not ours to fix directly** — but it is ours to stop it costing an hour. ## Why this is filed rather than fixed Two separable pieces, and the second is the valuable one: 1. **Diagnose the inversion** — needs a reliable reproduction; 1-in-2 locally is workable but the JVM's own deadlock report already names both monitors, so most of the diagnosis is done. Worth checking whether a newer `work-runtime`/`room` pairing has fixed it before investing (both float on minor+patch here, so the pair can change without a commit). 2. **Bound it.** The JVM suite has no test timeout at all. A global timeout would turn "hangs until someone notices" into a normal failure with a stack — the same argument as #122, which is a hang in the instrumented suite for an unrelated reason. Two different hangs, one missing guardrail. **Do not respond by deleting or `@Ignore`-ing the test.** It is asserting real cleanup behaviour and it passes; the hang is a scheduling accident, not a fault in what it checks. ## Evidence Full `jstack` dump attached below as a gist-able file; the deadlock section and both stacks are quoted above verbatim. Captured 2026-08-25 on `main` at `b49295d`, JDK 25, Robolectric `sdk=36`. **It is not caused by any PR in flight** — the run that hung was gating a two-`.sh`-plus-one-`.md` diff, the same commit passed CI's Unit tests job, and the local retry passed 444/0.
JMR-dev commented 2026-08-26 05:25:28 +00:00 (Migrated from github.com)

Where Room actually comes from — and a trap for whoever tries the obvious fix

The deadlock's frames are all Room's coroutine path: TransactionExecutor.android.kt,
FlowUtil$createFlow, DBUtil.performSuspending. So "try a newer Room" is the natural first move.
Two things make that harder than it looks, and one of them will waste your time.

1. Nothing in this project requests Room. It arrives only through WorkManager:

androidx.room:room-runtime:2.7.0
\--- androidx.work:work-runtime:2.11.2
     \--- androidx.work:work-runtime-ktx:2.11.2
          +--- debugRuntimeClasspath (requested androidx.work:work-runtime-ktx:2.+)

:app:dependencyInsight shows no other path. work = "2.+" resolves to 2.11.2, and 2.11.2 brings
Room 2.7.0.

2. The catalog's Room entries are dead, and editing them does nothing. gradle/libs.versions.toml
carries room = "2.+" plus androidx-room-runtime, androidx-room-ktx and androidx-room-compiler
— and no build script references any of them. grep -rn "libs\.androidx\.room" across the repo
returns nothing, ksp is not applied to the module (so room-compiler could not run anyway), and no
source file imports androidx.room.

So bumping room = "2.+" looks like it should raise Room and has no effect whatsoever. That is a
convincing dead end and it is sitting right where someone debugging this would step. Filed separately
as a hygiene item; noting it here because this ticket is where it bites.

The cheap experiment, for whoever picks this up

Room 2.8.4 is stable and published (2.7.0, 2.7.1, 2.7.2, 2.8.0 … 2.8.4 on Google Maven). The
lock-ordering bug is in code Room rewrote for coroutines in 2.7, so a fix landing in 2.8 is plausible
— not established.

The test is cheap and needs no device: force androidx.room:room-runtime to 2.8.x via a resolution
strategy (not the catalog, per above), then run the repro. #127 landed a synthetic deadlock probe
technique — a two-monitor lock-order inversion — but for this bug the real reproduction is
ConversionViewModelCleanupTest."start over on a finished conversion deletes the staged file", which
hit it 1 run in 2 locally.

Two cautions: forcing a transitive version can break WorkManager, so the suite passing matters as
much as the deadlock not reproducing; and 1-in-2 means a handful of green runs is weak evidence —
the arithmetic on #49's comment applies here too.

Bounded, not fixed

#127 landed a 10-minute task timeout plus a jstack watchdog at 8 minutes, so this now ends in minutes
with a stack instead of burning the job's 30-minute cap anonymously. That is containment. This
ticket stays open for the cause.

## Where Room actually comes from — and a trap for whoever tries the obvious fix The deadlock's frames are all Room's coroutine path: `TransactionExecutor.android.kt`, `FlowUtil$createFlow`, `DBUtil.performSuspending`. So "try a newer Room" is the natural first move. **Two things make that harder than it looks, and one of them will waste your time.** **1. Nothing in this project requests Room.** It arrives only through WorkManager: ``` androidx.room:room-runtime:2.7.0 \--- androidx.work:work-runtime:2.11.2 \--- androidx.work:work-runtime-ktx:2.11.2 +--- debugRuntimeClasspath (requested androidx.work:work-runtime-ktx:2.+) ``` `:app:dependencyInsight` shows no other path. `work = "2.+"` resolves to 2.11.2, and 2.11.2 brings Room **2.7.0**. **2. The catalog's Room entries are dead, and editing them does nothing.** `gradle/libs.versions.toml` carries `room = "2.+"` plus `androidx-room-runtime`, `androidx-room-ktx` and `androidx-room-compiler` — and **no build script references any of them**. `grep -rn "libs\.androidx\.room"` across the repo returns nothing, `ksp` is not applied to the module (so `room-compiler` could not run anyway), and no source file imports `androidx.room`. So bumping `room = "2.+"` looks like it should raise Room and **has no effect whatsoever**. That is a convincing dead end and it is sitting right where someone debugging this would step. Filed separately as a hygiene item; noting it here because this ticket is where it bites. ## The cheap experiment, for whoever picks this up **Room 2.8.4 is stable and published** (2.7.0, 2.7.1, 2.7.2, 2.8.0 … 2.8.4 on Google Maven). The lock-ordering bug is in code Room rewrote for coroutines in 2.7, so a fix landing in 2.8 is plausible — not established. The test is cheap and needs no device: force `androidx.room:room-runtime` to 2.8.x via a resolution strategy (**not** the catalog, per above), then run the repro. #127 landed a synthetic deadlock probe technique — a two-monitor lock-order inversion — but for *this* bug the real reproduction is `ConversionViewModelCleanupTest."start over on a finished conversion deletes the staged file"`, which hit it 1 run in 2 locally. Two cautions: forcing a transitive version can break WorkManager, so the suite passing matters as much as the deadlock not reproducing; and 1-in-2 means a handful of green runs is weak evidence — the arithmetic on #49's comment applies here too. ## Bounded, not fixed #127 landed a 10-minute task timeout plus a jstack watchdog at 8 minutes, so this now ends in minutes with a stack instead of burning the job's 30-minute cap anonymously. **That is containment.** This ticket stays open for the cause.
JMR-dev commented 2026-08-26 05:44:33 +00:00 (Migrated from github.com)

Tried to reproduce it: 16 runs, zero hangs — and two things that narrow it

All on main (dc8b7c3/099b7fd), JDK 25, --rerun-tasks every time so nothing was served from cache.

configuration runs hangs per-run
ConversionViewModelCleanupTest alone 6 0 ~13 s
full :app:testDebugUnitTest (454 tests) 6 0 ~29 s
the whole gate command, --continue 4 0 54-72 s

This does not clear the bug. There is a jstack naming both monitors; it happened. What the 16
runs do is narrow where to look.

1. It is not internal to the test whose frame appears in the dump

Six isolated runs of ConversionViewModelCleanupTest never hang. The class named in the stack is
where the deadlock was observed, not where it lives — which fits the mechanism, since a lock-order
inversion needs a second party already holding the other lock, and that party is another test. Do
not try to fix this by editing that test.

2. "1 in 2" is a point estimate from n=2, and should not be planned against

The original observation was 1 of 2 runs. That is a 50% point estimate with a confidence interval
spanning roughly 3-97%. Pooling with these 16 gives 1 in 18 ≈ 6%, at which zero-in-16 is an
unremarkable outcome (0.94^16 ≈ 0.37). Nothing here is in tension; the first number was just very
small.

Practical consequence: a fix cannot be validated by "it stopped happening" in any run count anyone
will sit through. That is the same arithmetic recorded on #49, and it is why the Room-version
experiment below needs a deterministic probe rather than repetition.

3. The one variable I could not reproduce: concurrent load

The original hang happened while three other agents were running full Gradle builds in parallel
worktrees
. My runs were sequential on an otherwise-idle machine (load average 8.34 on 8 CPUs was
my own build, not four competing ones). Thread scheduling under heavy contention is exactly what
decides whether a lock-order inversion actually interleaves, so that is the most likely missing
ingredient — and it is not something the suite controls.

If anyone wants a reliable repro, that is where I would start: several concurrent
:app:testDebugUnitTest invocations, not more sequential ones.

Status of the experiment I proposed above

Not run. Forcing Room 2.8.x only means something against a baseline that reproduces, and this
machine does not reproduce it today. Running it now would produce a green result that proves nothing
— precisely the failure mode this ticket's own arithmetic warns about. The dependency findings in my
previous comment stand and are the useful half; the version bump is worth trying only once someone
has a repro that fires.

## Tried to reproduce it: 16 runs, zero hangs — and two things that narrow it All on `main` (`dc8b7c3`/`099b7fd`), JDK 25, `--rerun-tasks` every time so nothing was served from cache. | configuration | runs | hangs | per-run | |---|---|---|---| | `ConversionViewModelCleanupTest` alone | 6 | **0** | ~13 s | | full `:app:testDebugUnitTest` (454 tests) | 6 | **0** | ~29 s | | the whole gate command, `--continue` | 4 | **0** | 54-72 s | **This does not clear the bug.** There is a `jstack` naming both monitors; it happened. What the 16 runs do is narrow where to look. ## 1. It is not internal to the test whose frame appears in the dump Six isolated runs of `ConversionViewModelCleanupTest` never hang. The class named in the stack is where the deadlock was *observed*, not where it lives — which fits the mechanism, since a lock-order inversion needs a second party already holding the other lock, and that party is another test. **Do not try to fix this by editing that test.** ## 2. "1 in 2" is a point estimate from n=2, and should not be planned against The original observation was 1 of 2 runs. That is a 50% point estimate with a confidence interval spanning roughly 3-97%. Pooling with these 16 gives **1 in 18 ≈ 6%**, at which zero-in-16 is an unremarkable outcome (`0.94^16 ≈ 0.37`). Nothing here is in tension; the first number was just very small. Practical consequence: a fix cannot be validated by "it stopped happening" in any run count anyone will sit through. That is the same arithmetic recorded on #49, and it is why the Room-version experiment below needs a *deterministic* probe rather than repetition. ## 3. The one variable I could not reproduce: concurrent load The original hang happened while **three other agents were running full Gradle builds in parallel worktrees**. My runs were sequential on an otherwise-idle machine (`load average 8.34` on 8 CPUs was my own build, not four competing ones). Thread scheduling under heavy contention is exactly what decides whether a lock-order inversion actually interleaves, so that is the most likely missing ingredient — and it is not something the suite controls. If anyone wants a reliable repro, that is where I would start: several concurrent `:app:testDebugUnitTest` invocations, not more sequential ones. ## Status of the experiment I proposed above **Not run.** Forcing Room 2.8.x only means something against a baseline that reproduces, and this machine does not reproduce it today. Running it now would produce a green result that proves nothing — precisely the failure mode this ticket's own arithmetic warns about. The dependency findings in my previous comment stand and are the useful half; the version bump is worth trying only once someone has a repro that fires.
JMR-dev commented 2026-08-26 06:12:22 +00:00 (Migrated from github.com)

Concurrent-load reproduction: 9 more runs, still zero. Total 25.

Last comment named concurrent load as the likeliest missing ingredient, so I tested it rather than
leaving it as a guess: three simultaneous :app:testDebugUnitTest invocations from three separate
worktrees, three rounds.

round wt1 wt2 wt3
1 BUILD FAILED 3m40s SUCCESS 3m42s BUILD FAILED 3m31s
2 SUCCESS 1m08s SUCCESS 1m09s SUCCESS 1m07s
3 SUCCESS 1m06s SUCCESS 1m07s SUCCESS 1m05s

No deadlock, no hang signal, in any of the nine. The contention was real — 3m40s against ~29s for
the same task run alone.

Running total across every configuration: 25 runs, 0 reproductions.

The two red builds are my harness, not this repo

java.nio.file.NoSuchFileException:
  .../app/build/test-results/testDebugUnitTest/binary/in-progress-results-generic.bin

No test failed (the XML carries zero failures), memory was not tight (15 GiB free), and it did not
recur in rounds 2 or 3. Three concurrent builds sharing one Gradle user home and daemon pool is not a
configuration this project ever runs in, so I am not filing it.

I explicitly ruled out #127's new jstack watchdog before saying that, since it is the newest code
touching the test worker and its author documented a misfire mode: it never fired (it dumps at 8
minutes; these runs were 1-4), wrote no report in any worktree, and is written to be incapable of
failing a build — "it reads a live process and writes a file. Nothing here kills, interrupts or
signals anything."

Honesty about how this evidence was produced

My first attempt at this experiment silently did not run. wait $pids does not word-split in zsh,
so it returned immediately, the script exited, and all nine builds were killed ~2 s in. The results
file read hang=0 nine times and looked exactly like a clean negative result. I caught it only
because the logs were 183 bytes and no exit codes had been written.

Recording it because a null result is only worth as much as the proof the experiment ran, and this
one nearly went in without that proof.

Where that leaves the ticket

Still open, still real — there is a jstack naming both monitors. But it has now survived 25
attempts across four configurations, including the one I predicted would catch it. Anyone picking
this up should assume it is rarer than the original 1-of-2 suggested and plan for a deterministic
probe rather than repetition. The dependency facts two comments up are the more useful half.

## Concurrent-load reproduction: 9 more runs, still zero. Total 25. Last comment named concurrent load as the likeliest missing ingredient, so I tested it rather than leaving it as a guess: **three simultaneous `:app:testDebugUnitTest` invocations from three separate worktrees, three rounds.** | round | wt1 | wt2 | wt3 | |---|---|---|---| | 1 | BUILD FAILED 3m40s | SUCCESS 3m42s | BUILD FAILED 3m31s | | 2 | SUCCESS 1m08s | SUCCESS 1m09s | SUCCESS 1m07s | | 3 | SUCCESS 1m06s | SUCCESS 1m07s | SUCCESS 1m05s | **No deadlock, no hang signal, in any of the nine.** The contention was real — 3m40s against ~29s for the same task run alone. Running total across every configuration: **25 runs, 0 reproductions.** ## The two red builds are my harness, not this repo ``` java.nio.file.NoSuchFileException: .../app/build/test-results/testDebugUnitTest/binary/in-progress-results-generic.bin ``` No test failed (the XML carries zero failures), memory was not tight (15 GiB free), and it did not recur in rounds 2 or 3. Three concurrent builds sharing one Gradle user home and daemon pool is not a configuration this project ever runs in, so I am **not** filing it. **I explicitly ruled out #127's new jstack watchdog** before saying that, since it is the newest code touching the test worker and its author documented a misfire mode: it never fired (it dumps at 8 minutes; these runs were 1-4), wrote no report in any worktree, and is written to be incapable of failing a build — *"it reads a live process and writes a file. Nothing here kills, interrupts or signals anything."* ## Honesty about how this evidence was produced My first attempt at this experiment **silently did not run**. `wait $pids` does not word-split in zsh, so it returned immediately, the script exited, and all nine builds were killed ~2 s in. The results file read `hang=0` nine times and looked exactly like a clean negative result. I caught it only because the logs were 183 bytes and no exit codes had been written. Recording it because a null result is only worth as much as the proof the experiment ran, and this one nearly went in without that proof. ## Where that leaves the ticket Still open, still real — there is a `jstack` naming both monitors. But it has now survived 25 attempts across four configurations, including the one I predicted would catch it. Anyone picking this up should assume it is **rarer than the original 1-of-2 suggested** and plan for a deterministic probe rather than repetition. The dependency facts two comments up are the more useful half.
JMR-dev commented 2026-09-06 01:11:12 +00:00 (Migrated from github.com)

First CI occurrence since the watchdog landed — and it invalidates this ticket's own prediction

PR #214, run 34001741668, Unit tests leg, during the wave-4 merge train.

This ticket says:

On CI it would burn the job's 60-minute cap and report as a timeout with no cause.

That is no longer what happens, and the difference is #118's jstack watchdog. The leg failed in
10m57s, and the log carries the attribution:

:app:testDebugUnitTest is still running after 8 minutes and is about to be timed out.
Thread dump of pid 2720 ... look for 'Found one Java-level deadlock' (that is #125).

"Test worker" #3 ... waiting on condition
    at org.libremediaconverter.convert.ConversionViewModel.observe(ConversionViewModel.kt:551)
    at org.libremediaconverter.convert.ConversionViewModel.convert(ConversionViewModel.kt:530)
    at org.libremediaconverter.convert.ConversionViewModelNamingTest.convertedViewModel(...:116)
    at ...ConversionViewModelNamingTest.and the saved name is the job's, not the picker's as it stands now(...:82)
    at org.libremediaconverter.work.JobSnapshotsKt$jobSnapshots$2.invokeSuspend(JobSnapshots.kt:27)

Found one Java-level deadlock is present in the dump, so this is this ticket and not a slow test.

Two things worth recording.

  1. The hung test is named, which the ticket assumed impossible on CI. ConversionViewModelNamingTest
    → ConversionViewModel.observe → jobSnapshots, i.e. the Room/WorkManager pair this ticket
    already identifies. The watchdog is doing exactly what it was added for, and the "no cause"
    sentence above should be struck when this is next edited.
  2. It cost a merge-train stop. #214's diff is a single new JVM test file that touches neither
    Room nor WorkManager, so this was a pure false signal on an unrelated PR — the same tax #190
    documents for the emulator legs, now demonstrated on the JVM leg too.

Frequency datum: one occurrence in the ~14 full CI runs of this merge train.

## First CI occurrence since the watchdog landed — and it invalidates this ticket's own prediction PR #214, run `34001741668`, **Unit tests** leg, during the wave-4 merge train. This ticket says: > On CI it would burn the job's 60-minute cap and report as a timeout with no cause. That is no longer what happens, and the difference is #118's jstack watchdog. The leg failed in **10m57s**, and the log carries the attribution: ``` :app:testDebugUnitTest is still running after 8 minutes and is about to be timed out. Thread dump of pid 2720 ... look for 'Found one Java-level deadlock' (that is #125). "Test worker" #3 ... waiting on condition at org.libremediaconverter.convert.ConversionViewModel.observe(ConversionViewModel.kt:551) at org.libremediaconverter.convert.ConversionViewModel.convert(ConversionViewModel.kt:530) at org.libremediaconverter.convert.ConversionViewModelNamingTest.convertedViewModel(...:116) at ...ConversionViewModelNamingTest.and the saved name is the job's, not the picker's as it stands now(...:82) at org.libremediaconverter.work.JobSnapshotsKt$jobSnapshots$2.invokeSuspend(JobSnapshots.kt:27) ``` `Found one Java-level deadlock` is present in the dump, so this is this ticket and not a slow test. **Two things worth recording.** 1. **The hung test is named**, which the ticket assumed impossible on CI. `ConversionViewModelNamingTest` → `ConversionViewModel.observe` → `jobSnapshots`, i.e. the Room/WorkManager pair this ticket already identifies. The watchdog is doing exactly what it was added for, and the "no cause" sentence above should be struck when this is next edited. 2. **It cost a merge-train stop.** #214's diff is a single new JVM test file that touches neither Room nor WorkManager, so this was a pure false signal on an unrelated PR — the same tax #190 documents for the emulator legs, now demonstrated on the JVM leg too. Frequency datum: one occurrence in the ~14 full CI runs of this merge train.
JMR-dev commented 2026-09-06 02:05:13 +00:00 (Migrated from github.com)

Closing: both pieces are answered, and the answers point opposite ways

This ticket asked for two separable things and said the second was the valuable one. Taking them in
that order.

Piece 2 — "bound it" — is done, and has now been exercised in anger

The ticket was filed when "the JVM suite has no test timeout at all". It has one now: the timeout
on the Test tasks plus the jstack watchdog beside it in app/build.gradle.kts, with
app/src/test/java/org/libremediaconverter/ci/HangBoundTest.kt guarding both numbers.

That guardrail met this deadlock on 2026-09-06, on PR #214's Unit tests leg (run 34001741668), and
did exactly what was asked:

:app:testDebugUnitTest is still running after 8 minutes and is about to be timed out.
Thread dump of pid 2720 ... look for 'Found one Java-level deadlock' (that is #125).

"Test worker" #3 ... waiting on condition
    at ConversionViewModel.observe(ConversionViewModel.kt:551)
    at ConversionViewModel.convert(ConversionViewModel.kt:530)
    at ConversionViewModelNamingTest.convertedViewModel(...:116)
    at JobSnapshotsKt$jobSnapshots$2.invokeSuspend(JobSnapshots.kt:27)

BUILD FAILED in 10m57s, with the hung test named.

So this ticket's own prediction is now false and should not be re-derived by the next reader:

On CI it would burn the job's 60-minute cap and report as a timeout with no cause.

It cost 11 minutes and reported the cause. Consider that sentence struck.

Piece 1 — "is it fixed by a newer pairing?" — no, and it is not reachable from our code

Still live at the current floating pair: androidx.work:work-runtime 2.11.2, androidx.room 2.7.0.
The occurrence above is on that pair, so the "worth checking whether a newer pairing has fixed it"
question is answered without needing a bisect.

And there is no seam on our side. Checked rather than assumed:

  • ConversionViewModel.observe collects on viewModelScope (ConversionViewModel.kt:551), i.e.
    Main.immediate, and JoinViewModel.kt:351 is the same shape. That is the SDK 36 Main Thread
    of the deadlock report. There is no dispatcher parameter here and adding one would not be a
    test-only change: the collect assigns the state the screen renders, so it has to resume on main.
    The class already injects cleanupDispatcher and the probe dispatcher where a seam was
    legitimate; this is not one of those.
  • The other side of the inversion, DefaultDispatcher-worker-N inside
    SerialExecutorImpl$Task.run, is reached through Room's own FlowUtil$createFlow →
    DBUtil.performSuspending hop. That thread is not ours and is not configurable from here.
  • The test fixture is already as serial as the API allows: installWorkManager
    (StagingCleanupSupport.kt:123-133) sets both setExecutor(SynchronousExecutor()) and
    setTaskExecutor(SynchronousExecutor()). The inversion survives that, which is the strongest
    evidence that it is internal to the two libraries rather than a threading choice of ours.

So the diagnosis the ticket hoped for lands where the ticket guessed it would — "most likely not
ours to fix directly
" — and the mitigation it actually wanted is in place.

Closing, with a condition

Closing as bounded rather than fixed. Reopen if the watchdog-attributed rate rises above what a
retry absorbs
— one occurrence in the ~14 full CI runs of the wave-4 merge train is the only
frequency datum, and it is not a rate. If it becomes one, the next thing to try is pinning
work-runtime/room to a pair that does not invert, which is a dependency experiment rather than a
code change, and needs its own ticket because those two float on minor+patch here.

The instruction not to delete or @Ignore the affected tests still stands and nothing here changes
it: they assert real behaviour and they pass.

## Closing: both pieces are answered, and the answers point opposite ways This ticket asked for two separable things and said the second was the valuable one. Taking them in that order. ### Piece 2 — "bound it" — is done, and has now been exercised in anger The ticket was filed when "the JVM suite has no test timeout at all". It has one now: the `timeout` on the `Test` tasks plus the jstack watchdog beside it in `app/build.gradle.kts`, with `app/src/test/java/org/libremediaconverter/ci/HangBoundTest.kt` guarding both numbers. That guardrail met this deadlock on 2026-09-06, on PR #214's Unit tests leg (run `34001741668`), and did exactly what was asked: ``` :app:testDebugUnitTest is still running after 8 minutes and is about to be timed out. Thread dump of pid 2720 ... look for 'Found one Java-level deadlock' (that is #125). "Test worker" #3 ... waiting on condition at ConversionViewModel.observe(ConversionViewModel.kt:551) at ConversionViewModel.convert(ConversionViewModel.kt:530) at ConversionViewModelNamingTest.convertedViewModel(...:116) at JobSnapshotsKt$jobSnapshots$2.invokeSuspend(JobSnapshots.kt:27) ``` `BUILD FAILED in 10m57s`, with the hung test named. **So this ticket's own prediction is now false and should not be re-derived by the next reader:** > On CI it would burn the job's 60-minute cap and report as a timeout with no cause. It cost 11 minutes and reported the cause. Consider that sentence struck. ### Piece 1 — "is it fixed by a newer pairing?" — no, and it is not reachable from our code **Still live at the current floating pair: `androidx.work:work-runtime 2.11.2`, `androidx.room 2.7.0`.** The occurrence above is on that pair, so the "worth checking whether a newer pairing has fixed it" question is answered without needing a bisect. **And there is no seam on our side.** Checked rather than assumed: - `ConversionViewModel.observe` collects on `viewModelScope` (`ConversionViewModel.kt:551`), i.e. `Main.immediate`, and `JoinViewModel.kt:351` is the same shape. That is the `SDK 36 Main Thread` of the deadlock report. There is no dispatcher parameter here and adding one would not be a test-only change: the collect assigns the state the screen renders, so it has to resume on main. The class already injects `cleanupDispatcher` and the probe dispatcher where a seam was legitimate; this is not one of those. - The other side of the inversion, `DefaultDispatcher-worker-N` inside `SerialExecutorImpl$Task.run`, is reached through Room's own `FlowUtil$createFlow` → `DBUtil.performSuspending` hop. That thread is not ours and is not configurable from here. - The test fixture is already as serial as the API allows: `installWorkManager` (`StagingCleanupSupport.kt:123-133`) sets **both** `setExecutor(SynchronousExecutor())` and `setTaskExecutor(SynchronousExecutor())`. The inversion survives that, which is the strongest evidence that it is internal to the two libraries rather than a threading choice of ours. So the diagnosis the ticket hoped for lands where the ticket guessed it would — "most likely **not ours to fix directly**" — and the mitigation it actually wanted is in place. ### Closing, with a condition Closing as bounded rather than fixed. **Reopen if the watchdog-attributed rate rises above what a retry absorbs** — one occurrence in the ~14 full CI runs of the wave-4 merge train is the only frequency datum, and it is not a rate. If it becomes one, the next thing to try is pinning `work-runtime`/`room` to a pair that does not invert, which is a dependency experiment rather than a code change, and needs its own ticket because those two float on minor+patch here. The instruction not to delete or `@Ignore` the affected tests still stands and nothing here changes it: they assert real behaviour and they pass.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: JMR-dev/LibreMediaConverter#125