The JVM unit test suite can deadlock in Room/WorkManager and hang forever #125
Closed
opened 2026-08-26 03:48:36 +00:00 by JMR-dev
·
5 comments
No Branch/Tag Specified
main
fix/102-picker-back-press-overshoot
fix/268-saf-picker-determinism
feat/ogg-vorbis-libvorbis
feat/expedited-conversion-work
fix/gate-cache-in-worktrees
chore/gate-runs-shellcheck
test/publish-delete-arm-real-provider
docs/e8-instrumented-coverage
docs/api37-carrier-count-drift
docs/e7-second-constraint
test/publish-to-a-real-saf-destination
fix/api37-report-match-line
fix/api37-task-snapshot-crash
test/join-failure-message-on-device
docs/e2e-read-findings-e7
test/cancelling-a-running-export
test/reattach-to-a-running-job
test/content-uri-reaches-ffmpeg
fix/launcher-wiring-waits-for-the-pick
fix/cancel-tests-need-a-slower-encode
test/cancelling-a-running-join
test/cancelling-a-running-session
test/notification-cancel-action
test/ffmpeg-progress-is-observed
test/fallback-asserts-the-path
test/flac-and-opus-assert-their-format
docs/e2e-read-findings
docs/wave4-coverage-numbers
fix/injectable-startup-sweep-scope
test/session-outcome-seam
test/launcher-callback-identity
test/theme-follows-system-dark
test/audio-drop-arm
fix/rotation-waits-for-recreation
fix/convert-guards-on-ready
test/retry-save-mime
test/hardware-progress-reaches-workmanager
test/ffprobe-mapping-seam
test/device-codec-enumeration-seam
test/unknown-container-row
test/null-message-fallbacks
test/cancel-reaches-workmanager
docs/coverage-wave3-recovery
test/concat-engine-seam
docs/coverage-wave3
test/mediaprobe-merge-seam
test/adaptive-shell-wiring
test/aac-audio-args
test/notification-progress-text
test/media3-muxer-guard
test/hardware-fallback-and-cancellation
test/unprobeable-join-clip
test/one-branch-outcomes
test/foreground-type-regimes
fix/bound-wedge-diagnostics
docs/coverage-wave2
test/screen-wiring
test/viewmodel-setters
test/join-state-mapping
test/conversion-state-mapping
test/dedupe-user-messages
fix/restore-stack-merges
test/refused-jobs
test/concatworker-failure-arms
test/container-capabilities-audio
test/readspec-enum-fallbacks
test/outputpublisher-seams
test/mediaprobe-track-seam
test/outputpublisher-partial-branches
test/fake-provider-scaffolding
docs/coverage-read-findings
chore/gitignore-kotlin
test/bound-the-hangs
docs/coverage-remeasure
ci/baseline-counter-precision
fix/invalid-suggestion-chip
ci/wedged-leg-report
fix/reattachment-overwrites-pick
test/theme-live-branches
fix/failed-save-retry
fix/empty-composition-crash
ci/advisory-failure-report
docs/seven-run-counts
test/release-permission-guard
ci/build-workflow-permissions
docs/api37-point-release
docs/benchmark-populate-path
fix/dead-assertion-probe-test
ci/actionlint
test/device-codecs-encode-consequence
fix/sdkmanager-pipefail
docs/readme-restart-claim
fix/saf-picker-root-discovery
fix/probe-dispatcher-seam
test/media3engine-mime-tables
test/mediaprobe-pure-helpers
fix/codec-vocabulary-drift
docs/robolectric-rationale-correction
docs/api37-advisory-counts
test/r38-8-saf-e2e
test/r38-7-join-states
test/r38-6-conversion-states
test/r38-5-state-seam
fix/jacoco-robolectric-coverage
docs/instrumented-tests-correction
test/r38-2-filecard
test/r38-4-advanced-picker
test/r38-3-pickers
tools/file-issue-script
tools/api-37-emulator
fix/review-app-gaps
docs/review-corrections
No results found.
Labels
Clear labels
above-cut
accessibility
backlog
bug
confirmed
documentation
duplicate
enhancement
good first issue
help wanted
invalid
plausible
question
sev:high
sev:low
sev:medium
wontfix
Worked autonomously overnight: local, JVM-verifiable, no product decision
Barrier affecting people with disabilities
Held for manual review: product/UX call, CI/workflow, hardware, or unverifiable here
Something isn't working
Reviewer demonstrated the defect
Improvements or additions to documentation
This issue or pull request already exists
New feature or request
Good for newcomers
Extra attention is needed
This doesn't seem right
Reviewer could not fully demonstrate it; treat as unproven
Further information is requested
High severity
Low severity
Medium severity
This will not be worked on
Milestone
No items
No Milestone
Projects
Clear projects
No projects
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: JMR-dev/LibreMediaConverter#125
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
:app:testDebugUnitTestcan deadlock and hang forever. Caught locally withjstackduring unrelated work on #118; reproduced on 1 of 2 runs. The suite hasno per-test timeout, so nothing breaks the hang — one local run sat in it for
47 minutes before being killed. On CI it would burn the job's 60-minute cap
and report as a timeout with no cause.
It is a real lock-order inversion, named by the JVM
Two library executors, each holding the lock the other wants:
SDK 36 Main Threaded9501d0—androidx.room.TransactionExecutored94cfe0—androidx.work.impl.utils.SerialExecutorImpl.execute:49DefaultDispatcher-worker-3ed94cfe0— insideSerialExecutorImpl$Task.run:96ed9501d0—androidx.room.TransactionExecutor.execute:34Both arrive through Room's
FlowUtil$createFlow→DBUtil.performSuspending→withContext, i.e. WorkManager's Room-backedWorkInfoflow being collected while WorkManager's own serial executor isdraining a task.
Where it enters our code
Nothing here is obviously wrong: the test drives a real
ConversionViewModelagainst a real
WorkManagerTestInitHelperand collects the WorkInfo flow, whichis exactly what the production code does. The inversion is between two androidx
libraries, so this is most likely not ours to fix directly — but it is ours
to stop it costing an hour.
Why this is filed rather than fixed
Two separable pieces, and the second is the valuable one:
workable but the JVM's own deadlock report already names both monitors, so
most of the diagnosis is done. Worth checking whether a newer
work-runtime/roompairing has fixed it before investing (both float onminor+patch here, so the pair can change without a commit).
turn "hangs until someone notices" into a normal failure with a stack — the
same argument as #122, which is a hang in the instrumented suite for an
unrelated reason. Two different hangs, one missing guardrail.
Do not respond by deleting or
@Ignore-ing the test. It is asserting realcleanup behaviour and it passes; the hang is a scheduling accident, not a fault
in what it checks.
Evidence
Full
jstackdump attached below as a gist-able file; the deadlock section andboth stacks are quoted above verbatim. Captured 2026-08-25 on
mainatb49295d, JDK 25, Robolectricsdk=36.It is not caused by any PR in flight — the run that hung was gating a
two-
.sh-plus-one-.mddiff, the same commit passed CI's Unit tests job, and thelocal retry passed 444/0.
Where Room actually comes from — and a trap for whoever tries the obvious fix
The deadlock's frames are all Room's coroutine path:
TransactionExecutor.android.kt,FlowUtil$createFlow,DBUtil.performSuspending. So "try a newer Room" is the natural first move.Two things make that harder than it looks, and one of them will waste your time.
1. Nothing in this project requests Room. It arrives only through WorkManager:
:app:dependencyInsightshows no other path.work = "2.+"resolves to 2.11.2, and 2.11.2 bringsRoom 2.7.0.
2. The catalog's Room entries are dead, and editing them does nothing.
gradle/libs.versions.tomlcarries
room = "2.+"plusandroidx-room-runtime,androidx-room-ktxandandroidx-room-compiler— and no build script references any of them.
grep -rn "libs\.androidx\.room"across the reporeturns nothing,
kspis not applied to the module (soroom-compilercould not run anyway), and nosource file imports
androidx.room.So bumping
room = "2.+"looks like it should raise Room and has no effect whatsoever. That is aconvincing dead end and it is sitting right where someone debugging this would step. Filed separately
as a hygiene item; noting it here because this ticket is where it bites.
The cheap experiment, for whoever picks this up
Room 2.8.4 is stable and published (2.7.0, 2.7.1, 2.7.2, 2.8.0 … 2.8.4 on Google Maven). The
lock-ordering bug is in code Room rewrote for coroutines in 2.7, so a fix landing in 2.8 is plausible
— not established.
The test is cheap and needs no device: force
androidx.room:room-runtimeto 2.8.x via a resolutionstrategy (not the catalog, per above), then run the repro. #127 landed a synthetic deadlock probe
technique — a two-monitor lock-order inversion — but for this bug the real reproduction is
ConversionViewModelCleanupTest."start over on a finished conversion deletes the staged file", whichhit it 1 run in 2 locally.
Two cautions: forcing a transitive version can break WorkManager, so the suite passing matters as
much as the deadlock not reproducing; and 1-in-2 means a handful of green runs is weak evidence —
the arithmetic on #49's comment applies here too.
Bounded, not fixed
#127 landed a 10-minute task timeout plus a jstack watchdog at 8 minutes, so this now ends in minutes
with a stack instead of burning the job's 30-minute cap anonymously. That is containment. This
ticket stays open for the cause.
Tried to reproduce it: 16 runs, zero hangs — and two things that narrow it
All on
main(dc8b7c3/099b7fd), JDK 25,--rerun-tasksevery time so nothing was served from cache.ConversionViewModelCleanupTestalone:app:testDebugUnitTest(454 tests)--continueThis does not clear the bug. There is a
jstacknaming both monitors; it happened. What the 16runs do is narrow where to look.
1. It is not internal to the test whose frame appears in the dump
Six isolated runs of
ConversionViewModelCleanupTestnever hang. The class named in the stack iswhere the deadlock was observed, not where it lives — which fits the mechanism, since a lock-order
inversion needs a second party already holding the other lock, and that party is another test. Do
not try to fix this by editing that test.
2. "1 in 2" is a point estimate from n=2, and should not be planned against
The original observation was 1 of 2 runs. That is a 50% point estimate with a confidence interval
spanning roughly 3-97%. Pooling with these 16 gives 1 in 18 ≈ 6%, at which zero-in-16 is an
unremarkable outcome (
0.94^16 ≈ 0.37). Nothing here is in tension; the first number was just verysmall.
Practical consequence: a fix cannot be validated by "it stopped happening" in any run count anyone
will sit through. That is the same arithmetic recorded on #49, and it is why the Room-version
experiment below needs a deterministic probe rather than repetition.
3. The one variable I could not reproduce: concurrent load
The original hang happened while three other agents were running full Gradle builds in parallel
worktrees. My runs were sequential on an otherwise-idle machine (
load average 8.34on 8 CPUs wasmy own build, not four competing ones). Thread scheduling under heavy contention is exactly what
decides whether a lock-order inversion actually interleaves, so that is the most likely missing
ingredient — and it is not something the suite controls.
If anyone wants a reliable repro, that is where I would start: several concurrent
:app:testDebugUnitTestinvocations, not more sequential ones.Status of the experiment I proposed above
Not run. Forcing Room 2.8.x only means something against a baseline that reproduces, and this
machine does not reproduce it today. Running it now would produce a green result that proves nothing
— precisely the failure mode this ticket's own arithmetic warns about. The dependency findings in my
previous comment stand and are the useful half; the version bump is worth trying only once someone
has a repro that fires.
Concurrent-load reproduction: 9 more runs, still zero. Total 25.
Last comment named concurrent load as the likeliest missing ingredient, so I tested it rather than
leaving it as a guess: three simultaneous
:app:testDebugUnitTestinvocations from three separateworktrees, three rounds.
No deadlock, no hang signal, in any of the nine. The contention was real — 3m40s against ~29s for
the same task run alone.
Running total across every configuration: 25 runs, 0 reproductions.
The two red builds are my harness, not this repo
No test failed (the XML carries zero failures), memory was not tight (15 GiB free), and it did not
recur in rounds 2 or 3. Three concurrent builds sharing one Gradle user home and daemon pool is not a
configuration this project ever runs in, so I am not filing it.
I explicitly ruled out #127's new jstack watchdog before saying that, since it is the newest code
touching the test worker and its author documented a misfire mode: it never fired (it dumps at 8
minutes; these runs were 1-4), wrote no report in any worktree, and is written to be incapable of
failing a build — "it reads a live process and writes a file. Nothing here kills, interrupts or
signals anything."
Honesty about how this evidence was produced
My first attempt at this experiment silently did not run.
wait $pidsdoes not word-split in zsh,so it returned immediately, the script exited, and all nine builds were killed ~2 s in. The results
file read
hang=0nine times and looked exactly like a clean negative result. I caught it onlybecause the logs were 183 bytes and no exit codes had been written.
Recording it because a null result is only worth as much as the proof the experiment ran, and this
one nearly went in without that proof.
Where that leaves the ticket
Still open, still real — there is a
jstacknaming both monitors. But it has now survived 25attempts across four configurations, including the one I predicted would catch it. Anyone picking
this up should assume it is rarer than the original 1-of-2 suggested and plan for a deterministic
probe rather than repetition. The dependency facts two comments up are the more useful half.
First CI occurrence since the watchdog landed — and it invalidates this ticket's own prediction
PR #214, run
34001741668, Unit tests leg, during the wave-4 merge train.This ticket says:
That is no longer what happens, and the difference is #118's jstack watchdog. The leg failed in
10m57s, and the log carries the attribution:
Found one Java-level deadlockis present in the dump, so this is this ticket and not a slow test.Two things worth recording.
ConversionViewModelNamingTest→
ConversionViewModel.observe→jobSnapshots, i.e. the Room/WorkManager pair this ticketalready identifies. The watchdog is doing exactly what it was added for, and the "no cause"
sentence above should be struck when this is next edited.
Room nor WorkManager, so this was a pure false signal on an unrelated PR — the same tax #190
documents for the emulator legs, now demonstrated on the JVM leg too.
Frequency datum: one occurrence in the ~14 full CI runs of this merge train.
Closing: both pieces are answered, and the answers point opposite ways
This ticket asked for two separable things and said the second was the valuable one. Taking them in
that order.
Piece 2 — "bound it" — is done, and has now been exercised in anger
The ticket was filed when "the JVM suite has no test timeout at all". It has one now: the
timeouton the
Testtasks plus the jstack watchdog beside it inapp/build.gradle.kts, withapp/src/test/java/org/libremediaconverter/ci/HangBoundTest.ktguarding both numbers.That guardrail met this deadlock on 2026-09-06, on PR #214's Unit tests leg (run
34001741668), anddid exactly what was asked:
BUILD FAILED in 10m57s, with the hung test named.So this ticket's own prediction is now false and should not be re-derived by the next reader:
It cost 11 minutes and reported the cause. Consider that sentence struck.
Piece 1 — "is it fixed by a newer pairing?" — no, and it is not reachable from our code
Still live at the current floating pair:
androidx.work:work-runtime 2.11.2,androidx.room 2.7.0.The occurrence above is on that pair, so the "worth checking whether a newer pairing has fixed it"
question is answered without needing a bisect.
And there is no seam on our side. Checked rather than assumed:
ConversionViewModel.observecollects onviewModelScope(ConversionViewModel.kt:551), i.e.Main.immediate, andJoinViewModel.kt:351is the same shape. That is theSDK 36 Main Threadof the deadlock report. There is no dispatcher parameter here and adding one would not be a
test-only change: the collect assigns the state the screen renders, so it has to resume on main.
The class already injects
cleanupDispatcherand the probe dispatcher where a seam waslegitimate; this is not one of those.
DefaultDispatcher-worker-NinsideSerialExecutorImpl$Task.run, is reached through Room's ownFlowUtil$createFlow→DBUtil.performSuspendinghop. That thread is not ours and is not configurable from here.installWorkManager(
StagingCleanupSupport.kt:123-133) sets bothsetExecutor(SynchronousExecutor())andsetTaskExecutor(SynchronousExecutor()). The inversion survives that, which is the strongestevidence that it is internal to the two libraries rather than a threading choice of ours.
So the diagnosis the ticket hoped for lands where the ticket guessed it would — "most likely not
ours to fix directly" — and the mitigation it actually wanted is in place.
Closing, with a condition
Closing as bounded rather than fixed. Reopen if the watchdog-attributed rate rises above what a
retry absorbs — one occurrence in the ~14 full CI runs of the wave-4 merge train is the only
frequency datum, and it is not a rate. If it becomes one, the next thing to try is pinning
work-runtime/roomto a pair that does not invert, which is a dependency experiment rather than acode change, and needs its own ticket because those two float on minor+patch here.
The instruction not to delete or
@Ignorethe affected tests still stands and nothing here changesit: they assert real behaviour and they pass.