Skip to content

[BUG] intermittent hang in 26.10 nightly testing #1159

Description

@grlee77

Describe the bug

Steps/Code to reproduce bug
The following run took 3 attempts to pass the specific case:
conda-python-tests / 12.2.2, 3.13, amd64, rockylinux8, v100, earliest-driver, latest-deps

No failed test cases show up in the log, but with AI assisted reviews of the logs it indicates the specific case that is hanging:

TestBatchDecodingCUDA::test_batch_read_cuda_memory_cleanup (python/cucim/tests/unit/clara/test_batch_decoding.py:405)

In both failed attempts:

  • That test printed its start line but never reported PASSED or FAILED.
  • test_get_nocache was the next queued test and never started.
  • The other xdist workers continued running, so unrelated blob tests became the last visible results.

In attempt 3, the batch-decoding test passed in about 0.30 seconds, and the complete suite passed in 10m30s.

This confirms an intermittent concurrency/resource hang. The suspect test performs repeated CUDA batch decoding with two internal workers while eight pytest-xdist processes share one V100, and calls mempool.free_all_blocks() between iterations. A random seed would not help because neither this test nor the blob test uses random input.

If it recurs, I would focus on serializing the CUDA batch-decoding tests or running them outside the eight-worker xdist invocation, then investigate synchronization/resource cleanup in read_region(..., num_workers=2, device="cuda").

Expected behavior
In a passing run this test takes < 1 second.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't working

Type

Projects

  • Status
    No status

Relationships

None yet

Development

No branches or pull requests

Issue actions