Skip to content

Fix hang after rapid continue with multiple threads - #2081

Merged
Rich Chiodo (rchiodo) merged 1 commit into
microsoft:mainfrom
Om-singhaI:fix/rapid-continue-hang
Oct 8, 2026
Merged

Rich Chiodo (rchiodo) merged 1 commit into
microsoft:mainfrom
Om-singhaI:fix/rapid-continue-hang

Conversation

@Om-singhaI

Copy link
Copy Markdown
Contributor

Fixes #2048.

In single notification mode (the debugpy default), on_continue_request sends the continue response only when the next resume notification goes out. Two cases never get one: a continue that arrives after the last stop was already reported as resumed (the second of two quick F5 presses), and a continue whose stop is replaced by another thread's breakpoint before any suspended thread leaves its wait loop. The adapter waits for that response inside its message loop, so it stops answering every later request: the frozen session in the issue.

The pydevd log in the issue shows the first case. Right after continue (seq 122) is processed, the next thread to leave its wait loop logs Resume not sent (it was already sent). Last resume 18 >= Last suspend 18. The resume for that stop had already gone out, so this continue had nothing left to wait for.

What changed, all in pydevd.py:

  • The on resumed callbacks move from ThreadsSuspendedSingleNotification to AbstractSingleNotificationBehavior and use its _lock, so the check below can't race a resume.
  • add_on_resumed_callback calls the callback right away when the last suspend notification was already followed by a resume notification.
  • When callbacks are pending and a new suspend notification is about to go out, the resume notification is sent first, so the client gets continued, the continue response, then stopped.

A continue for a stop not yet reported as resumed still waits, as the comment in on_continue_request asks.

New tests: test_single_notification_5 and _6 cover the two cases, and test_continue_while_running sends a second continue through the adapter while the debuggee runs.

Testing on macOS 26.6.2 with PYDEVD_USE_CYTHON=NO:

  • test_single_notification.py: 6 passed on CPython 3.10.6 and 3.13.15, and 20 repeated runs passed on 3.10.6.
  • tests/debugpy/test_threads.py on 3.10.6: 9 passed. test_continue_while_running with DEBUGPY_TESTS_FULL=1: 14 passed.
  • 12 related pydevd tests in test_debugger_json.py and test_debugger.py (suspend notification, continue, pause, logpoints and others): all passed.
  • On main with only the tests added, test_continue_while_running hangs on the second continue until the 90 second timeout. The two unit tests fail with AttributeError, or with main's callback handling patched in, assert [1] == [1, 2] and assert 'suspend' == 'resume'.
  • A raw DAP client against a variant of the issue's script, 7 runs with one or two quick continue requests per stop: all 7 hang on main and all finish with this change, on 3.10.6 and 3.13.15.
  • ruff check . passes.

The adapter still blocks its message loop on the continue response; I can look at that separately.

In single notification mode the continue response waits for the next
resume notification. A continue that arrives after the last stop was
already reported as resumed, or one whose stop is replaced by another
thread's breakpoint before any suspended thread leaves its wait loop,
never got one, so the adapter blocked on it and stopped answering every
later request.

Keep the callbacks in AbstractSingleNotificationBehavior under its lock,
call them right away when no suspend notification is waiting for a
resume, and send the pending resume before a new suspend notification.
@Om-singhaI
om singhal (Om-singhaI) requested a review from a team as a code owner October 7, 2026 19:23
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
There may be pipelines that require an authorized user to comment /azp run to run.

@rchiodo

Copy link
Copy Markdown
Contributor

/azp run

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 1 pipeline(s).

@heejaechang

Heejae Chang (heejaechang) commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

🔒 Automated review in progress — Heejae Chang (@heejaechang) is auto-reviewing this PR.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved via Review Center.

@heejaechang Heejae Chang (heejaechang) added the review-auto:approved Automated review: no blocking findings (approval posted). label Oct 7, 2026

@rchiodo Rich Chiodo (rchiodo) left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved via Review Center.

@rchiodo

Copy link
Copy Markdown
Contributor

/azp run

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 1 pipeline(s).

@Om-singhaI

Copy link
Copy Markdown
Contributor Author

The failures in both runs are Windows tests that also fail on main without this change (test_thread_count and test_systemexit in build 4660, test_unsupported_configure_from_environment in build 4648), and test_continue_while_running passed on every job. Let me know if you would like me to look into any of them.

@rchiodo
Rich Chiodo (rchiodo) merged commit 6ee39d2 into microsoft:main Oct 8, 2026
24 of 26 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

review-auto:approved Automated review: no blocking findings (approval posted).

Projects

None yet

Development

Successfully merging this pull request may close these issues.

#debugpy becomes unresponsive after rapid breakpoint suspend/resume cycles in multithreaded Python programs

3 participants