DEV Community

Cover image for Hunting Down a Zombie Python Thread: My First PR to Webots πŸ€–πŸŒΈ
Rin Katsuragi
Rin Katsuragi

Posted on

Hunting Down a Zombie Python Thread: My First PR to Webots πŸ€–πŸŒΈ

How a subtle threading deadlock on Windows led to my first open source contribution to the Webots robotics simulator.

Following up on my intro post from quiet Ikoma, and while hype still runs through me, I wanted to share a fun little milestone as I just submitted my first Pull Request to Webots!

If you work in robotics or embodied AI, you probably know Webots, it’s an incredible open-source 3D robot simulator used worldwide for prototyping everything from multi-jointed arms to autonomous rovers before touching physical hardware.

Here is the story of a sneaky deadlock on Windows, and how I solved it.


The Mystery of the Zombie Process πŸ§Ÿβ€β™‚οΈ

While working with external Python controllers on Windows, I ran into an issue where scripts wouldn't terminate.

When running a controller with stdout/stderr redirection enabled (--stdout-redirect):

from controller import Robot

robot = Robot()
print("Hello Webots!")
# simulation loop finishes...
Enter fullscreen mode Exit fullscreen mode

The simulation ended, the message printed cleanly to the Webots console, but python.exe never exited but stayed stuck in the background as a zombie process until forcefully killed. I also noticed there is an open issue that specified the exact error from the maintainer.

Why did that happen?


The Anatomy of the Deadlock πŸ”

Digging into Webots' Python API (robot.py), stdout redirection is handled by a class called StdStreamRedirect. It sets up an OS pipe (os.pipe()) and a background helper thread that reads lines from the pipe and forwards them directly to the Universal C Runtime (ucrtbase), allowing Webots' C++ backend to capture the output.

Here was the catch:

  1. Python shutdown order: When a Python script reaches EOF, Python's runtime (threading._shutdown) waits to join all non-daemon threads before executing atexit cleanup callbacks.
  2. The pipe blocking: The helper thread was created with daemon=False. It was blocked inside self._r.readline(), waiting for an EOF on the pipe.
  3. The stalemate: An OS pipe only sends EOF when all write ends are closed. But the write end was scheduled to close inside an atexit callback!

The interpreter waited for the thread to finish before running atexit, but the thread couldn't finish until atexit ran resulting to classic deadlock.


The Solution πŸ› οΈ

The fix was clean and minimal:

  1. Daemonize the helper thread (daemon=True): This allows Python to exit the main thread and proceed smoothly into the atexit stage.
  2. Explicit close() in atexit: We close the write pipe (sending EOF down the pipe), then call self._thread.join() to guarantee all pending prints are flushed to Webots before final termination.
  3. Restore streams: Restored sys.stdout and sys.stderr before closing to ensure any late interpreter shutdown messages don't raise ValueError: I/O operation on closed file.

Tested on Windows 11 β€” zero hangs, instant exit with code 0.


Small Steps Forward 🌱

Contributing to the core tools you rely on daily in your research is such a rewarding feeling.

If you've been hesitating to contribute to a large open-source project, start by tackling a small bug or reading through issue trackers. You might find a curious little puzzle waiting for you as I did!

What was your first open-source bug fix like? I'd love to hear your stories and get inspired to be not afraid to post about my journey in robotics and AI, even if my contributions look small and simple to advanced devs! πŸ΅πŸ’…

Top comments (3)

Collapse
 
reidmarlow profile image
Reid Marlow •

The shutdown ordering in Python catches so many people because threading._shutdown joins non-daemon threads before running atexit hooks. If the thread is waiting on an open pipe that only atexit closes, you get a clean deadlock every time.

Restoring sys.stdout before closing was a smart catch too. During late interpreter finalization, if Python tries to flush internal buffers to a closed file descriptor, it throws an unhandled exception right as the process dies.

My first open-source fix had a similar flavor, patching an unhandled SIGPIPE in a CLI wrapper that assumed the downstream pager stayed alive to take stdout.

Collapse
 
rinrinrinrin profile image
Rin Katsuragi •

Thank you so much and I couldn't agree more about how sneaky interpreter shutdown can be and also fixing an unhandled SIGPIPE in a pager must have been an interesting patch as well. What are you working on now?

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to