如何获取卡住的Dataflow Python进程的线程转储(threadz dump)?
Got it, when your Dataflow Python workers grind to a halt, grabbing a thread dump is one of the most effective ways to diagnose the root cause—whether it's a deadlock, an infinite loop, a blocked I/O call, or a stuck external API request. Here's a practical, step-by-step guide to get you the data you need:
1. Locate the Stuck Worker Node and SSH Into It
First, you need access to the problematic worker instance:
- Head to the Dataflow Jobs page in GCP Console, select your stuck job, then go to the Workers tab to find the stuck worker's instance ID.
- SSH into the worker using either the GCP Console's SSH button or the
gcloudcommand:gcloud compute ssh <WORKER_INSTANCE_ID> --zone <ZONE>
2. Identify the Stuck Python Process
Once you're on the worker, find the specific Python process tied to Dataflow:
- Use
toporhtopto spot processes with unusual CPU/memory behavior (stuck processes often show 0% CPU for deadlocks, or 100% CPU for infinite loops). - Narrow down to Dataflow-related Python processes with this command:
Look for the process ID (PID) of the worker process—this is the one you'll target for the thread dump.ps aux | grep python | grep dataflow
Note: Modern Dataflow workers run in containers, so if the above command doesn't show the process, you'll need to first enter the worker container:
# List running containers to find the Dataflow worker container docker ps | grep dataflow # Enter the container shell docker exec -it <CONTAINER_ID> /bin/bashThen repeat the
ps auxcommand inside the container to get the PID.
3. Capture the Thread Dump (3 Reliable Methods)
Choose one of these methods based on your setup and access:
Method 1: Use py-spy (Recommended for Detailed Output)
py-spy is a low-overhead profiling tool that generates clear thread dumps without needing to modify your code:
- Install it on the worker (or container):
sudo apt-get update && sudo apt-get install -y py-spy - Generate the thread dump to a file:
The output file will include every thread's call stack, making it easy to spot where execution is stuck.sudo py-spy dump --pid <TARGET_PID> --output thread_dump.txt
Method 2: Use Python's Built-in faulthandler
If your Dataflow job uses the official Python SDK, faulthandler is likely enabled by default. You can trigger a thread dump by sending a signal to the process:
- Send the
SIGUSR1signal to the target PID:
The thread dump will be printed to the worker's stderr, which you can find in GCP Logging (filter by the worker instance ID and log levelkill -SIGUSR1 <TARGET_PID>ERROR/DEBUG) or directly in the SSH session if you're attached to the process output.
Method 3: Use gdb for Low-Level Debugging
For deeper debugging, gdb can capture Python thread stacks:
- Install
gdband Python debugging symbols:sudo apt-get update && sudo apt-get install -y gdb python3-dbg - Attach
gdbto the target process:gdb -p <TARGET_PID> - Inside the
gdbshell, run this command to print all Python thread stacks:
For full stack details (including C-level frames), usepy-btpy-bt full.
Pro Tips
- Act fast: Dataflow workers may be replaced automatically if they're marked unhealthy, so capture the dump as soon as you spot the stuck worker.
- Enable debug mode: When launching your Dataflow job, add
--enable-debugto keep workers running longer after failures, giving you more time to debug.
内容的提问来源于stack exchange,提问作者Traeger Meyer

