You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取卡住的Dataflow Python进程的线程转储(threadz dump)?

How to Identify and Capture Thread Dumps for Stuck Dataflow Python Worker Processes

Got it, when your Dataflow Python workers grind to a halt, grabbing a thread dump is one of the most effective ways to diagnose the root cause—whether it's a deadlock, an infinite loop, a blocked I/O call, or a stuck external API request. Here's a practical, step-by-step guide to get you the data you need:

1. Locate the Stuck Worker Node and SSH Into It

First, you need access to the problematic worker instance:

  • Head to the Dataflow Jobs page in GCP Console, select your stuck job, then go to the Workers tab to find the stuck worker's instance ID.
  • SSH into the worker using either the GCP Console's SSH button or the gcloud command:
    gcloud compute ssh <WORKER_INSTANCE_ID> --zone <ZONE>
    

2. Identify the Stuck Python Process

Once you're on the worker, find the specific Python process tied to Dataflow:

  • Use top or htop to spot processes with unusual CPU/memory behavior (stuck processes often show 0% CPU for deadlocks, or 100% CPU for infinite loops).
  • Narrow down to Dataflow-related Python processes with this command:
    ps aux | grep python | grep dataflow
    
    Look for the process ID (PID) of the worker process—this is the one you'll target for the thread dump.

Note: Modern Dataflow workers run in containers, so if the above command doesn't show the process, you'll need to first enter the worker container:

# List running containers to find the Dataflow worker container
docker ps | grep dataflow
# Enter the container shell
docker exec -it <CONTAINER_ID> /bin/bash

Then repeat the ps aux command inside the container to get the PID.

3. Capture the Thread Dump (3 Reliable Methods)

Choose one of these methods based on your setup and access:

py-spy is a low-overhead profiling tool that generates clear thread dumps without needing to modify your code:

  • Install it on the worker (or container):
    sudo apt-get update && sudo apt-get install -y py-spy
    
  • Generate the thread dump to a file:
    sudo py-spy dump --pid <TARGET_PID> --output thread_dump.txt
    
    The output file will include every thread's call stack, making it easy to spot where execution is stuck.

Method 2: Use Python's Built-in faulthandler

If your Dataflow job uses the official Python SDK, faulthandler is likely enabled by default. You can trigger a thread dump by sending a signal to the process:

  • Send the SIGUSR1 signal to the target PID:
    kill -SIGUSR1 <TARGET_PID>
    
    The thread dump will be printed to the worker's stderr, which you can find in GCP Logging (filter by the worker instance ID and log level ERROR/DEBUG) or directly in the SSH session if you're attached to the process output.

Method 3: Use gdb for Low-Level Debugging

For deeper debugging, gdb can capture Python thread stacks:

  • Install gdb and Python debugging symbols:
    sudo apt-get update && sudo apt-get install -y gdb python3-dbg
    
  • Attach gdb to the target process:
    gdb -p <TARGET_PID>
    
  • Inside the gdb shell, run this command to print all Python thread stacks:
    py-bt
    
    For full stack details (including C-level frames), use py-bt full.

Pro Tips

  • Act fast: Dataflow workers may be replaced automatically if they're marked unhealthy, so capture the dump as soon as you spot the stuck worker.
  • Enable debug mode: When launching your Dataflow job, add --enable-debug to keep workers running longer after failures, giving you more time to debug.

内容的提问来源于stack exchange,提问作者Traeger Meyer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:02:09