You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GCP Dataflow运行本地正常管道时遇CalledProcessError(退出状态2)

Troubleshooting DATAFLOW CalledProcessError (Exit Status 2) When Running Pipeline on GCP

Exit status 2 from a subprocess call in Dataflow typically points to issues with how your script2.py is being accessed or executed in the worker environment. Let’s walk through the most common causes and actionable fixes:

1. script2.py isn’t accessible to Dataflow workers

When running locally, your script lives on your machine, but Dataflow workers run in isolated VMs—they can’t access your local files, and you can’t directly reference a GCS path in a subprocess call. Workers need a local copy of the script to execute it.

Fix: Use Beam’s FileSystems API to download the script from GCS to the worker’s temporary directory before running it:

import subprocess
from apache_beam.io.filesystems import FileSystems

def execute_script2(element):
    # Define your paths
    gcs_script_path = "gs://your-bucket/path/to/script2.py"
    local_script_path = "/tmp/script2.py"
    
    # Download from GCS to worker's local temp
    with FileSystems.open(gcs_script_path, 'r') as src_file, open(local_script_path, 'w') as dest_file:
        dest_file.write(src_file.read())
    
    # Make the script executable (optional but avoids permission issues)
    subprocess.run(["chmod", "+x", local_script_path], check=True)
    
    # Run the script using python3 (default on Dataflow workers)
    result = subprocess.run(
        ["python3", local_script_path, "your-arguments-here"],
        capture_output=True,
        text=True,
        check=True
    )
    return result.stdout

# Integrate this function into your pipeline
with beam.Pipeline(options=pipeline_options) as p:
    p | beam.Create([...]) | beam.Map(execute_script2)

2. Missing dependencies for script2.py

If script2.py uses libraries that aren’t pre-installed on the default Dataflow worker image (like pandas, requests, or custom modules), the subprocess call will fail with exit code 2 (usually accompanied by a "module not found" error in stderr).

Fix:

  • Create a requirements.txt file listing all dependencies for both test.py and script2.py.
  • Submit your pipeline with the --requirements_file flag to install these packages on workers:
    python test.py --runner=DataflowRunner --project=your-project-id --region=your-region --requirements_file=requirements.txt ...
    
  • For system-level dependencies (like libxml2), use a custom Docker container image tailored to your pipeline’s needs.

3. Incorrect command syntax or paths

Dataflow workers use python3 by default (not python), and relative paths can break in the isolated worker environment. Using python instead of python3 will trigger a "command not found" error (exit code 2).

Fix: Always use python3 (or the absolute path /usr/bin/python3) in your subprocess calls, and verify all arguments (especially file paths) are correctly formatted.

4. Permissions issues accessing GCS

If the Dataflow service account doesn’t have read access to the bucket hosting script2.py, the download step will fail, leading to a missing file error.

Fix:

  • Navigate to your GCS bucket’s permissions page in the GCP Console.
  • Add the Dataflow default service account (format: [your-project-number]@cloudservices.gserviceaccount.com) with the Storage Object Viewer role.
  • Alternatively, specify a custom service account with the necessary permissions when submitting the pipeline using the --service_account_email flag.

Critical Step: Check Dataflow Logs

To get the exact error details, go to your Dataflow job in the GCP Console, navigate to the Logs tab, and search for the stderr output from the subprocess call. This will tell you precisely why the command failed (e.g., "file not found", "missing module", etc.).

内容的提问来源于stack exchange,提问作者Guillem Puig comerma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:15:56