GCP Dataflow运行本地正常管道时遇CalledProcessError(退出状态2)
Exit status 2 from a subprocess call in Dataflow typically points to issues with how your script2.py is being accessed or executed in the worker environment. Let’s walk through the most common causes and actionable fixes:
1. script2.py isn’t accessible to Dataflow workers
When running locally, your script lives on your machine, but Dataflow workers run in isolated VMs—they can’t access your local files, and you can’t directly reference a GCS path in a subprocess call. Workers need a local copy of the script to execute it.
Fix: Use Beam’s FileSystems API to download the script from GCS to the worker’s temporary directory before running it:
import subprocess from apache_beam.io.filesystems import FileSystems def execute_script2(element): # Define your paths gcs_script_path = "gs://your-bucket/path/to/script2.py" local_script_path = "/tmp/script2.py" # Download from GCS to worker's local temp with FileSystems.open(gcs_script_path, 'r') as src_file, open(local_script_path, 'w') as dest_file: dest_file.write(src_file.read()) # Make the script executable (optional but avoids permission issues) subprocess.run(["chmod", "+x", local_script_path], check=True) # Run the script using python3 (default on Dataflow workers) result = subprocess.run( ["python3", local_script_path, "your-arguments-here"], capture_output=True, text=True, check=True ) return result.stdout # Integrate this function into your pipeline with beam.Pipeline(options=pipeline_options) as p: p | beam.Create([...]) | beam.Map(execute_script2)
2. Missing dependencies for script2.py
If script2.py uses libraries that aren’t pre-installed on the default Dataflow worker image (like pandas, requests, or custom modules), the subprocess call will fail with exit code 2 (usually accompanied by a "module not found" error in stderr).
Fix:
- Create a
requirements.txtfile listing all dependencies for bothtest.pyandscript2.py. - Submit your pipeline with the
--requirements_fileflag to install these packages on workers:python test.py --runner=DataflowRunner --project=your-project-id --region=your-region --requirements_file=requirements.txt ... - For system-level dependencies (like
libxml2), use a custom Docker container image tailored to your pipeline’s needs.
3. Incorrect command syntax or paths
Dataflow workers use python3 by default (not python), and relative paths can break in the isolated worker environment. Using python instead of python3 will trigger a "command not found" error (exit code 2).
Fix: Always use python3 (or the absolute path /usr/bin/python3) in your subprocess calls, and verify all arguments (especially file paths) are correctly formatted.
4. Permissions issues accessing GCS
If the Dataflow service account doesn’t have read access to the bucket hosting script2.py, the download step will fail, leading to a missing file error.
Fix:
- Navigate to your GCS bucket’s permissions page in the GCP Console.
- Add the Dataflow default service account (format:
[your-project-number]@cloudservices.gserviceaccount.com) with theStorage Object Viewerrole. - Alternatively, specify a custom service account with the necessary permissions when submitting the pipeline using the
--service_account_emailflag.
Critical Step: Check Dataflow Logs
To get the exact error details, go to your Dataflow job in the GCP Console, navigate to the Logs tab, and search for the stderr output from the subprocess call. This will tell you precisely why the command failed (e.g., "file not found", "missing module", etc.).
内容的提问来源于stack exchange,提问作者Guillem Puig comerma

