Apache Beam Dataflow运行报错:NameError: name 'beam' is not defined
Below are the most common causes and corresponding fixes for your issue:
1. Import is Shadowed by a Local Variable
Check if any local variable, function parameter, or class in your code shares the name beam, which would overwrite the imported Apache Beam module. For example:
import apache_beam as beam def transform_data(beam): # Parameter shadows the imported module return beam.Row(id=1, value=element) # Throws NameError here
Fix: Rename the conflicting local identifier (e.g., change beam to input_element).
2. Import Not Accessible in Serialized Code
When using beam.Row inside a DoFn, lambda, or other code that gets serialized for Dataflow workers, ensure the apache_beam as beam import is present in the same module where the code is defined. If your DoFn lives in a separate file, add the import there too—workers need access to the import at runtime to resolve beam.Row.
3. Outdated Apache Beam Version
The beam.Row class was introduced in Apache Beam 2.16.0. If you're using an older version, the class won't exist, and you might see this error. Fix: Upgrade to a recent version:
pip install --upgrade apache-beam[gcp]
4. Lambda Serialization Issues
Lambdas can sometimes fail to capture imports properly during serialization. If you're using a lambda to create beam.Row, replace it with a named DoFn instead:
Before:
pipeline | beam.Map(lambda elem: beam.Row(key=elem[0], count=elem[1]))
After:
class CreateAggregatedRow(beam.DoFn): def process(self, element): key, count = element return [beam.Row(key=key, count=count)] pipeline | beam.ParDo(CreateAggregatedRow())
5. Incorrect Environment Setup
Ensure you're running your code in the virtual environment where Apache Beam is installed. If you have multiple Python environments, the import may not be available in the one you're using. Verify with:
pip show apache-beam
This should display the installed version and location—confirm it matches your active environment.
内容的提问来源于stack exchange,提问作者Ajay S Pal

