PySpark执行Join语句报错:'DataFrame' object is not callable 求解决
Let's break down what's causing this error and how to fix your code step by step.
Why You're Seeing This Error
The TypeError: 'DataFrame' object is not callable almost always happens when you accidentally treat a DataFrame like a function (e.g., calling df() instead of accessing a column with df.col). In your case, the most likely culprit is missing the import for PySpark's functions module while using F.isnull() and F.lit().
If you didn't import from pyspark.sql import functions as F, there's a chance that F is accidentally defined as a DataFrame elsewhere in your code (or environment), so when you try to call F.isnull(), you're actually trying to invoke a DataFrame as a function—hence the error.
Fixed & Optimized Code
First, add the critical import at the top of your script. Then, we can clean up your code to be more efficient (no need for unnecessary temp views and re-queries):
from pyspark.sql import functions as F # Read your CSV files directly into DataFrames (skip redundant temp views) df1 = spark.read.option("header","true").option("delimiter", ",").csv("s3://mybucket/file1.csv") df2 = spark.read.option("header","true").option("delimiter", ",").csv("s3://mybucket/file2.csv") # Option 1: Fix your original left join + filter logic df_d = df1.join(df2, df1.NO == df2.NO, "left") \ .filter(F.isnull(df2.NO)) \ .select(df1.NO, df1.NAME, df1.LAT, df1.LONG, F.lit('DELETE').alias('FLAG')) # Option 2: Use leftanti join (more efficient for your use case) # Leftanti join directly gets records present in df1 but not in df2, no need to filter nulls df_d = df1.join(df2, on="NO", how="leftanti") \ .select("NO", "NAME", "LAT", "LONG", F.lit('DELETE').alias('FLAG'))
Key Improvements
- Added the required functions import: This ensures
F.isnull()andF.lit()are referencing PySpark's built-in functions, not a stray DataFrame. - Removed redundant temp views: You don't need to create temp views and re-query them—just use the DataFrames you get directly from
spark.read.csv(). This saves unnecessary computation. - Optional leftanti join: This is a more efficient approach for your goal (finding records in df1 not present in df2) because it avoids bringing in columns from df2 entirely, unlike a left join followed by filtering nulls.
内容的提问来源于stack exchange,提问作者RK.

