PySpark删除空值超阈值行异常排查:为何仍有空值行留存?
dropna Behavior for Thresholded Null Removal Let's break down what's happening with your code and get this sorted out!
First, Clarify PySpark's dropna(thresh) Parameter
The thresh argument in dropna doesn't count the number of nulls allowed — instead, it defines the minimum number of non-null values required for a row to be kept.
When you run:
df3 = df3.dropna(thresh=len(df3.columns) - na_threshold)
With na_threshold=2, you're telling PySpark: "Keep any row that has at least (total_columns - 2) non-null values". This translates directly to: "Allow up to 2 nulls per row".
That’s why you’re still seeing a row with 1 null in df_null — that row is supposed to be retained under your current logic!
Why Increasing na_threshold Isn’t Fixing It
If you bump na_threshold to 3, you’re now allowing up to 3 nulls per row. Rows with 1 null still fall under that allowed limit, so they stay in the DataFrame. The only way those 1-null rows would get removed is if you set na_threshold=0 — which would require all columns to be non-null (since thresh=total_columns - 0 = total_columns).
If You Want to Delete Rows With Any Nulls (Not Just Over Threshold)
If your actual goal is to remove all rows that have even one null (instead of just rows with more than 2 nulls), simplify your code to:
df3 = df3.dropna() # Equivalent to dropna(how="any")
If You Truly Want to Delete Rows With Nulls Exceeding the Threshold
Wait, maybe you meant to delete rows where the number of nulls is greater than or equal to the threshold (e.g., delete rows with 2 or more nulls)? If so, adjust the thresh calculation:
# Allow at most (na_threshold - 1) nulls → require at least (total_columns - (na_threshold -1)) non-nulls df3 = df3.dropna(thresh=len(df3.columns) - na_threshold + 1)
For na_threshold=2, this becomes total_columns - 2 +1 = total_columns -1 — meaning rows need at least total_columns -1 non-nulls (i.e., max 1 null allowed). This would remove any row with 2 or more nulls, while keeping rows with 0 or 1 nulls.
Verify the Result
To check exactly how many nulls are in each remaining row (instead of just checking if any null exists), run this:
from pyspark.sql import functions as f from functools import reduce # Add a column counting nulls per row df3_with_null_count = df3.withColumn( "null_count", reduce(lambda x, y: x + y, [f.col(c).isNull().cast("int") for c in df3.columns]) ) # Show rows with nulls and their null count df3_with_null_count.filter(f.col("null_count") > 0).show()
This will let you confirm if your dropna logic is working as intended.
内容的提问来源于stack exchange,提问作者Tiger_Stripes

