Parquet数据过滤异常求助:NOT条件下结果行数不符预期
问题原因与解决方法
问题根源
你的问题出在test_column包含NULL值。SQL遵循三值逻辑:NULL与任何值比较的结果都是UNKNOWN,而WHERE/filter子句只会保留结果为TRUE的行,因此NULL行在NOT test_column = 'xxx'或test_column != 'xxx'的过滤中会被排除。
计算缺失行数:175371 - 1 - 175362 = 8,这8行就是test_column为NULL的数据。
PySpark 解决方法
1. 使用filter API
df = spark.read.parquet("s3://test/parquet/date=2023-01-31") # 同时包含不等于目标值和NULL的行 df.filter((df["test_column"] != "test_column_test") | df["test_column"].isNull()).count()
2. 使用where SQL表达式
df.where("test_column != 'test_column_test' OR test_column IS NULL").count()
Impala 解决方法
修改查询语句,显式包含NULL行:
SELECT count(*) FROM test.test_table WHERE date='2023-01-31' AND (test_column != 'test_column_test' OR test_column IS NULL);
或者更简洁的写法:
SELECT count(*) FROM test.test_table WHERE date='2023-01-31' AND NOT (test_column = 'test_column_test' AND test_column IS NOT NULL);
内容的提问来源于stack exchange,提问作者user19666823
相关产品推荐
相关产品推荐

