You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Parquet数据过滤异常求助:NOT条件下结果行数不符预期

问题原因与解决方法

问题根源

你的问题出在test_column包含NULL值。SQL遵循三值逻辑:NULL与任何值比较的结果都是UNKNOWN,而WHERE/filter子句只会保留结果为TRUE的行,因此NULL行在NOT test_column = 'xxx'或test_column != 'xxx'的过滤中会被排除。

计算缺失行数:175371 - 1 - 175362 = 8,这8行就是test_column为NULL的数据。


PySpark 解决方法

1. 使用filter API

df = spark.read.parquet("s3://test/parquet/date=2023-01-31")
# 同时包含不等于目标值和NULL的行
df.filter((df["test_column"] != "test_column_test") | df["test_column"].isNull()).count()

2. 使用where SQL表达式

df.where("test_column != 'test_column_test' OR test_column IS NULL").count()

Impala 解决方法

修改查询语句,显式包含NULL行:

SELECT count(*) FROM test.test_table 
WHERE date='2023-01-31' 
AND (test_column != 'test_column_test' OR test_column IS NULL);

或者更简洁的写法:

SELECT count(*) FROM test.test_table 
WHERE date='2023-01-31' 
AND NOT (test_column = 'test_column_test' AND test_column IS NOT NULL);

内容的提问来源于stack exchange,提问作者user19666823

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 08:20:42