使用Great Expectations获取Spark DataFrame无效记录失败求助
问题排查与修复方案
1. 修正result_format参数结构
你定义的result_format存在嵌套错误,正确写法是将格式类型和额外参数放在同一层级,无需嵌套result_format键:
result_format = { "result_format": "COMPLETE", "include_unexpected_rows": True }
也可以直接拆分参数传入(更直观):
result = df_ge.expect_column_values_to_be_in_type_list( "page_title", [DateType], result_format="COMPLETE", include_unexpected_rows=True )
2. 修正数据类型引用方式
你传入的["DateType"]是字符串,Spark无法识别为有效类型,需导入pyspark的DateType类并直接引用:
from pyspark.sql.types import DateType
之后在expect方法中使用[DateType]替代字符串列表。
3. 正确查看无效记录
SparkDFDataset返回的unexpected_rows是Spark DataFrame对象,直接print整个result无法看到具体内容,需调用show()方法:
if not result["success"]: result["result"]["unexpected_rows"].show()
修复后的完整代码
import great_expectations as ge from great_expectations.dataset.sparkdf_dataset import SparkDFDataset from pyspark.sql.types import DateType df = spark.read.table("data_quality_test") df_ge = SparkDFDataset(df) result_format = { "result_format": "COMPLETE", "include_unexpected_rows": True } result = df_ge.expect_column_values_to_be_in_type_list( "page_title", [DateType], result_format=result_format ) # 查看违反规则的记录 if not result["success"]: result["result"]["unexpected_rows"].show()
额外注意事项
- 若需获取查询语句而非直接返回行,可使用
return_unexpected_index_query=True,对应查询语句会存于result["result"]["unexpected_index_query"]中。 - 确保Great Expectations为最新版本,旧版本可能存在参数支持不全的问题。
内容的提问来源于stack exchange,提问作者wander
相关产品推荐
相关产品推荐

