You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Great Expectations获取Spark DataFrame无效记录失败求助

问题排查与修复方案

1. 修正result_format参数结构

你定义的result_format存在嵌套错误,正确写法是将格式类型和额外参数放在同一层级,无需嵌套result_format键:

result_format = {
    "result_format": "COMPLETE",
    "include_unexpected_rows": True
}

也可以直接拆分参数传入(更直观):

result = df_ge.expect_column_values_to_be_in_type_list(
    "page_title", 
    [DateType],  
    result_format="COMPLETE",
    include_unexpected_rows=True
)

2. 修正数据类型引用方式

你传入的["DateType"]是字符串,Spark无法识别为有效类型,需导入pyspark的DateType类并直接引用:

from pyspark.sql.types import DateType

之后在expect方法中使用[DateType]替代字符串列表。

3. 正确查看无效记录

SparkDFDataset返回的unexpected_rows是Spark DataFrame对象,直接print整个result无法看到具体内容,需调用show()方法:

if not result["success"]:
    result["result"]["unexpected_rows"].show()

修复后的完整代码

import great_expectations as ge
from great_expectations.dataset.sparkdf_dataset import SparkDFDataset
from pyspark.sql.types import DateType

df = spark.read.table("data_quality_test")
df_ge = SparkDFDataset(df)

result_format = {
    "result_format": "COMPLETE",
    "include_unexpected_rows": True
}

result = df_ge.expect_column_values_to_be_in_type_list(
    "page_title", 
    [DateType], 
    result_format=result_format
)

# 查看违反规则的记录
if not result["success"]:
    result["result"]["unexpected_rows"].show()

额外注意事项

  • 若需获取查询语句而非直接返回行,可使用return_unexpected_index_query=True,对应查询语句会存于result["result"]["unexpected_index_query"]中。
  • 确保Great Expectations为最新版本,旧版本可能存在参数支持不全的问题。

内容的提问来源于stack exchange,提问作者wander

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 02:10:02