You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Great Expectations中向验证器传入自定义Batch?

如何将自定义扁平化后的Batch与已有Great Expectations期望套件结合验证

我正在处理嵌套JSON数据,需要先扁平化才能操作。已经修改了Batch(新增了嵌套字段展开后的列),但不知道怎么把这个自定义Batch和指定的expectation suite一起传入验证器。

当前尝试的代码(扁平化部分已完成,但验证套件未关联):

# Import necessary libraries
import great_expectations as ge
import datetime

# Load the Great Expectations context
context = ge.data_context.DataContext("../.")

# Load the JSON data into a Pandas DataFrame
data_file_path = "../../data/nested.json"
df = ge.read_json(data_file_path)

# Create a batch of data
batch = ge.dataset.PandasDataset(df)

# Create new columns for nested values
batch["details_age"] = batch["details"].apply(lambda x: x.get("age"))
batch["details_address_city"] = batch["details"].apply(lambda x: x.get("address").get("city"))
batch["details_address_state"] = batch["details"].apply(lambda x: x.get("address").get("state"))

# Load the expectation suite
expectation_suite_name = 'nestedjson_expectations_suite'
suite = context.get_expectation_suite(expectation_suite_name)

# Validate the batch against the expectation suite
results = context.run_validation_operator(
    "action_list_operator", 
    assets_to_validate=[batch],
    run_name = "abcd1",
    run_time = datetime.datetime.now(datetime.timezone.utc),
)

# Print the validation results
print(results)

context.build_data_docs()
context.open_data_docs(resource_identifier=results.list_validation_result_identifiers()[0])

查看run_validation_operator文档后,知道assets_to_validate参数可接受Batch列表或(batch_kwargs, expectation_suite_name)元组,但不清楚如何搭配自定义Batch与已有期望套件进行验证。目前测试过直接在Batch上定义期望的方式,但希望复用已有的nestedjson_expectations_suite。


解决方案

有两种简单的方式可以实现自定义Batch与已有期望套件的绑定:

方法1:直接为自定义Batch加载期望套件

核心是给创建的PandasDataset对象设置expectation_suite属性,这样验证器会自动使用该套件对Batch进行校验。

修改后的完整代码:

# Import necessary libraries
import great_expectations as ge
import datetime

# Load the Great Expectations context
context = ge.data_context.DataContext("../.")

# Load the JSON data into a Pandas DataFrame
data_file_path = "../../data/nested.json"
df = ge.read_json(data_file_path)

# Create a batch of data
batch = ge.dataset.PandasDataset(df)

# 扁平化嵌套字段,新增列
batch["details_age"] = batch["details"].apply(lambda x: x.get("age"))
batch["details_address_city"] = batch["details"].apply(lambda x: x.get("address").get("city"))
batch["details_address_state"] = batch["details"].apply(lambda x: x.get("address").get("state"))

# 加载已有期望套件
expectation_suite_name = 'nestedjson_expectations_suite'
suite = context.get_expectation_suite(expectation_suite_name)

# *关键步骤*:将期望套件绑定到自定义Batch上
batch.expectation_suite = suite

# 验证Batch与套件
results = context.run_validation_operator(
    "action_list_operator", 
    assets_to_validate=[batch],
    run_name = "abcd1",
    run_time = datetime.datetime.now(datetime.timezone.utc),
)

# 输出结果并生成数据文档
print(results)
context.build_data_docs()
context.open_data_docs(resource_identifier=results.list_validation_result_identifiers()[0])

方法2:使用(batch, expectation_suite)元组传入

如果方法1不兼容你的Great Expectations版本,可以直接在assets_to_validate中传入包含Batch和套件的元组:

替换原验证部分的代码:

# 替换原有的run_validation_operator调用
results = context.run_validation_operator(
    "action_list_operator", 
    assets_to_validate=[(batch, suite)],  # 传入Batch与套件的元组
    run_name = "abcd1",
    run_time = datetime.datetime.now(datetime.timezone.utc),
)

注意:确保你的nestedjson_expectations_suite中已经包含了details_age、details_address_city等扁平化后字段的验证规则,否则这些字段的校验会被跳过或触发错误。

内容的提问来源于stack exchange,提问作者Shivam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 20:37:02