如何在Great Expectations中向验证器传入自定义Batch?
如何将自定义扁平化后的Batch与已有Great Expectations期望套件结合验证
我正在处理嵌套JSON数据,需要先扁平化才能操作。已经修改了Batch(新增了嵌套字段展开后的列),但不知道怎么把这个自定义Batch和指定的expectation suite一起传入验证器。
当前尝试的代码(扁平化部分已完成,但验证套件未关联):
# Import necessary libraries import great_expectations as ge import datetime # Load the Great Expectations context context = ge.data_context.DataContext("../.") # Load the JSON data into a Pandas DataFrame data_file_path = "../../data/nested.json" df = ge.read_json(data_file_path) # Create a batch of data batch = ge.dataset.PandasDataset(df) # Create new columns for nested values batch["details_age"] = batch["details"].apply(lambda x: x.get("age")) batch["details_address_city"] = batch["details"].apply(lambda x: x.get("address").get("city")) batch["details_address_state"] = batch["details"].apply(lambda x: x.get("address").get("state")) # Load the expectation suite expectation_suite_name = 'nestedjson_expectations_suite' suite = context.get_expectation_suite(expectation_suite_name) # Validate the batch against the expectation suite results = context.run_validation_operator( "action_list_operator", assets_to_validate=[batch], run_name = "abcd1", run_time = datetime.datetime.now(datetime.timezone.utc), ) # Print the validation results print(results) context.build_data_docs() context.open_data_docs(resource_identifier=results.list_validation_result_identifiers()[0])
查看run_validation_operator文档后,知道assets_to_validate参数可接受Batch列表或(batch_kwargs, expectation_suite_name)元组,但不清楚如何搭配自定义Batch与已有期望套件进行验证。目前测试过直接在Batch上定义期望的方式,但希望复用已有的nestedjson_expectations_suite。
解决方案
有两种简单的方式可以实现自定义Batch与已有期望套件的绑定:
方法1:直接为自定义Batch加载期望套件
核心是给创建的PandasDataset对象设置expectation_suite属性,这样验证器会自动使用该套件对Batch进行校验。
修改后的完整代码:
# Import necessary libraries import great_expectations as ge import datetime # Load the Great Expectations context context = ge.data_context.DataContext("../.") # Load the JSON data into a Pandas DataFrame data_file_path = "../../data/nested.json" df = ge.read_json(data_file_path) # Create a batch of data batch = ge.dataset.PandasDataset(df) # 扁平化嵌套字段,新增列 batch["details_age"] = batch["details"].apply(lambda x: x.get("age")) batch["details_address_city"] = batch["details"].apply(lambda x: x.get("address").get("city")) batch["details_address_state"] = batch["details"].apply(lambda x: x.get("address").get("state")) # 加载已有期望套件 expectation_suite_name = 'nestedjson_expectations_suite' suite = context.get_expectation_suite(expectation_suite_name) # *关键步骤*:将期望套件绑定到自定义Batch上 batch.expectation_suite = suite # 验证Batch与套件 results = context.run_validation_operator( "action_list_operator", assets_to_validate=[batch], run_name = "abcd1", run_time = datetime.datetime.now(datetime.timezone.utc), ) # 输出结果并生成数据文档 print(results) context.build_data_docs() context.open_data_docs(resource_identifier=results.list_validation_result_identifiers()[0])
方法2:使用(batch, expectation_suite)元组传入
如果方法1不兼容你的Great Expectations版本,可以直接在assets_to_validate中传入包含Batch和套件的元组:
替换原验证部分的代码:
# 替换原有的run_validation_operator调用 results = context.run_validation_operator( "action_list_operator", assets_to_validate=[(batch, suite)], # 传入Batch与套件的元组 run_name = "abcd1", run_time = datetime.datetime.now(datetime.timezone.utc), )
注意:确保你的nestedjson_expectations_suite中已经包含了details_age、details_address_city等扁平化后字段的验证规则,否则这些字段的校验会被跳过或触发错误。
内容的提问来源于stack exchange,提问作者Shivam
相关产品推荐
相关产品推荐

