Kedro集成Great Expectations执行ETL列验证时出现AttributeError问题求助,求端到端实现示例
Kedro集成Great Expectations执行ETL列验证时出现AttributeError问题求助,求端到端实现示例
嗨,我完全理解你现在被这个Kedro+Great Expectations的集成错误卡得有多烦躁——这个AttributeError: 'Datasource' object has no attribute 'get_batch'本质是版本不兼容导致的:新版本Great Expectations已经废弃了旧的get_batch方法,而你用的是Kedro官方文档里适配旧版GE的hooks代码,自然会报错。
下面我给你一套针对你的ecom_analytics项目的端到端可运行方案,一步步来解决问题:
一、先对齐环境依赖版本
首先确保你的requirements.txt里的版本是兼容的,建议用这个组合:
kedro~=0.18.12 great-expectations~=0.17.12
执行pip install -r requirements.txt更新依赖。
二、替换适配新版GE的hooks.py代码
删掉原来从Kedro文档复制的hooks代码,换成下面适配新版GE的实现:
from kedro.framework.hooks import hook_impl from great_expectations.data_context import DataContext from great_expectations.core.batch import BatchRequest class GreatExpectationsHooks: def __init__(self): # 初始化GE上下文,指向项目根目录下的great_expectations文件夹 self.context = DataContext(context_root_dir="./great_expectations") @hook_impl def after_dataset_loaded(self, dataset_name: str, data) -> None: # 只对需要验证的数据集触发验证,比如你的dataset_raw if dataset_name == "dataset_raw": # 1. 加载预先定义的期望套件 expectation_suite = self.context.get_expectation_suite("data.raw") # 2. 构建批量请求(适配新版GE的方式) batch_request = BatchRequest( datasource_name="main_datasource", data_connector_name="default_inferred_data_connector_name", data_asset_name="dataset_raw", # 对应data/01_raw下的dataset_raw.csv(不带后缀) batch_identifiers={"default_identifier_name": "default"} ) # 3. 获取批量数据并执行验证 batch_list = self.context.get_batch_list(batch_request=batch_request) validation_results = self.context.run_validation_operator( "action_list_operator", assets_to_validate=batch_list, expectation_suite_name="data.raw" ) # 4. 如果验证失败,直接抛出异常终止Kedro流程 if not validation_results["success"]: raise ValueError(f"数据集 {dataset_name} 未通过Great Expectations验证,请检查数据或期望规则!")
三、修正Great Expectations的数据源配置
打开great_expectations/great_expectations.yml,确保你的main_datasource配置能正确识别data/01_raw下的CSV文件:
datasources: main_datasource: class_name: Datasource module_name: great_expectations.datasource execution_engine: class_name: PandasExecutionEngine module_name: great_expectations.execution_engine data_connectors: default_inferred_data_connector_name: class_name: InferredAssetFilesystemDataConnector base_directory: ./data/01_raw # 指向你的原始数据目录 default_regex: pattern: (.*)\.csv group_names: - data_asset_name
四、确认Kedro Catalog配置
打开conf/base/catalog.yml,确保dataset_raw的配置正确:
dataset_raw: type: pandas.CSVDataSet filepath: data/01_raw/dataset_raw.csv load_args: sep: ',' # 根据你的CSV分隔符调整
五、注册hooks到Kedro
在src/ecom_analytics/settings.py里添加hooks的注册:
from ecom_analytics.hooks import GreatExpectationsHooks HOOKS = [GreatExpectationsHooks()]
六、重新生成并验证期望套件
- 重新创建适配当前数据源的期望套件:
great_expectations suite new --suite-name data.raw --datasource main_datasource --data-asset-name dataset_raw
- 编辑套件添加你需要的列验证规则(比如非空、数据类型、取值范围等):
great_expectations suite edit data.raw
- 可以先单独用GE验证数据,确保套件能正常运行:
# 先创建一个checkpoint great_expectations checkpoint new my_checkpoint # 运行验证 great_expectations checkpoint run my_checkpoint
七、运行Kedro流程
最后执行kedro run,现在应该能正常触发Great Expectations的列验证,验证失败会终止流程,成功则继续执行ETL。
额外排查要点
- 确保
great_expectations文件夹在项目根目录下,hooks里的context_root_dir路径正确 data_asset_name必须和data/01_raw下的CSV文件名一致(不带.csv后缀)- 如果还是有问题,可以检查GE的日志文件
great_expectations/uncommitted/logs/great_expectations.log,里面会有更详细的错误信息
备注:内容来源于stack exchange,提问作者Dhaval Thakkar
相关产品推荐
相关产品推荐

