You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Kedro集成Great Expectations执行ETL列验证时出现AttributeError问题求助,求端到端实现示例

Kedro集成Great Expectations执行ETL列验证时出现AttributeError问题求助,求端到端实现示例

嗨,我完全理解你现在被这个Kedro+Great Expectations的集成错误卡得有多烦躁——这个AttributeError: 'Datasource' object has no attribute 'get_batch'本质是版本不兼容导致的:新版本Great Expectations已经废弃了旧的get_batch方法,而你用的是Kedro官方文档里适配旧版GE的hooks代码,自然会报错。

下面我给你一套针对你的ecom_analytics项目的端到端可运行方案,一步步来解决问题:

一、先对齐环境依赖版本

首先确保你的requirements.txt里的版本是兼容的,建议用这个组合:

kedro~=0.18.12
great-expectations~=0.17.12

执行pip install -r requirements.txt更新依赖。

二、替换适配新版GE的hooks.py代码

删掉原来从Kedro文档复制的hooks代码,换成下面适配新版GE的实现:

from kedro.framework.hooks import hook_impl
from great_expectations.data_context import DataContext
from great_expectations.core.batch import BatchRequest

class GreatExpectationsHooks:
    def __init__(self):
        # 初始化GE上下文,指向项目根目录下的great_expectations文件夹
        self.context = DataContext(context_root_dir="./great_expectations")

    @hook_impl
    def after_dataset_loaded(self, dataset_name: str, data) -> None:
        # 只对需要验证的数据集触发验证,比如你的dataset_raw
        if dataset_name == "dataset_raw":
            # 1. 加载预先定义的期望套件
            expectation_suite = self.context.get_expectation_suite("data.raw")
            
            # 2. 构建批量请求(适配新版GE的方式)
            batch_request = BatchRequest(
                datasource_name="main_datasource",
                data_connector_name="default_inferred_data_connector_name",
                data_asset_name="dataset_raw",  # 对应data/01_raw下的dataset_raw.csv(不带后缀)
                batch_identifiers={"default_identifier_name": "default"}
            )
            
            # 3. 获取批量数据并执行验证
            batch_list = self.context.get_batch_list(batch_request=batch_request)
            validation_results = self.context.run_validation_operator(
                "action_list_operator",
                assets_to_validate=batch_list,
                expectation_suite_name="data.raw"
            )
            
            # 4. 如果验证失败,直接抛出异常终止Kedro流程
            if not validation_results["success"]:
                raise ValueError(f"数据集 {dataset_name} 未通过Great Expectations验证,请检查数据或期望规则!")

三、修正Great Expectations的数据源配置

打开great_expectations/great_expectations.yml,确保你的main_datasource配置能正确识别data/01_raw下的CSV文件:

datasources:
  main_datasource:
    class_name: Datasource
    module_name: great_expectations.datasource
    execution_engine:
      class_name: PandasExecutionEngine
      module_name: great_expectations.execution_engine
    data_connectors:
      default_inferred_data_connector_name:
        class_name: InferredAssetFilesystemDataConnector
        base_directory: ./data/01_raw  # 指向你的原始数据目录
        default_regex:
          pattern: (.*)\.csv
          group_names:
            - data_asset_name

四、确认Kedro Catalog配置

打开conf/base/catalog.yml,确保dataset_raw的配置正确:

dataset_raw:
  type: pandas.CSVDataSet
  filepath: data/01_raw/dataset_raw.csv
  load_args:
    sep: ','  # 根据你的CSV分隔符调整

五、注册hooks到Kedro

在src/ecom_analytics/settings.py里添加hooks的注册:

from ecom_analytics.hooks import GreatExpectationsHooks

HOOKS = [GreatExpectationsHooks()]

六、重新生成并验证期望套件

  1. 重新创建适配当前数据源的期望套件:
great_expectations suite new --suite-name data.raw --datasource main_datasource --data-asset-name dataset_raw
  1. 编辑套件添加你需要的列验证规则(比如非空、数据类型、取值范围等):
great_expectations suite edit data.raw
  1. 可以先单独用GE验证数据,确保套件能正常运行:
# 先创建一个checkpoint
great_expectations checkpoint new my_checkpoint
# 运行验证
great_expectations checkpoint run my_checkpoint

七、运行Kedro流程

最后执行kedro run,现在应该能正常触发Great Expectations的列验证,验证失败会终止流程,成功则继续执行ETL。

额外排查要点

  • 确保great_expectations文件夹在项目根目录下,hooks里的context_root_dir路径正确
  • data_asset_name必须和data/01_raw下的CSV文件名一致(不带.csv后缀)
  • 如果还是有问题,可以检查GE的日志文件great_expectations/uncommitted/logs/great_expectations.log,里面会有更详细的错误信息

备注:内容来源于stack exchange,提问作者Dhaval Thakkar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.23 07:44:10