You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Great Expectations的MetricResolutionError列名称未赋值错误?

解决Great Expectations中expect_compound_columns_to_be_unique的MetricResolutionError问题

问题场景

在使用Great Expectations连接Impala数据源,调用validator.expect_compound_columns_to_be_unique(['column1', 'column2'])时,出现以下错误:

MetricResolutionError: Cannot compile Column object until its 'name' is assigned.

核心代码模板如下:

import datetime
import pandas as pd
import great_expectations as ge
import great_expectations.jupyter_ux
from great_expectations.core.batch import BatchRequest
from great_expectations.checkpoint import SimpleCheckpoint
from great_expectations.exceptions import DataContextError

context = ge.data_context.DataContext()

batch_request = {'datasource_name': 'impala_okh', 'data_connector_name': 'default_inferred_data_connector_name', 'data_asset_name': 'okh.okh_forecast_prod', 'limit': 1000}

expectation_suite_name = "okh_forecast_prod"
try:
    suite = context.get_expectation_suite(expectation_suite_name=expectation_suite_name)
    print(f'Loaded ExpectationSuite "{suite.expectation_suite_name}" containing {len(suite.expectations)} expectations.')
except DataContextError:
    suite = context.create_expectation_suite(expectation_suite_name=expectation_suite_name)
    print(f'Created ExpectationSuite "{suite.expectation_suite_name}".')

validator = context.get_validator(
    batch_request=BatchRequest(**batch_request),
    expectation_suite_name=expectation_suite_name
)
column_names = [f'"{column_name}"' for column_name in validator.columns()]
print(f"Columns: {', '.join(column_names)}.")
validator.head(n_rows=5, fetch_all=False)

解决方案

这个错误通常是因为Validator未完成列对象的初始化,或者列名匹配出现问题,可尝试以下几种方法解决:

1. 强制加载数据完成列初始化

在调用expect_compound_columns_to_be_unique前,执行validator.load_batch()或者修改head方法的fetch_all参数为True,强制Validator加载完整批次数据,确保列对象被正确初始化:

# 强制加载数据
validator.load_batch()
# 或者修改head方法
validator.head(n_rows=5, fetch_all=True)

# 之后再调用期望函数
validator.expect_compound_columns_to_be_unique(['column1', 'column2'])

2. 确认列名完全匹配

检查传递的列名是否与数据集实际列名完全一致(注意Impala列名是否区分大小写),可以通过打印validator.columns()的原始结果确认:

print("实际列名:", validator.columns())
# 确保['column1', 'column2']中的列名完全出现在上述结果中

3. 改用SQL查询指定列

如果使用自动推断的数据连接器存在元数据解析问题,可在batch request中直接指定SQL查询,明确获取需要的列,避免列对象初始化异常:

batch_request = {
    'datasource_name': 'impala_okh',
    'data_connector_name': 'default_inferred_data_connector_name',
    'query': 'SELECT column1, column2, 其他需要的列 FROM okh.okh_forecast_prod LIMIT 1000'
}

4. 检查数据源配置

确认Impala数据源的配置是否正确,特别是数据连接器的类型是否匹配,确保元数据(包括列名)能被正确读取。

内容的提问来源于stack exchange,提问作者Sevval Kahraman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 20:35:28