如何解决Great Expectations的MetricResolutionError列名称未赋值错误?
解决Great Expectations中expect_compound_columns_to_be_unique的MetricResolutionError问题
问题场景
在使用Great Expectations连接Impala数据源,调用validator.expect_compound_columns_to_be_unique(['column1', 'column2'])时,出现以下错误:
MetricResolutionError: Cannot compile Column object until its 'name' is assigned.
核心代码模板如下:
import datetime import pandas as pd import great_expectations as ge import great_expectations.jupyter_ux from great_expectations.core.batch import BatchRequest from great_expectations.checkpoint import SimpleCheckpoint from great_expectations.exceptions import DataContextError context = ge.data_context.DataContext() batch_request = {'datasource_name': 'impala_okh', 'data_connector_name': 'default_inferred_data_connector_name', 'data_asset_name': 'okh.okh_forecast_prod', 'limit': 1000} expectation_suite_name = "okh_forecast_prod" try: suite = context.get_expectation_suite(expectation_suite_name=expectation_suite_name) print(f'Loaded ExpectationSuite "{suite.expectation_suite_name}" containing {len(suite.expectations)} expectations.') except DataContextError: suite = context.create_expectation_suite(expectation_suite_name=expectation_suite_name) print(f'Created ExpectationSuite "{suite.expectation_suite_name}".') validator = context.get_validator( batch_request=BatchRequest(**batch_request), expectation_suite_name=expectation_suite_name ) column_names = [f'"{column_name}"' for column_name in validator.columns()] print(f"Columns: {', '.join(column_names)}.") validator.head(n_rows=5, fetch_all=False)
解决方案
这个错误通常是因为Validator未完成列对象的初始化,或者列名匹配出现问题,可尝试以下几种方法解决:
1. 强制加载数据完成列初始化
在调用expect_compound_columns_to_be_unique前,执行validator.load_batch()或者修改head方法的fetch_all参数为True,强制Validator加载完整批次数据,确保列对象被正确初始化:
# 强制加载数据 validator.load_batch() # 或者修改head方法 validator.head(n_rows=5, fetch_all=True) # 之后再调用期望函数 validator.expect_compound_columns_to_be_unique(['column1', 'column2'])
2. 确认列名完全匹配
检查传递的列名是否与数据集实际列名完全一致(注意Impala列名是否区分大小写),可以通过打印validator.columns()的原始结果确认:
print("实际列名:", validator.columns()) # 确保['column1', 'column2']中的列名完全出现在上述结果中
3. 改用SQL查询指定列
如果使用自动推断的数据连接器存在元数据解析问题,可在batch request中直接指定SQL查询,明确获取需要的列,避免列对象初始化异常:
batch_request = { 'datasource_name': 'impala_okh', 'data_connector_name': 'default_inferred_data_connector_name', 'query': 'SELECT column1, column2, 其他需要的列 FROM okh.okh_forecast_prod LIMIT 1000' }
4. 检查数据源配置
确认Impala数据源的配置是否正确,特别是数据连接器的类型是否匹配,确保元数据(包括列名)能被正确读取。
内容的提问来源于stack exchange,提问作者Sevval Kahraman
相关产品推荐
相关产品推荐

