You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure Databricks多输出列PyFunc模型MLflow评估Latency指标报错

解决MLflow评估多列输出PyFunc模型时的重复索引错误

问题根源

MLflow 2.8.0在处理返回多列DataFrame的PyFunc模型时,若输入数据或模型输出的DataFrame存在重复索引,结合latency、token_count等需要实时计算的额外指标时,会触发重索引操作,进而抛出ValueError: cannot reindex on an axis with duplicate labels异常。

解决方案

1. 确保输入与输出数据的索引唯一

首先检查并修复输入数据和模型输出的索引问题:

  • 重置输入数据的索引:
    # 处理输入数据,移除重复索引
    data = data.reset_index(drop=True)
    
  • 修改模型的predict方法,确保返回的DataFrame索引唯一:
    def predict(self, context, model_input):
        # 原有逻辑生成answers、sources、prompts列表
        answers = ...
        sources = ...
        prompts = ...
        
        # 构建DataFrame并重置索引
        result_df = pd.DataFrame({
            'answers': answers,
            'sources': sources,
            'prompts': prompts
        }).reset_index(drop=True)
        
        return result_df
    

2. 调整MLflow评估参数(适配多列输出)

修改评估代码,确保MLflow能正确识别多列输出,并避免索引冲突:

evaluation_results = mlflow.evaluate(
    model=f'models:/{model_name}/{model_version}',
    data=data.reset_index(drop=True),  # 确保输入无重复索引
    predictions="answers",
    extra_metrics=[
        mlflow.metrics.latency(),
        mlflow.metrics.token_count(predictions_col="answers")  # 明确指定token统计列
    ],
    evaluator_config={"log_model_explainability": False}  # 可选:关闭可解释性以避免额外索引操作
)

3. 手动计算并记录指标(兜底方案)

如果上述方法仍无法解决,可手动运行预测并记录所有指标,完全控制数据流程:

import mlflow
import time
from your_tokenizer_module import YourTokenizer  # 替换为实际的tokenizer

# 加载模型
model = mlflow.pyfunc.load_model(f'models:/{model_name}/{model_version}')

# 执行预测并计算延迟
start_time = time.perf_counter()
predictions_df = model.predict(data.reset_index(drop=True))
total_latency = time.perf_counter() - start_time
avg_latency = total_latency / len(predictions_df)

# 计算token数量
tokenizer = YourTokenizer()
predictions_df['token_count'] = predictions_df['answers'].apply(lambda x: len(tokenizer.encode(x)))
avg_token_count = predictions_df['token_count'].mean()

# 记录到MLflow
with mlflow.start_run(run_name=f"eval_model_v{model_version}"):
    # 记录数值指标
    mlflow.log_metric("avg_latency", avg_latency)
    mlflow.log_metric("avg_token_count", avg_token_count)
    # 记录审计用的完整预测结果(支持UI中查看对比)
    mlflow.log_table(predictions_df, artifact_file="full_predictions.csv")

4. 升级MLflow版本(推荐)

MLflow 2.8.0存在多列输出模型评估的索引处理bug,升级到2.10.0及以上版本可直接解决该问题,同时获得更完善的多列输出评估支持。

实现核心需求

  • 数值指标对比:latency和token_count会被记录为MLflow运行的指标,可在UI的"Metrics"标签页中直接对比不同模型版本的数值。
  • 审计数据查看:通过mlflow.log_table或模型输出的artifact,可在UI的"Artifacts"标签页中查看完整的answers、sources、prompts数据,支持下载对比。

内容的提问来源于stack exchange,提问作者gmatthews88

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 19:05:19