Azure Databricks多输出列PyFunc模型MLflow评估Latency指标报错
解决MLflow评估多列输出PyFunc模型时的重复索引错误
问题根源
MLflow 2.8.0在处理返回多列DataFrame的PyFunc模型时,若输入数据或模型输出的DataFrame存在重复索引,结合latency、token_count等需要实时计算的额外指标时,会触发重索引操作,进而抛出ValueError: cannot reindex on an axis with duplicate labels异常。
解决方案
1. 确保输入与输出数据的索引唯一
首先检查并修复输入数据和模型输出的索引问题:
- 重置输入数据的索引:
# 处理输入数据,移除重复索引 data = data.reset_index(drop=True) - 修改模型的
predict方法,确保返回的DataFrame索引唯一:def predict(self, context, model_input): # 原有逻辑生成answers、sources、prompts列表 answers = ... sources = ... prompts = ... # 构建DataFrame并重置索引 result_df = pd.DataFrame({ 'answers': answers, 'sources': sources, 'prompts': prompts }).reset_index(drop=True) return result_df
2. 调整MLflow评估参数(适配多列输出)
修改评估代码,确保MLflow能正确识别多列输出,并避免索引冲突:
evaluation_results = mlflow.evaluate( model=f'models:/{model_name}/{model_version}', data=data.reset_index(drop=True), # 确保输入无重复索引 predictions="answers", extra_metrics=[ mlflow.metrics.latency(), mlflow.metrics.token_count(predictions_col="answers") # 明确指定token统计列 ], evaluator_config={"log_model_explainability": False} # 可选:关闭可解释性以避免额外索引操作 )
3. 手动计算并记录指标(兜底方案)
如果上述方法仍无法解决,可手动运行预测并记录所有指标,完全控制数据流程:
import mlflow import time from your_tokenizer_module import YourTokenizer # 替换为实际的tokenizer # 加载模型 model = mlflow.pyfunc.load_model(f'models:/{model_name}/{model_version}') # 执行预测并计算延迟 start_time = time.perf_counter() predictions_df = model.predict(data.reset_index(drop=True)) total_latency = time.perf_counter() - start_time avg_latency = total_latency / len(predictions_df) # 计算token数量 tokenizer = YourTokenizer() predictions_df['token_count'] = predictions_df['answers'].apply(lambda x: len(tokenizer.encode(x))) avg_token_count = predictions_df['token_count'].mean() # 记录到MLflow with mlflow.start_run(run_name=f"eval_model_v{model_version}"): # 记录数值指标 mlflow.log_metric("avg_latency", avg_latency) mlflow.log_metric("avg_token_count", avg_token_count) # 记录审计用的完整预测结果(支持UI中查看对比) mlflow.log_table(predictions_df, artifact_file="full_predictions.csv")
4. 升级MLflow版本(推荐)
MLflow 2.8.0存在多列输出模型评估的索引处理bug,升级到2.10.0及以上版本可直接解决该问题,同时获得更完善的多列输出评估支持。
实现核心需求
- 数值指标对比:
latency和token_count会被记录为MLflow运行的指标,可在UI的"Metrics"标签页中直接对比不同模型版本的数值。 - 审计数据查看:通过
mlflow.log_table或模型输出的artifact,可在UI的"Artifacts"标签页中查看完整的answers、sources、prompts数据,支持下载对比。
内容的提问来源于stack exchange,提问作者gmatthews88
相关产品推荐
相关产品推荐

