MLFlow按test_rmse指标排序运行记录失败求助
MLFlow按test_rmse排序返回空结果的解决办法
我在用MLFlow客户端按test_rmse指标对运行记录排序时,order_by参数识别不了这个指标,返回空结果,怀疑是search_runs函数的嵌套返回结构导致的。
代码示例
- 可用代码(按运行结束时间排序正常):
runs = mlflow_client.search_runs(experiment_ids = experiment_ids, order_by=["info.end_time"]) - 不可用代码(完全按照官方文档写法,但返回空结果):
runs = mlflow_client.search_runs(experiment_ids = experiment_ids, order_by=["metrics.test_rmse"])
尝试过的无效语法
我试了各种写法指定test_rmse指标,包括:
data.metrics.test_rmsemetrics.'test_rmse'data.metrics.'test_rmse'metrics['test_rmse']、data.metrics.['test_rmse']metrics[test_rmse]、data.metrics.[test_rmse]
但全部都没效果。
未指定排序时的返回示例
不设置order_by参数时,返回的单个Run元素结构如下:
<Run: data=<RunData: metrics={'test_mape': 20.545290384026003, 'test_rmse': 7722.505535457056, 'training_mae': 1317.665098010665, 'training_mse': 2600325.7268625204, 'training_r2_score': 0.9597774286238024, 'training_rmse': 1612.5525501088391, 'training_score': 0.9597774286238023}, params={'alpha': '0.9', 'criterion': 'friedman_mse', 'init': 'None', 'learning_rate': '0.1', 'loss': 'ls', 'max_depth': '3', 'max_features': 'None', 'max_leaf_nodes': 'None', 'min_impurity_decrease': '0.0', 'min_impurity_split': 'None', 'min_samples_leaf': '1', 'min_samples_split': '2', 'min_weight_fraction_leaf': '0.0', 'n_estimators': '100', 'n_iter_no_change': 'None', 'presort': 'auto', 'random_state': '12', 'subsample': '1.0', 'tol': '0.0001', 'validation_fraction': '0.1', 'verbose': '0', 'warm_start': 'False'}, tags={'brand': 'dominicks', 'estimator_class': 'sklearn.ensemble.gradient_boosting.GradientBoostingRegressor', 'estimator_name': 'GradientBoostingRegressor', 'mlflow.parentRunId': '2bccc81e-a97c-4704-bb06-e889701172e2', 'mlflow.rootRunId': 'red_planet_0b4kbtckls', 'mlflow.runName': 'plum_neck_0svr7057', 'mlflow.user': 'Zack Soenen', 'store': '1000'}>, info=<RunInfo: artifact_uri='', end_time=1704397763872, experiment_id='e6a75da7-2df2-4211-8f15-b4e8d9ee8e23', lifecycle_stage='active', run_id='65dcfeea-aa57-4c83-80a4-6b1ee9ac8c9c', run_name='plum_neck_0svr7057', run_uuid='65dcfeea-aa57-4c83-80a4-6b1ee9ac8c9c', start_time=1704397755468, status='FINISHED'>
解决办法
方法1:添加排序方向声明
部分MLFlow版本对无排序方向的参数解析有问题,必须显式指定ASC(升序)或DESC(降序):
# 按test_rmse降序排序,升序则改ASC runs = mlflow_client.search_runs(experiment_ids=experiment_ids, order_by=["metrics.test_rmse DESC"])
方法2:手动排序(兜底方案)
如果API排序依然无效,直接获取所有记录后手动排序:
runs = mlflow_client.search_runs(experiment_ids=experiment_ids) # 按test_rmse升序排列,需要降序就加reverse=True sorted_runs = sorted(runs, key=lambda run: run.data.metrics.get('test_rmse', float('inf')))
这个方法还能处理部分运行记录缺失test_rmse指标的情况,用float('inf')把缺失指标的记录排到最后。
内容的提问来源于stack exchange,提问作者user23214789
相关产品推荐
相关产品推荐

