Azure ML中MLFlow模型Checkpoint覆盖功能开启方法咨询
问题:Azure ML环境下MLFlow能否覆盖已存在的模型Checkpoint?
我在训练模型时,希望用当前训练轮次(epoch)的最佳模型checkpoint覆盖旧版本,而非累积存储所有过时文件,因此会在模型性能提升时重复调用MLFlow往同一路径记录模型。这个逻辑在本地环境正常运行,但在Azure ML中默认会触发资源冲突报错,且官方文档未提及相关配置方法。想确认是否可以开启Azure Machine Learning Studio模型注册表的覆盖功能?
报错信息
Traceback (most recent call last): File "scripts/train.py", line 30, in save_checkpoint mlflow.pytorch.log_model(model, "model/checkpoint") File "/opt/miniconda/lib/python3.8/site-packages/mlflow/pytorch/__init__.py", line 310, in log_model return Model.log( File "/opt/miniconda/lib/python3.8/site-packages/mlflow/models/model.py", line 487, in log mlflow.tracking.fluent.log_artifacts(local_path, mlflow_model.artifact_path) File "/opt/miniconda/lib/python3.8/site-packages/mlflow/tracking/fluent.py", line 810, in log_artifacts MlflowClient().log_artifacts(run_id, local_dir, artifact_path) File "/opt/miniconda/lib/python3.8/site-packages/mlflow/tracking/client.py", line 1048, in log_artifacts self._tracking_client.log_artifacts(run_id, local_dir, artifact_path) File "/opt/miniconda/lib/python3.8/site-packages/mlflow/tracking/_tracking_service/client.py", line 448, in log_artifacts self._get_artifact_repo(run_id).log_artifacts(local_dir, artifact_path) File "/opt/miniconda/lib/python3.8/site-packages/azureml/mlflow/_store/artifact/artifact_repo.py", line 88, in log_artifacts self.artifacts.upload_dir(local_dir, dest_path) File "/opt/miniconda/lib/python3.8/site-packages/azureml/mlflow/_client/artifact/run_artifact_client.py", line 97, in upload_dir result = self._upload_files( File "/opt/miniconda/lib/python3.8/site-packages/azureml/mlflow/_client/artifact/base_artifact_client.py", line 34, in _upload_files empty_artifact_content = self._create_empty_artifacts(paths=batch_remote_paths) File "/opt/miniconda/lib/python3.8/site-packages/azureml/mlflow/_client/artifact/run_artifact_client.py", line 170, in _create_empty_artifacts raise Exception("\n".join(error_messages)) Exception: UserError: Resource Conflict: ArtifactId ExperimentRun/dcid.99783d0b-d340-4f2c-ab02-f4dd1c9d7dc0/model/checkpoint/python_env.yaml already exists. UserError: Resource Conflict: ArtifactId ExperimentRun/dcid.99783d0b-d340-4f2c-ab02-f4dd1c9d7dc0/model/checkpoint/requirements.txt already exists. UserError: Resource Conflict: ArtifactId ExperimentRun/dcid.99783d0b-d340-4f2c-ab02-f4dd1c9d7dc0/model/checkpoint/MLmodel already exists. UserError: Resource Conflict: ArtifactId ExperimentRun/dcid.99783d0b-d340-4f2c-ab02-f4dd1c9d7dc0/model/checkpoint/conda.yaml already exists. UserError: Resource Conflict: ArtifactId ExperimentRun/dcid.99783d0b-d340-4f2c-ab02-f4dd1c9d7dc0/model/checkpoint/data/model.pth already exists. UserError: Resource Conflict: ArtifactId ExperimentRun/dcid.99783d0b-d340-4f2c-ab02-f4dd1c9d7dc0/model/checkpoint/data/pickle_module_info.txt already exists.
触发报错代码
mlflow.pytorch.log_model(model, f"model/checkpoint/")
临时替代方案(存在缺陷)
可以按轮次命名路径避免冲突,但会累积大量过时checkpoint浪费存储空间:
mlflow.pytorch.log_model(model, f"model/checkpoint/epoch_{epoch_index}")
内容的提问来源于stack exchange,提问作者lo tolmencre
相关产品推荐
相关产品推荐

