You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure ML中MLFlow模型Checkpoint覆盖功能开启方法咨询

问题:Azure ML环境下MLFlow能否覆盖已存在的模型Checkpoint?

我在训练模型时,希望用当前训练轮次(epoch)的最佳模型checkpoint覆盖旧版本,而非累积存储所有过时文件,因此会在模型性能提升时重复调用MLFlow往同一路径记录模型。这个逻辑在本地环境正常运行,但在Azure ML中默认会触发资源冲突报错,且官方文档未提及相关配置方法。想确认是否可以开启Azure Machine Learning Studio模型注册表的覆盖功能?

报错信息

Traceback (most recent call last):
  File "scripts/train.py", line 30, in save_checkpoint
    mlflow.pytorch.log_model(model, "model/checkpoint")
  File "/opt/miniconda/lib/python3.8/site-packages/mlflow/pytorch/__init__.py", line 310, in log_model
    return Model.log(
  File "/opt/miniconda/lib/python3.8/site-packages/mlflow/models/model.py", line 487, in log
    mlflow.tracking.fluent.log_artifacts(local_path, mlflow_model.artifact_path)
  File "/opt/miniconda/lib/python3.8/site-packages/mlflow/tracking/fluent.py", line 810, in log_artifacts
    MlflowClient().log_artifacts(run_id, local_dir, artifact_path)
  File "/opt/miniconda/lib/python3.8/site-packages/mlflow/tracking/client.py", line 1048, in log_artifacts
    self._tracking_client.log_artifacts(run_id, local_dir, artifact_path)
  File "/opt/miniconda/lib/python3.8/site-packages/mlflow/tracking/_tracking_service/client.py", line 448, in log_artifacts
    self._get_artifact_repo(run_id).log_artifacts(local_dir, artifact_path)
  File "/opt/miniconda/lib/python3.8/site-packages/azureml/mlflow/_store/artifact/artifact_repo.py", line 88, in log_artifacts
    self.artifacts.upload_dir(local_dir, dest_path)
  File "/opt/miniconda/lib/python3.8/site-packages/azureml/mlflow/_client/artifact/run_artifact_client.py", line 97, in upload_dir
    result = self._upload_files(
  File "/opt/miniconda/lib/python3.8/site-packages/azureml/mlflow/_client/artifact/base_artifact_client.py", line 34, in _upload_files
    empty_artifact_content = self._create_empty_artifacts(paths=batch_remote_paths)
  File "/opt/miniconda/lib/python3.8/site-packages/azureml/mlflow/_client/artifact/run_artifact_client.py", line 170, in _create_empty_artifacts
    raise Exception("\n".join(error_messages))
Exception: UserError: Resource Conflict: ArtifactId ExperimentRun/dcid.99783d0b-d340-4f2c-ab02-f4dd1c9d7dc0/model/checkpoint/python_env.yaml already exists.
UserError: Resource Conflict: ArtifactId ExperimentRun/dcid.99783d0b-d340-4f2c-ab02-f4dd1c9d7dc0/model/checkpoint/requirements.txt already exists.
UserError: Resource Conflict: ArtifactId ExperimentRun/dcid.99783d0b-d340-4f2c-ab02-f4dd1c9d7dc0/model/checkpoint/MLmodel already exists.
UserError: Resource Conflict: ArtifactId ExperimentRun/dcid.99783d0b-d340-4f2c-ab02-f4dd1c9d7dc0/model/checkpoint/conda.yaml already exists.
UserError: Resource Conflict: ArtifactId ExperimentRun/dcid.99783d0b-d340-4f2c-ab02-f4dd1c9d7dc0/model/checkpoint/data/model.pth already exists.
UserError: Resource Conflict: ArtifactId ExperimentRun/dcid.99783d0b-d340-4f2c-ab02-f4dd1c9d7dc0/model/checkpoint/data/pickle_module_info.txt already exists.

触发报错代码

mlflow.pytorch.log_model(model, f"model/checkpoint/")

临时替代方案(存在缺陷)

可以按轮次命名路径避免冲突,但会累积大量过时checkpoint浪费存储空间:

mlflow.pytorch.log_model(model, f"model/checkpoint/epoch_{epoch_index}")

内容的提问来源于stack exchange,提问作者lo tolmencre

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 19:47:18