You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用MLflow向S3存储桶上传模型时遭遇503错误求助

解决MLflow上传模型至Yandex Cloud S3时的503错误

先看几个核心问题和修复步骤:

  • 修正MLFLOW_TRACKING_URI配置错误
    你当前设置的os.environ["MLFLOW_TRACKING_URI"]='http://:8000'缺失了具体的主机IP/域名,这直接导致请求无法定位到MLflow Tracking Server,引发重试耗尽的503错误。改成包含实际地址的格式:

    os.environ["MLFLOW_TRACKING_URI"]='http://62.84.121.234:8000' # 替换为你的Tracking Server真实IP/域名
    
  • 确保MLflow Tracking Server正确配置S3代理
    场景4是通过Tracking Server代理上传模型到S3,启动Tracking Server时必须指定--default-artifact-root为你的Yandex Cloud S3桶路径,同时保证Server所在环境已配置好S3凭证且安装了S3依赖:

    1. 安装依赖:
      pip install s3fs
      
    2. 启动Tracking Server(替换为你的桶路径):
      mlflow server --host 0.0.0.0 --port 8000 --default-artifact-root s3://your-bucket-name/mlflow-artifacts/
      
  • 验证S3连通性与权限

    1. 在本地或Tracking Server机器上,用AWS CLI测试Yandex Cloud S3访问:
      aws s3 ls --endpoint-url=https://storage.yandexcloud.net
      
    2. 检查.aws/credentials的格式是否正确(Yandex Cloud兼容AWS签名):
      [default]
      aws_access_key_id=你的Yandex访问密钥ID
      aws_secret_access_key=你的Yandex私有密钥
      
    3. 确认密钥拥有目标S3桶的读写权限,同时网络能正常访问Yandex存储端点和Tracking Server:
      curl https://storage.yandexcloud.net
      curl http://62.84.121.234:8000/api/2.0/mlflow/experiments/list
      
  • 临时调整重试策略(针对服务波动场景)
    如果是偶尔的服务不稳定,可以增加MLflow的重试次数:

    os.environ["MLFLOW_HTTP_REQUEST_MAX_RETRIES"] = "5"
    os.environ["MLFLOW_HTTP_REQUEST_RETRY_DELAY"] = "2" # 重试间隔2秒
    

内容的提问来源于stack exchange,提问作者sergzemsk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 18:57:43