You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决使用MLflow上传大体积PCA模型时的超时错误?

大尺寸PCA模型上传至EC2部署的MLflow服务器超时连接中断问题

问题概述

尝试将训练好的1.5GB大小的sklearn PCA模型上传到EC2服务器部署的MLflow服务,即便将超时时间调整至900秒仍出现连接中断错误。已尝试GitHub issue#2401提及的环境变量设置方案,未生效。

服务器配置(systemd)

[Unit]
Description=Mlflow server configuration
[Service]
Type=simple
User=ubuntu
ExecStart=/bin/bash -c 'PATH=/home/ubuntu/.local/bin/:$PATH exec mlflow server --host 0.0.0.0 --port 8883 --artifacts-destination s3://icad-mlflow-artifacts --serve-artifacts --backend-store-uri /home/ubuntu/mlruns --gunicorn-opts "--log-level debug --timeout 900 --graceful-timeout 120"'
WorkingDirectory=/home/ubuntu
Restart=always
[Install]
WantedBy=multi-user.target

问题复现代码

import os
import cv2
import numpy as np
import json
import mlflow.pytorch
import numpy as np
from sklearn.decomposition import PCA
import matplotlib.pyplot as plt
import pickle as pk

folder = 'enter the path of folder'
images = []
image_names = []
count = 0
for filename in os.listdir(folder):
    img = cv2.imread(os.path.join(folder, filename))
    img = np.resize(img, (90, 360))
    if img is not None:
        images.append(img.flatten()/255)
        image_names.append(filename)

mlflow.set_tracking_uri("mention your tracking uri")
mlflow.set_experiment(experiment_name="your experiment name")
exp = 'some random name'

with mlflow.start_run(run_name=exp) as run:
    print("run started")
    pca_2000 = PCA(n_components=7000)
    pca_2000_reduced = pca_2000.fit_transform(images)
    pk.dump(pca_2000, open("pca_test_cat.pkl","wb"))
    print("started upload")
    model_info = mlflow.sklearn.log_model(
        sk_model=pca_2000, 
        artifact_path="pca_cat_dog",
        registered_model_name= "pca_model_test_notuseful"
    )

复现条件:使用Kaggle猫狗数据集训练,PCA模型大小≥1.5GB即可触发问题。

报错栈信息

2023/05/25 19:22:47 WARNING mlflow.sklearn: Model was missing function: predict. Not logging python_function flavor!
Traceback (most recent call last):
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\urllib3\connectionpool.py", line 703, in urlopen
    httplib_response = self._make_request(
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\urllib3\connectionpool.py", line 398, in _make_request
    conn.request(method, url, **httplib_request_kw)
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\urllib3\connection.py", line 239, in request
    super(HTTPConnection, self).request(method, url, body=body, headers=headers)
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\http\client.py", line 1285, in request
    self._send_request(method, url, body, headers, encode_chunked)
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\http\client.py", line 1331, in _send_request
    self.endheaders(body, encode_chunked=encode_chunked)
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\http\client.py", line 1280, in endheaders
    self._send_output(message_body, encode_chunked=encode_chunked)
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\http\client.py", line 1079, in _send_output
    self.send(chunk)
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\http\client.py", line 1001, in send
    self.sock.sendall(data)
ConnectionResetError: [WinError 10054] An existing connection was forcibly closed by the remote host

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\requests\adapters.py", line 489, in send
    resp = conn.urlopen(
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\urllib3\connectionpool.py", line 815, in urlopen
    return self.urlopen(
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\urllib3\connectionpool.py", line 815, in urlopen
    return self.urlopen(
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\urllib3\connectionpool.py", line 815, in urlopen
    return self.urlopen(
  [Previous line repeated 2 more times]
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\urllib3\connectionpool.py", line 787, in urlopen
    retries = retries.increment(
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\urllib3\util\retry.py", line 592, in increment
    raise MaxRetryError(_pool, url, error or ResponseError(cause))
urllib3.exceptions.MaxRetryError: HTTPConnectionPool(host='<server url>', port=8883): Max retries exceeded with url: /api/2.0/mlflow-artifacts/artifacts/586086446372044084/031b643c33c54003a364e1a26369ff7b/artifacts/pca_cat_dog/model.pkl (Caused by ProtocolError('Connection aborted.', ConnectionResetError(10054, 'An existing connection was forcibly closed by the remote host', None, 10054, None)))

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\mlflow\utils\rest_utils.py", line 167, in http_request
    return _get_http_response_with_retries(
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\mlflow\utils\rest_utils.py", line 98, in _get_http_response_with_retries
    return session.request(method, url, **kwargs)
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\requests\sessions.py", line 587, in request
    resp = self.send(prep, **send_kwargs)
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\requests\sessions.py", line 701, in send
    r = adapter.send(request, **kwargs)
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\requests\adapters.py", line 565, in send
    raise ConnectionError(e, request=request)
requests.exceptions.ConnectionError: HTTPConnectionPool(host='<server url>', port=8883): Max retries exceeded with url: /api/2.0/mlflow-artifacts/artifacts/586086446372044084/031b643c33c54003a364e1a26369ff7b/artifacts/pca_cat_dog/model.pkl (Caused by ProtocolError('Connection aborted.', ConnectionResetError(10054, 'An existing connection was forcibly closed by the remote host', None, 10054, None)))

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "C:\Users\someting\OneDrive - iCAD Dental\Desktop\git_igps\clustering\training_code_pca\pca_mnist.py", line 39, in <module>
    model_info = mlflow.sklearn.log_model(sk_model=pca_2000, artifact_path="pca_cat_dog",registered_model_name= "pca_model_test_notuseful")
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\mlflow\sklearn\__init__.py", line 424, in log_model
    return Model.log(
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\mlflow\models\model.py", line 552, in log
    mlflow.tracking.fluent.log_artifacts(local_path, mlflow_model.artifact_path)
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\mlflow\tracking\fluent.py", line 817, in log_artifacts
    MlflowClient().log_artifacts(run_id, local_dir, artifact_path)
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\mlflow\tracking\client.py", line 1069, in log_artifacts
    self._tracking_client.log_artifacts(run_id, local_dir, artifact_path)
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\mlflow\tracking\_tracking_service\client.py", line 448, in log_artifacts
    self._get_artifact_repo(run_id).log_artifacts(local_dir, artifact_path)
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\mlflow\store\artifact\http_artifact_repo.py", line 40, in log_artifacts
    self.log_artifact(os.path.join(root, f), artifact_dir)
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\mlflow\store\artifact\http_artifact_repo.py", line 25, in log_artifact
    resp = http_request(self._host_creds, endpoint, "PUT", data=f)
  File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\mlflow\utils\rest_utils.py", line 185, in http_request
    raise MlflowException(f"API request to {url} failed with exception {e}")
mlflow.exceptions.MlflowException: API request to <server url>:8883/api/2.0/mlflow-artifacts/artifacts/586086446372044084/031b643c33c54003a364e1a26369ff7b/artifacts/pca_cat_dog/model.pkl failed with exception HTTPConnectionPool(host='<server url>', port=8883): Max retries exceeded with url: /api/2.0/mlflow-artifacts/artifacts/586086446372044084/031b643c33c54003a364e1a26369ff7b/artifacts/pca_cat_dog/model.pkl (Caused by ProtocolError('Connection aborted.', ConnectionResetError(10054, 'An existing connection was forcibly closed by the remote host', None, 10054, None)))

解决方案建议

  • 调整Gunicorn请求限制:在--gunicorn-opts中添加请求体大小限制参数,比如--limit-request-field_size 65536 --limit-request-line 8192,避免大文件触发请求截断。
  • 跳过MLflow中转直接传S3:客户端配置AWS凭证(环境变量或~/.aws/credentials),MLflow启动时移除--serve-artifacts参数,让客户端直接将模型上传到S3,绕开EC2的HTTP中转。
  • 压缩模型文件:保存模型时启用压缩,减小文件体积:
    pk.dump(pca_2000, open("pca_test_cat.pkl","wb"), protocol=4, compress=('gzip', 9))
    
  • 延长客户端超时:在客户端代码中设置更长的超时时间:
    from mlflow.tracking import MlflowClient
    client = MlflowClient(tracking_uri="your_tracking_uri", timeout=1800)
    
    或设置环境变量MLFLOW_HTTP_REQUEST_TIMEOUT=1800。
  • 检查EC2网络配置:确认安全组允许足够MTU值,无防火墙限制长连接;同时检查实例带宽是否满足大文件上传需求。

内容的提问来源于stack exchange,提问作者Rohan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 14:37:02