如何解决使用MLflow上传大体积PCA模型时的超时错误?
大尺寸PCA模型上传至EC2部署的MLflow服务器超时连接中断问题
问题概述
尝试将训练好的1.5GB大小的sklearn PCA模型上传到EC2服务器部署的MLflow服务,即便将超时时间调整至900秒仍出现连接中断错误。已尝试GitHub issue#2401提及的环境变量设置方案,未生效。
服务器配置(systemd)
[Unit] Description=Mlflow server configuration [Service] Type=simple User=ubuntu ExecStart=/bin/bash -c 'PATH=/home/ubuntu/.local/bin/:$PATH exec mlflow server --host 0.0.0.0 --port 8883 --artifacts-destination s3://icad-mlflow-artifacts --serve-artifacts --backend-store-uri /home/ubuntu/mlruns --gunicorn-opts "--log-level debug --timeout 900 --graceful-timeout 120"' WorkingDirectory=/home/ubuntu Restart=always [Install] WantedBy=multi-user.target
问题复现代码
import os import cv2 import numpy as np import json import mlflow.pytorch import numpy as np from sklearn.decomposition import PCA import matplotlib.pyplot as plt import pickle as pk folder = 'enter the path of folder' images = [] image_names = [] count = 0 for filename in os.listdir(folder): img = cv2.imread(os.path.join(folder, filename)) img = np.resize(img, (90, 360)) if img is not None: images.append(img.flatten()/255) image_names.append(filename) mlflow.set_tracking_uri("mention your tracking uri") mlflow.set_experiment(experiment_name="your experiment name") exp = 'some random name' with mlflow.start_run(run_name=exp) as run: print("run started") pca_2000 = PCA(n_components=7000) pca_2000_reduced = pca_2000.fit_transform(images) pk.dump(pca_2000, open("pca_test_cat.pkl","wb")) print("started upload") model_info = mlflow.sklearn.log_model( sk_model=pca_2000, artifact_path="pca_cat_dog", registered_model_name= "pca_model_test_notuseful" )
复现条件:使用Kaggle猫狗数据集训练,PCA模型大小≥1.5GB即可触发问题。
报错栈信息
2023/05/25 19:22:47 WARNING mlflow.sklearn: Model was missing function: predict. Not logging python_function flavor! Traceback (most recent call last): File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\urllib3\connectionpool.py", line 703, in urlopen httplib_response = self._make_request( File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\urllib3\connectionpool.py", line 398, in _make_request conn.request(method, url, **httplib_request_kw) File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\urllib3\connection.py", line 239, in request super(HTTPConnection, self).request(method, url, body=body, headers=headers) File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\http\client.py", line 1285, in request self._send_request(method, url, body, headers, encode_chunked) File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\http\client.py", line 1331, in _send_request self.endheaders(body, encode_chunked=encode_chunked) File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\http\client.py", line 1280, in endheaders self._send_output(message_body, encode_chunked=encode_chunked) File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\http\client.py", line 1079, in _send_output self.send(chunk) File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\http\client.py", line 1001, in send self.sock.sendall(data) ConnectionResetError: [WinError 10054] An existing connection was forcibly closed by the remote host During handling of the above exception, another exception occurred: Traceback (most recent call last): File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\requests\adapters.py", line 489, in send resp = conn.urlopen( File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\urllib3\connectionpool.py", line 815, in urlopen return self.urlopen( File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\urllib3\connectionpool.py", line 815, in urlopen return self.urlopen( File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\urllib3\connectionpool.py", line 815, in urlopen return self.urlopen( [Previous line repeated 2 more times] File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\urllib3\connectionpool.py", line 787, in urlopen retries = retries.increment( File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\urllib3\util\retry.py", line 592, in increment raise MaxRetryError(_pool, url, error or ResponseError(cause)) urllib3.exceptions.MaxRetryError: HTTPConnectionPool(host='<server url>', port=8883): Max retries exceeded with url: /api/2.0/mlflow-artifacts/artifacts/586086446372044084/031b643c33c54003a364e1a26369ff7b/artifacts/pca_cat_dog/model.pkl (Caused by ProtocolError('Connection aborted.', ConnectionResetError(10054, 'An existing connection was forcibly closed by the remote host', None, 10054, None))) During handling of the above exception, another exception occurred: Traceback (most recent call last): File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\mlflow\utils\rest_utils.py", line 167, in http_request return _get_http_response_with_retries( File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\mlflow\utils\rest_utils.py", line 98, in _get_http_response_with_retries return session.request(method, url, **kwargs) File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\requests\sessions.py", line 587, in request resp = self.send(prep, **send_kwargs) File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\requests\sessions.py", line 701, in send r = adapter.send(request, **kwargs) File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\requests\adapters.py", line 565, in send raise ConnectionError(e, request=request) requests.exceptions.ConnectionError: HTTPConnectionPool(host='<server url>', port=8883): Max retries exceeded with url: /api/2.0/mlflow-artifacts/artifacts/586086446372044084/031b643c33c54003a364e1a26369ff7b/artifacts/pca_cat_dog/model.pkl (Caused by ProtocolError('Connection aborted.', ConnectionResetError(10054, 'An existing connection was forcibly closed by the remote host', None, 10054, None))) During handling of the above exception, another exception occurred: Traceback (most recent call last): File "C:\Users\someting\OneDrive - iCAD Dental\Desktop\git_igps\clustering\training_code_pca\pca_mnist.py", line 39, in <module> model_info = mlflow.sklearn.log_model(sk_model=pca_2000, artifact_path="pca_cat_dog",registered_model_name= "pca_model_test_notuseful") File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\mlflow\sklearn\__init__.py", line 424, in log_model return Model.log( File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\mlflow\models\model.py", line 552, in log mlflow.tracking.fluent.log_artifacts(local_path, mlflow_model.artifact_path) File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\mlflow\tracking\fluent.py", line 817, in log_artifacts MlflowClient().log_artifacts(run_id, local_dir, artifact_path) File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\mlflow\tracking\client.py", line 1069, in log_artifacts self._tracking_client.log_artifacts(run_id, local_dir, artifact_path) File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\mlflow\tracking\_tracking_service\client.py", line 448, in log_artifacts self._get_artifact_repo(run_id).log_artifacts(local_dir, artifact_path) File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\mlflow\store\artifact\http_artifact_repo.py", line 40, in log_artifacts self.log_artifact(os.path.join(root, f), artifact_dir) File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\mlflow\store\artifact\http_artifact_repo.py", line 25, in log_artifact resp = http_request(self._host_creds, endpoint, "PUT", data=f) File "C:\Users\something\AppData\Local\Programs\Python\Python39\lib\site-packages\mlflow\utils\rest_utils.py", line 185, in http_request raise MlflowException(f"API request to {url} failed with exception {e}") mlflow.exceptions.MlflowException: API request to <server url>:8883/api/2.0/mlflow-artifacts/artifacts/586086446372044084/031b643c33c54003a364e1a26369ff7b/artifacts/pca_cat_dog/model.pkl failed with exception HTTPConnectionPool(host='<server url>', port=8883): Max retries exceeded with url: /api/2.0/mlflow-artifacts/artifacts/586086446372044084/031b643c33c54003a364e1a26369ff7b/artifacts/pca_cat_dog/model.pkl (Caused by ProtocolError('Connection aborted.', ConnectionResetError(10054, 'An existing connection was forcibly closed by the remote host', None, 10054, None)))
解决方案建议
- 调整Gunicorn请求限制:在
--gunicorn-opts中添加请求体大小限制参数,比如--limit-request-field_size 65536 --limit-request-line 8192,避免大文件触发请求截断。 - 跳过MLflow中转直接传S3:客户端配置AWS凭证(环境变量或~/.aws/credentials),MLflow启动时移除
--serve-artifacts参数,让客户端直接将模型上传到S3,绕开EC2的HTTP中转。 - 压缩模型文件:保存模型时启用压缩,减小文件体积:
pk.dump(pca_2000, open("pca_test_cat.pkl","wb"), protocol=4, compress=('gzip', 9)) - 延长客户端超时:在客户端代码中设置更长的超时时间:
或设置环境变量from mlflow.tracking import MlflowClient client = MlflowClient(tracking_uri="your_tracking_uri", timeout=1800)MLFLOW_HTTP_REQUEST_TIMEOUT=1800。 - 检查EC2网络配置:确认安全组允许足够MTU值,无防火墙限制长连接;同时检查实例带宽是否满足大文件上传需求。
内容的提问来源于stack exchange,提问作者Rohan
相关产品推荐
相关产品推荐

