You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Kubernetes部署下批量API请求遇连接重置问题求助

Kubernetes环境下批量API请求出现ConnectionResetError问题

我把移动应用后端打包成Docker镜像,分别用Docker Compose和Kubernetes部署,镜像在两个环境都能正常运行。为了测试,我写了个Python脚本发送1000次REST API请求,在Docker Compose环境下跑完全没问题,就算发5000次也不会出错,但在Kubernetes环境里,大概900次成功请求后就会抛出ConnectionResetError。我只需要切换脚本里的服务器IP就能在两个环境间切换测试,网上搜过相关错误但没找到解决办法,希望脚本在两个环境都能正常运行。

错误信息

Exception in thread Thread-1120 (run_hc_call):
Traceback (most recent call last):
  File "D:\python3115\Lib\site-packages\urllib3\connectionpool.py", line 790, in urlopen
    response = self._make_request(
               ^^^^^^^^^^^^^^^^^^^
  File "D:\python3115\Lib\site-packages\urllib3\connectionpool.py", line 536, in _make_request
    response = conn.getresponse()
               ^^^^^^^^^^^^^^^^^^
  File "D:\python3115\Lib\site-packages\urllib3\connection.py", line 461, in getresponse
    httplib_response = super().getresponse()
                       ^^^^^^^^^^^^^^^^^^^^^
  File "D:\python3115\Lib\http\client.py", line 1378, in getresponse
    response.begin()
  File "D:\python3115\Lib\http\client.py", line 318, in begin
    version, status, reason = self._read_status()
                              ^^^^^^^^^^^^^^^^^^^
  File "D:\python3115\Lib\http\client.py", line 279, in _read_status
    line = str(self.fp.readline(_MAXLINE + 1), "iso-8859-1")
               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "D:\python3115\Lib\socket.py", line 706, in readinto
    return self._sock.recv_into(b)
           ^^^^^^^^^^^^^^^^^^^^^^^
ConnectionResetError: [WinError 10054] An existing connection was forcibly closed by the remote host

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "D:\python3115\Lib\site-packages\requests\adapters.py", line 486, in send
    resp = conn.urlopen(
           ^^^^^^^^^^^^^
  File "D:\python3115\Lib\site-packages\urllib3\connectionpool.py", line 844, in urlopen
    retries = retries.increment(
              ^^^^^^^^^^^^^^^^^^
  File "D:\python3115\Lib\site-packages\urllib3\util\retry.py", line 470, in increment
    raise reraise(type(error), error, _stacktrace)
          ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "D:\python3115\Lib\site-packages\urllib3\util\util.py", line 38, in reraise
    raise value.with_traceback(tb)
  File "D:\python3115\Lib\site-packages\urllib3\connectionpool.py", line 790, in urlopen
    response = self._make_request(
               ^^^^^^^^^^^^^^^^^^^
  File "D:\python3115\Lib\site-packages\urllib3\connectionpool.py", line 536, in _make_request
    response = conn.getresponse()
               ^^^^^^^^^^^^^^^^^^
  File "D:\python3115\Lib\site-packages\urllib3\connection.py", line 461, in getresponse
    httplib_response = super().getresponse()
                       ^^^^^^^^^^^^^^^^^^^^^
  File "D:\python3115\Lib\http\client.py", line 1378, in getresponse
    response.begin()
  File "D:\python3115\Lib\http\client.py", line 318, in begin
    version, status, reason = self._read_status()
                              ^^^^^^^^^^^^^^^^^^^
  File "D:\python3115\Lib\http\client.py", line 279, in _read_status
    line = str(self.fp.readline(_MAXLINE + 1), "iso-8859-1")
               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "D:\python3115\Lib\socket.py", line 706, in readinto
    return self._sock.recv_into(b)
           ^^^^^^^^^^^^^^^^^^^^^^^
urllib3.exceptions.ProtocolError: ('Connection aborted.', ConnectionResetError(10054, 'An existing connection was forcibly closed by the remote host', None, 10054, None))

Python测试脚本

import requests as requests
import logging
import threading


thread_count = 1000

logging.basicConfig(
     filename='log_file_name.log',
     level=logging.ERROR,
     format= '[%(asctime)s] %(levelname)s - %(message)s',
     datefmt='%H:%M:%S',
     filemode='w'
 )
logger = logging.getLogger()

docker_server = '<public ip>'
kubernetes_server = '<public ip>'
#server = docker_server
server = kubernetes_server

port = "82"
login_ep = "hcweb/account/login"

params = {"ReturnUrl": "/hcweb"}

payload_username = {'username': 'username', 'password': 'password'}

session = requests.Session()
res = session.get(f"http://{server}:{port}/hcweb", allow_redirects=False)
url = res.headers["Location"]
res2 = session.get(url)
AntiforgeryToken = res2.headers['Set-Cookie'].split(";")[0]

headersAntiforgeryToken = {'Cookie': AntiforgeryToken}


pres = session.post(url, allow_redirects=False, headers=headersAntiforgeryToken, data=payload_username)

HealthCheckToken = pres.headers['Set-Cookie'].split(";")[0]

ping_url = f"http://{server}:{port}/hcweb/home/CheckApi"

paramsIdentity = {'serviceName': 'Identity Api'}
paramsPicking = {'serviceName': 'Picking Api'}
paramsOrganization = {'serviceName': 'Organization Api'}
paramsStocktake = {'serviceName': 'Stocktake Api'}
paramsGoodsin = {'serviceName': 'Goodsin Api'}
paramsTransfer = {'serviceName': 'Transfer Api'}
paramsStoreStockLocation = {'serviceName': 'Store Stock Location Api'}

service_names = ['Identity Api', 'Picking Api', 'Organization Api', 'Stocktake Api', 'Transfer Api', 'Store Stock Location Api']

headersHC = {'Cookie': f"{AntiforgeryToken};{HealthCheckToken}", "Cache-Control": "no-cache",
                "Pragma": "no-cache",
             'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'}


threads = []
def run_hc_call(service_name):
    res4 = session.get(ping_url, headers=headersHC, params={'serviceName': service_name})
    logger.error(f"{service_name}, {res4.status_code}")

for n in range(0,thread_count):
    sn = service_names[n % len(service_names)]
    t = threading.Thread(target=run_hc_call, args=(sn,))
    t.start()
    threads.append(t)

# Wait all threads to finish.
for t in threads:
    t.join()

问题分析与解决方法

可能原因

  1. Kubernetes网络限制:Ingress Controller、Service或Node层面可能有连接数阈值,比如Nginx Ingress的max_connections限制,或者Node的TCP连接跟踪表耗尽;Service的ClientIP会话亲和性会把同一客户端请求都路由到单个Pod,导致单Pod过载。
  2. 连接复用冲突:所有线程共享同一个requests.Session,高并发下Kubernetes网络组件会主动关闭过量复用的连接,而Docker Compose环境网络策略更宽松。
  3. Pod资源不足:Kubernetes中Pod的CPU、内存配额不够,高负载下服务无法处理更多请求,主动关闭连接。
  4. TCP参数差异:Kubernetes节点的TCP超时、keepalive参数与Docker Compose主机不同,导致连接被提前回收。

解决步骤

  1. 调整脚本连接策略

    • 给Session配置连接池与重试机制,避免连接耗尽:
      from requests.adapters import HTTPAdapter
      from urllib3.util.retry import Retry
      
      session = requests.Session()
      retry_strategy = Retry(
          total=3,
          backoff_factor=0.5,
          status_forcelist=[429, 500, 502, 503, 504],
          allowed_methods=["GET"]
      )
      adapter = HTTPAdapter(max_retries=retry_strategy, pool_connections=100, pool_maxsize=100)
      session.mount("http://", adapter)
      
    • 或者让每个线程使用独立Session(需复用登录后的Cookie,避免重复登录):
      def run_hc_call(service_name):
          with requests.Session() as thread_session:
              thread_session.headers.update(headersHC)
              res4 = thread_session.get(ping_url, params={'serviceName': service_name})
              logger.error(f"{service_name}, {res4.status_code}")
      
  2. 检查Kubernetes网络配置

    • 查看Ingress Controller配置,调高max_connections、延长keepalive_timeout;
    • 把Service的sessionAffinity改为None,避免单Pod过载;
    • 在Node上执行ss -s查看TCP连接状态,调整net.ipv4.tcp_max_tw_buckets、net.ipv4.tcp_tw_reuse等参数,优化连接回收。
  3. 增加Pod资源配额
    修改Deployment的Pod资源配置,确保CPU和内存足够支撑高并发:

    resources:
      requests:
        cpu: "1"
        memory: "1Gi"
      limits:
        cpu: "2"
        memory: "2Gi"
    
  4. 添加请求延迟
    在脚本中加入少量延迟,避免瞬间请求量触发Kubernetes防护机制:

    import time
    
    def run_hc_call(service_name):
        res4 = session.get(ping_url, headers=headersHC, params={'serviceName': service_name})
        logger.error(f"{service_name}, {res4.status_code}")
        time.sleep(0.01)  # 增加10ms延迟
    

内容的提问来源于stack exchange,提问作者Ian Robertson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 23:10:57