You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

设置超时仍冻结的Python API请求问题排查与方案问询

问题描述

我有一个已稳定运行4年的高敏感生产脚本,其中runAPI函数每日调用约100次,此前从未出现问题,但今日执行时发生冻结,函数调用后无法继续执行后续代码,也无法退出。

我已在requests请求中设置timeout=2,预期出现问题时最多2秒内优雅失败,但实际未生效。

可能触发问题的两项变更:

  • 近期更换了GCP虚拟机,但已重新安装所有依赖库;
  • 目标API服务器可能存在故障(暂无法联系其支持)。

相关代码片段

import requests
import json

def runAPI():
    orderParams = {
        "someparameterhere": "example",
        "moreparameterhere": "example"
    }

    headerData = {
        'Content-type': 'application/json',
        'X-ClientLocalIP': '11.161.0.23',
        'X-ClientPublicIP': '11.22.55.66',
        'X-MACAddress': '32:01:0a:a0:00:16',
        'Accept': 'application/json',
        'X-PrivateKey': 'privatekeyhere',
        'X-UserType': 'USER',
        'X-SourceID': 'WEB',
        'Authorization': 'Bearer somelongstringhere.'
    }

    for j in range(1):
        try:
            requestid = requests.request(
                "POST",
                "https://apiurl.com/apidataurlexample",
                data=json.dumps(orderParams),
                headers=headerData,
                timeout=2  # expecting this to limit the freeze to 2 seconds
            ).json()
            return requestid
        except Exception as e:
            print(e)
    return 0

回溯信息

httplib_response = conn.getresponse()
                     ^^^^^^^^^^^^^^^^^^
File "/usr/lib/python3.11/http/client.py", line 1374, in getresponse
  response.begin()
File "/usr/lib/python3.11/http/client.py", line 318, in begin
  version, status, reason = self._read_status()
                            ^^^^^^^^^^^^^^^^^^^
File "/usr/lib/python3.11/http/client.py", line 279, in _read_status
  line = str(self.fp.readline(_MAXLINE + 1), "iso-8859-1")
             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/lib/python3.11/socket.py", line 706, in readinto
  return self._sock.recv_into(b)
         ^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/lib/python3.11/ssl.py", line 1311, in recv_into
  return self.read(nbytes, buffer)
         ^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/lib/python3.11/ssl.py", line 1167, in read
  return self._sslobj.read(len, buffer)
         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
TimeoutError: The read operation timed out

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
File "/home/usernamehere/venv3/lib/python3.11/site-packages/requests/adapters.py", line 440, in send
  resp = conn.urlopen(
         ^^^^^^^^^^^^^
File "/home/usernamehere/venv3/lib/python3.11/site-packages/urllib3/connectionpool.py", line 802, in urlopen

疑问

  1. 已显式定义超时,为何函数仍冻结?
  2. requests应在2秒后抛出异常或返回,为何未执行?
  3. 新GCP服务器环境是否可能导致该问题?
  4. 有无更优方式强制严格超时,避免完全冻结?

目前无法复现问题(生产代码今日无法再次运行),明日重试若问题复现将面临严重风险。

补充:API仍有10%概率返回“Connection reset by peer”错误,但该错误已被异常捕获,未导致程序崩溃;我使用Thread()同时调用该函数约20次,怀疑新GCP虚拟机性能过高导致多线程同时报错引发冻结,是否需在线程间设置微小间隔?


解答

1. 超时未生效的原因

requests的timeout=2是总超时(连接+读取的总和),但在极端场景下(比如目标服务器TCP连接已建立但完全不返回响应,或SSL握手阶段卡住),底层socket的超时可能无法被正确触发。另外从回溯看已经抛出了TimeoutError,冻结可能是因为异常处理后的逻辑阻塞,或是多线程场景下的资源竞争(如线程锁、IO阻塞)。

2. 为何未按预期抛出异常/返回

实际上已经抛出了TimeoutError,但可能存在两种情况:

  • 异常被捕获后,print(e)执行时发生阻塞(比如输出缓冲区满、日志系统卡住);
  • 多线程场景下,部分线程的超时异常触发后,其他线程因GIL锁或连接池耗尽陷入等待,表现为整体程序冻结。

3. 新GCP服务器的影响可能性

是的,存在几种可能:

  • GCP的网络配置(防火墙、NAT网关、TCP参数)与旧机器不同,导致socket超时行为变化;
  • 新机器性能更高,多线程并发请求时短时间内建立大量连接,触发目标API的限流/异常响应,导致连接挂起;
  • 依赖库版本虽重新安装,但与旧机器存在细微差异(如requests、urllib3的小版本更新),带来超时逻辑的变化。

4. 强制严格超时的方案

  • 分阶段超时:使用timeout=(连接超时, 读取超时),比如timeout=(1,1),明确限制连接和读取阶段的超时时间,避免某一阶段无限等待;
  • 配置连接池:用requests.adapters.HTTPAdapter限制最大连接数,避免多线程下连接耗尽导致阻塞,示例:
    from requests.adapters import HTTPAdapter
    from urllib3.util.retry import Retry
    
    session = requests.Session()
    retry = Retry(total=1, backoff_factor=0.1, status_forcelist=[500, 502, 503, 504])
    adapter = HTTPAdapter(max_retries=retry, pool_connections=10, pool_maxsize=10)
    session.mount("https://", adapter)
    # 用session发起请求
    response = session.post(url, data=..., headers=..., timeout=(1,1))
    
  • 线程级超时控制:使用concurrent.futures的超时机制强制终止超时线程,示例:
    from concurrent.futures import ThreadPoolExecutor
    
    def run_with_timeout(func, timeout):
        with ThreadPoolExecutor(max_workers=1) as executor:
            future = executor.submit(func)
            try:
                return future.result(timeout=timeout)
            except TimeoutError:
                executor.shutdown(wait=False)
                return 0
    
    # 调用方式
    result = run_with_timeout(runAPI, timeout=3)
    
  • 调整系统TCP参数:在GCP机器上配置TCP keepalive,让系统主动检测死连接并关闭,避免socket长期挂起:
    sysctl -w net.ipv4.tcp_keepalive_time=60
    sysctl -w net.ipv4.tcp_keepalive_intvl=10
    sysctl -w net.ipv4.tcp_keepalive_probes=3
    

关于多线程间隔的建议

建议添加50-100ms的微小间隔,原因:

  • 避免短时间内并发请求触发目标API的防护机制,导致服务器故意挂起连接;
  • 降低本地连接池压力,减少因连接竞争导致的阻塞;
  • 即使目标API故障,分散请求也能降低整体冻结的概率。

内容的提问来源于stack exchange,提问作者user376285

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 02:50:55