设置超时仍冻结的Python API请求问题排查与方案问询
问题描述
我有一个已稳定运行4年的高敏感生产脚本,其中runAPI函数每日调用约100次,此前从未出现问题,但今日执行时发生冻结,函数调用后无法继续执行后续代码,也无法退出。
我已在requests请求中设置timeout=2,预期出现问题时最多2秒内优雅失败,但实际未生效。
可能触发问题的两项变更:
- 近期更换了GCP虚拟机,但已重新安装所有依赖库;
- 目标API服务器可能存在故障(暂无法联系其支持)。
相关代码片段
import requests import json def runAPI(): orderParams = { "someparameterhere": "example", "moreparameterhere": "example" } headerData = { 'Content-type': 'application/json', 'X-ClientLocalIP': '11.161.0.23', 'X-ClientPublicIP': '11.22.55.66', 'X-MACAddress': '32:01:0a:a0:00:16', 'Accept': 'application/json', 'X-PrivateKey': 'privatekeyhere', 'X-UserType': 'USER', 'X-SourceID': 'WEB', 'Authorization': 'Bearer somelongstringhere.' } for j in range(1): try: requestid = requests.request( "POST", "https://apiurl.com/apidataurlexample", data=json.dumps(orderParams), headers=headerData, timeout=2 # expecting this to limit the freeze to 2 seconds ).json() return requestid except Exception as e: print(e) return 0
回溯信息
httplib_response = conn.getresponse() ^^^^^^^^^^^^^^^^^^ File "/usr/lib/python3.11/http/client.py", line 1374, in getresponse response.begin() File "/usr/lib/python3.11/http/client.py", line 318, in begin version, status, reason = self._read_status() ^^^^^^^^^^^^^^^^^^^ File "/usr/lib/python3.11/http/client.py", line 279, in _read_status line = str(self.fp.readline(_MAXLINE + 1), "iso-8859-1") ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/lib/python3.11/socket.py", line 706, in readinto return self._sock.recv_into(b) ^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/lib/python3.11/ssl.py", line 1311, in recv_into return self.read(nbytes, buffer) ^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/lib/python3.11/ssl.py", line 1167, in read return self._sslobj.read(len, buffer) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ TimeoutError: The read operation timed out During handling of the above exception, another exception occurred: Traceback (most recent call last): File "/home/usernamehere/venv3/lib/python3.11/site-packages/requests/adapters.py", line 440, in send resp = conn.urlopen( ^^^^^^^^^^^^^ File "/home/usernamehere/venv3/lib/python3.11/site-packages/urllib3/connectionpool.py", line 802, in urlopen
疑问
- 已显式定义超时,为何函数仍冻结?
- requests应在2秒后抛出异常或返回,为何未执行?
- 新GCP服务器环境是否可能导致该问题?
- 有无更优方式强制严格超时,避免完全冻结?
目前无法复现问题(生产代码今日无法再次运行),明日重试若问题复现将面临严重风险。
补充:API仍有10%概率返回“Connection reset by peer”错误,但该错误已被异常捕获,未导致程序崩溃;我使用Thread()同时调用该函数约20次,怀疑新GCP虚拟机性能过高导致多线程同时报错引发冻结,是否需在线程间设置微小间隔?
解答
1. 超时未生效的原因
requests的timeout=2是总超时(连接+读取的总和),但在极端场景下(比如目标服务器TCP连接已建立但完全不返回响应,或SSL握手阶段卡住),底层socket的超时可能无法被正确触发。另外从回溯看已经抛出了TimeoutError,冻结可能是因为异常处理后的逻辑阻塞,或是多线程场景下的资源竞争(如线程锁、IO阻塞)。
2. 为何未按预期抛出异常/返回
实际上已经抛出了TimeoutError,但可能存在两种情况:
- 异常被捕获后,
print(e)执行时发生阻塞(比如输出缓冲区满、日志系统卡住); - 多线程场景下,部分线程的超时异常触发后,其他线程因GIL锁或连接池耗尽陷入等待,表现为整体程序冻结。
3. 新GCP服务器的影响可能性
是的,存在几种可能:
- GCP的网络配置(防火墙、NAT网关、TCP参数)与旧机器不同,导致socket超时行为变化;
- 新机器性能更高,多线程并发请求时短时间内建立大量连接,触发目标API的限流/异常响应,导致连接挂起;
- 依赖库版本虽重新安装,但与旧机器存在细微差异(如requests、urllib3的小版本更新),带来超时逻辑的变化。
4. 强制严格超时的方案
- 分阶段超时:使用
timeout=(连接超时, 读取超时),比如timeout=(1,1),明确限制连接和读取阶段的超时时间,避免某一阶段无限等待; - 配置连接池:用
requests.adapters.HTTPAdapter限制最大连接数,避免多线程下连接耗尽导致阻塞,示例:from requests.adapters import HTTPAdapter from urllib3.util.retry import Retry session = requests.Session() retry = Retry(total=1, backoff_factor=0.1, status_forcelist=[500, 502, 503, 504]) adapter = HTTPAdapter(max_retries=retry, pool_connections=10, pool_maxsize=10) session.mount("https://", adapter) # 用session发起请求 response = session.post(url, data=..., headers=..., timeout=(1,1)) - 线程级超时控制:使用
concurrent.futures的超时机制强制终止超时线程,示例:from concurrent.futures import ThreadPoolExecutor def run_with_timeout(func, timeout): with ThreadPoolExecutor(max_workers=1) as executor: future = executor.submit(func) try: return future.result(timeout=timeout) except TimeoutError: executor.shutdown(wait=False) return 0 # 调用方式 result = run_with_timeout(runAPI, timeout=3) - 调整系统TCP参数:在GCP机器上配置TCP keepalive,让系统主动检测死连接并关闭,避免socket长期挂起:
sysctl -w net.ipv4.tcp_keepalive_time=60 sysctl -w net.ipv4.tcp_keepalive_intvl=10 sysctl -w net.ipv4.tcp_keepalive_probes=3
关于多线程间隔的建议
建议添加50-100ms的微小间隔,原因:
- 避免短时间内并发请求触发目标API的防护机制,导致服务器故意挂起连接;
- 降低本地连接池压力,减少因连接竞争导致的阻塞;
- 即使目标API故障,分散请求也能降低整体冻结的概率。
内容的提问来源于stack exchange,提问作者user376285
相关产品推荐
相关产品推荐

