为何httpx流式响应场景下超时设置不生效?
问题分析与解决方案
问题原因
你当前设置的LLM_TIMEOUT=60是作为单个数值传递给httpx.AsyncClient和stream方法,但httpx的数值型timeout默认仅控制连接超时和读取响应第一个字节的超时,并不控制整个请求的总耗时。对于LLM的流式响应,服务器会分块返回数据,只要每个块之间的间隔不超过读取超时,即使总请求时长超过60秒,也不会触发TimeoutException。
解决方案
要实现总请求时长超时控制,需要使用httpx.Timeout对象显式指定total参数,同时可以保留连接和读取超时的设置。另外,也可以在流式迭代过程中手动检查累计耗时,做双重保障。
方案1:使用httpx.Timeout对象设置总超时
修改代码中的超时参数,用httpx.Timeout替代单纯的数值,明确指定总超时时间:
import httpx import logging import json from httpx import TimeoutException, Timeout # 定义超时配置:总时长60秒,连接和读取每个块的超时也设为60秒 LLM_TIMEOUT = Timeout(60, connect=60, read=60, total=60) try: content = '' logging.info(f"request payload:{payload}") # 将Timeout对象传入AsyncClient async with httpx.AsyncClient(timeout=LLM_TIMEOUT) as client: # stream方法无需重复传timeout,会继承client的配置 async with client.stream("POST", MODEL_URL, headers=HEADERS, json=payload) as response: if response.status_code != 200: ans = await response.aread() logging.error(f"response.aread() error") raise Exception(f"Failed to generate completion stream: {ans}") async for line in response.aiter_lines(): if line.startswith("data:"): data_str = line[5:] if data_str.strip() == "[DONE]": break try: chunk_data = json.loads(data_str) if chunk_data['choices']: content += chunk_data["choices"][0]["delta"].get("content") except json.JSONDecodeError: logging.error("json parse error:", data_str) logging.info(f"request response:{content}") except TimeoutException as te: logging.warning("request timeout")
方案2:手动检查累计耗时(可选,双重保障)
如果担心httpx的总超时在某些场景下不生效,可以在请求开始时记录时间,每次迭代块时检查是否超过超时时间:
import httpx import logging import json import time from httpx import TimeoutException, Timeout LLM_TIMEOUT = 60 # 结合Timeout对象控制连接和单块读取超时 CLIENT_TIMEOUT = Timeout(LLM_TIMEOUT, connect=LLM_TIMEOUT, read=LLM_TIMEOUT) try: content = '' request_start_time = time.time() logging.info(f"request payload:{payload}") async with httpx.AsyncClient(timeout=CLIENT_TIMEOUT) as client: async with client.stream("POST", MODEL_URL, headers=HEADERS, json=payload) as response: if response.status_code != 200: ans = await response.aread() logging.error(f"response.aread() error") raise Exception(f"Failed to generate completion stream: {ans}") async for line in response.aiter_lines(): # 检查总耗时是否超过超时时间 if time.time() - request_start_time > LLM_TIMEOUT: raise TimeoutException("Total request duration exceeded timeout") if line.startswith("data:"): data_str = line[5:] if data_str.strip() == "[DONE]": break try: chunk_data = json.loads(data_str) if chunk_data['choices']: content += chunk_data["choices"][0]["delta"].get("content") except json.JSONDecodeError: logging.error("json parse error:", data_str) logging.info(f"request response:{content}") except TimeoutException as te: logging.warning("request timeout")
关键说明
httpx.Timeout的total参数会控制从请求发起至响应完全接收的总时长,适合LLM流式请求的整体超时需求。- 手动检查耗时的方式可以作为补充,确保在任何情况下都能触发超时。
- 避免重复在
AsyncClient和stream方法中传递超时参数,优先使用AsyncClient的全局配置,防止冲突。
内容的提问来源于stack exchange,提问作者haojie
相关产品推荐
相关产品推荐

