You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何httpx流式响应场景下超时设置不生效?

问题分析与解决方案

问题原因

你当前设置的LLM_TIMEOUT=60是作为单个数值传递给httpx.AsyncClient和stream方法,但httpx的数值型timeout默认仅控制连接超时和读取响应第一个字节的超时,并不控制整个请求的总耗时。对于LLM的流式响应,服务器会分块返回数据,只要每个块之间的间隔不超过读取超时,即使总请求时长超过60秒,也不会触发TimeoutException。

解决方案

要实现总请求时长超时控制,需要使用httpx.Timeout对象显式指定total参数,同时可以保留连接和读取超时的设置。另外,也可以在流式迭代过程中手动检查累计耗时,做双重保障。

方案1:使用httpx.Timeout对象设置总超时

修改代码中的超时参数,用httpx.Timeout替代单纯的数值,明确指定总超时时间:

import httpx
import logging
import json

from httpx import TimeoutException, Timeout

# 定义超时配置:总时长60秒,连接和读取每个块的超时也设为60秒
LLM_TIMEOUT = Timeout(60, connect=60, read=60, total=60)

try:
   content = ''
   logging.info(f"request payload:{payload}")
   # 将Timeout对象传入AsyncClient
   async with httpx.AsyncClient(timeout=LLM_TIMEOUT) as client:
      # stream方法无需重复传timeout,会继承client的配置
      async with client.stream("POST", MODEL_URL, headers=HEADERS, json=payload) as response:
          if response.status_code != 200:
              ans = await response.aread()
              logging.error(f"response.aread() error")
              raise Exception(f"Failed to generate completion stream: {ans}")
          async for line in response.aiter_lines():
              if line.startswith("data:"):
                  data_str = line[5:]
                  if data_str.strip() == "[DONE]":
                      break
                  try:
                      chunk_data = json.loads(data_str)
                      if chunk_data['choices']:
                          content += chunk_data["choices"][0]["delta"].get("content")
                  except json.JSONDecodeError:
                      logging.error("json parse error:", data_str)
      logging.info(f"request response:{content}")
except TimeoutException as te:
   logging.warning("request timeout")

方案2:手动检查累计耗时(可选,双重保障)

如果担心httpx的总超时在某些场景下不生效,可以在请求开始时记录时间,每次迭代块时检查是否超过超时时间:

import httpx
import logging
import json
import time

from httpx import TimeoutException, Timeout

LLM_TIMEOUT = 60
# 结合Timeout对象控制连接和单块读取超时
CLIENT_TIMEOUT = Timeout(LLM_TIMEOUT, connect=LLM_TIMEOUT, read=LLM_TIMEOUT)

try:
   content = ''
   request_start_time = time.time()
   logging.info(f"request payload:{payload}")
   async with httpx.AsyncClient(timeout=CLIENT_TIMEOUT) as client:
      async with client.stream("POST", MODEL_URL, headers=HEADERS, json=payload) as response:
          if response.status_code != 200:
              ans = await response.aread()
              logging.error(f"response.aread() error")
              raise Exception(f"Failed to generate completion stream: {ans}")
          async for line in response.aiter_lines():
              # 检查总耗时是否超过超时时间
              if time.time() - request_start_time > LLM_TIMEOUT:
                  raise TimeoutException("Total request duration exceeded timeout")
              
              if line.startswith("data:"):
                  data_str = line[5:]
                  if data_str.strip() == "[DONE]":
                      break
                  try:
                      chunk_data = json.loads(data_str)
                      if chunk_data['choices']:
                          content += chunk_data["choices"][0]["delta"].get("content")
                  except json.JSONDecodeError:
                      logging.error("json parse error:", data_str)
      logging.info(f"request response:{content}")
except TimeoutException as te:
   logging.warning("request timeout")

关键说明

  • httpx.Timeout的total参数会控制从请求发起至响应完全接收的总时长,适合LLM流式请求的整体超时需求。
  • 手动检查耗时的方式可以作为补充,确保在任何情况下都能触发超时。
  • 避免重复在AsyncClient和stream方法中传递超时参数,优先使用AsyncClient的全局配置,防止冲突。

内容的提问来源于stack exchange,提问作者haojie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 15:42:43