Llama-API二次请求后停止工作,请求超时问题求助
针对Llama-API循环调用超时的解决方案
添加调用间隔与速率控制:Llama-API存在请求频率限制,连续无间隔调用会触发限流导致超时。在每次API调用后添加延迟,示例:
import time def get_generated_text(...): time.sleep(1) response = client.chat.completions.create(model="llama3.3-70b", messages=[ {"role": "system", "content": "..."}, {"role": "user", "content": "..."} ], top_p=0.95) return response.choices[0].message.content设置请求超时并捕获异常:在
create方法中显式设置超时时间,同时捕获超时异常避免脚本卡住:from openai import APIConnectionError, APITimeoutError def get_generated_text(...): try: response = client.chat.completions.create(model="llama3.3-70b", messages=[ {"role": "system", "content": "..."}, {"role": "user", "content": "..."} ], top_p=0.95, timeout=30) return response.choices[0].message.content except (APITimeoutError, APIConnectionError) as e: print(f"请求异常: {e}") return None实现重试机制:针对超时或临时错误自动重试,示例用tenacity库:
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type from openai import APIConnectionError, APITimeoutError @retry( stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=2, max=10), retry=retry_if_exception_type((APITimeoutError, APIConnectionError)) ) def get_generated_text(...): response = client.chat.completions.create(model="llama3.3-70b", messages=[ {"role": "system", "content": "..."}, {"role": "user", "content": "..."} ], top_p=0.95) return response.choices[0].message.content测试小模型排除负载问题:llama3.3-70b属于大参数模型,服务端负载较高时容易超时。先切换为
llama3.3-8b测试,若不再超时,说明是大模型的服务端资源限制问题。复用客户端实例:确保
client是全局实例,不要在循环内重复创建,避免额外的连接开销。
内容的提问来源于stack exchange,提问作者padraig
相关产品推荐
相关产品推荐

