如何为DataFrame循环迭代设置超时处理并排查耗时异常?
解决方案
一、实现单次迭代超时处理
你可以通过concurrent.futures.ThreadPoolExecutor为process_row的执行添加超时限制,当任务执行超过30秒时捕获超时异常,再执行指定的写入逻辑。具体实现如下:
线程池方案(跨平台兼容)
这种方式能保持和原循环一致的串行执行逻辑,避免并发调用API触发限流:
from concurrent.futures import ThreadPoolExecutor, TimeoutError # 假设passed_writer已提前初始化 min_vote = 10 def process_row(index, row): print("Evaluating row index:", index) question = row["Question"] answer = row["Answer"] instruct = "..." instruct2 = "..." try: completion = openai.ChatCompletion.create( model="gpt-3.5-turbo", messages=[{"role": "user", "content": instruct}] ) response = completion["choices"][0]["message"]["content"] completion = openai.ChatCompletion.create( model="gpt-3.5-turbo", messages=[{"role": "user", "content": instruct2}] ) response2 = completion["choices"][0]["message"]["content"] # 原有的其他业务代码 .... OTHER CODE .... except Exception as e: print(e) # 遍历DataFrame并执行带超时的任务 with ThreadPoolExecutor(max_workers=1) as executor: for index, row in df.iterrows(): future = executor.submit(process_row, index, row) try: # 设置30秒超时阈值 future.result(timeout=30) except TimeoutError: print(f"Row index {index} execution timed out after 30 seconds") # 执行超时后的写入操作 row_with_vote = row.tolist() + [min_vote] passed_writer.writerow(row_with_vote)
信号量方案(仅Unix/Linux可用)
如果你的运行环境是类Unix系统,也可以用signal模块实现超时:
import signal class TimeoutException(Exception): pass def timeout_handler(signum, frame): raise TimeoutException("Execution timed out") min_vote = 10 for index, row in df.iterrows(): # 设置30秒超时信号 signal.signal(signal.SIGALRM, timeout_handler) signal.alarm(30) try: process_row(index, row) except TimeoutException: print(f"Row index {index} execution timed out after 30 seconds") row_with_vote = row.tolist() + [min_vote] passed_writer.writerow(row_with_vote) finally: # 取消超时信号,避免影响后续迭代 signal.alarm(0)
二、偶尔耗时十几分钟的原因分析
- API限流排队:当请求频率超过OpenAI的API限额(如并发数、每分钟请求数限制),系统会将请求放入队列等待处理,导致耗时剧增。
- 服务器负载高峰:OpenAI服务器在流量高峰时段负载过高,复杂的Prompt会占用更多计算资源,处理速度大幅下降。
- 网络波动:本地网络不稳定、中间节点延迟或丢包,会拉长API请求的往返时间。
- SDK自动重试:OpenAI Python SDK默认会对5xx等错误进行自动重试,多次重试会累计大量耗时。
- Prompt复杂度:如果
instruct或instruct2内容过长、逻辑复杂,模型需要更长时间生成响应。
内容的提问来源于stack exchange,提问作者Lorenzo Cutrupi
相关产品推荐
相关产品推荐

