生成器函数处理分页API数据时迭代变慢问题排查及优化
def clean_data(response): #do sth return response def extract_data(response): # do some transforms clean_response = clean_data(response) yield clean_response def fetch_paginated_data(first_url_of_paginated_response): response = requests.get(first_url_of_paginated_response) print(response.status_code) response_object = response.json() yield from extract_data(response_object) if isinstance(response_object, dict) and \ 'meta' in response_object and \ 'next' in response_object['meta'] and \ response_object['meta']['next'] is not None: yield from fetch_paginated_data(response_object['meta']['next']) api_endpoint = "www.api.com/?filtered=True" for data in fetch_paginated_data(api_endpoint): insert_data_into_db() pass
问题解答
1. 递归调用是否是栈溢出问题?
不是典型的栈溢出。Python生成器的yield from递归不会像普通递归那样在调用栈中持续积累帧——生成器是惰性执行的,每次yield后函数会暂停,控制权交回调用方,直到下一次迭代才会继续执行。30-50层的递归远低于Python默认的1000层递归限制,不会触发栈溢出错误,但递归结构会带来额外的栈帧维护开销,可能是后期速度变慢的原因之一。
2. 排查方法
- 跟踪递归深度与耗时:在
fetch_paginated_data中添加深度参数和计时,确认速度变慢是否和递归深度正相关:import time def fetch_paginated_data(url, depth=0): print(f"当前递归深度: {depth}") start = time.time() response = requests.get(url) # 原有逻辑 if next_url: yield from fetch_paginated_data(next_url, depth+1) print(f"深度{depth}执行耗时: {time.time()-start:.2f}s") - 拆分耗时环节:单独统计
requests.get和insert_data_into_db()的耗时,排查是API请求变慢还是数据库插入拖慢了整体速度。 - 检查栈使用情况:用
sys.getrecursionlimit()查看当前递归限制,结合traceback模块打印栈帧,确认栈空间占用是否异常。
3. 流程优化建议
(1)把递归改为迭代,消除递归开销
用循环替代递归分页逻辑,避免生成器递归带来的栈帧积累,性能更稳定:
def fetch_paginated_data(first_url): current_url = first_url while current_url is not None: response = requests.get(current_url) print(response.status_code) response_object = response.json() yield from extract_data(response_object) # 获取下一页URL current_url = response_object.get('meta', {}).get('next')
(2)优化API请求效率
- 添加超时与连接复用:用
requests.Session()复用TCP连接,减少握手开销,同时设置超时避免请求挂起:session = requests.Session() def fetch_paginated_data(first_url): current_url = first_url while current_url is not None: response = session.get(current_url, timeout=10) # 原有逻辑 - 处理API限流:如果API有速率限制,添加合理延迟(如
time.sleep(0.5)),避免被限流导致请求变慢。
(3)优化数据库插入
- 批量插入:积累一定数量的数据后批量插入,减少数据库IO次数:
batch_size = 50 data_batch = [] for data in fetch_paginated_data(api_endpoint): data_batch.append(data) if len(data_batch) >= batch_size: insert_batch_into_db(data_batch) data_batch = [] if data_batch: insert_batch_into_db(data_batch) - 事务优化:将批量插入放在单个事务中,减少事务提交的开销;检查表索引,避免插入时索引维护耗时过高。
(4)其他优化
- 若API支持,用
aiohttp实现异步并发请求,提升数据获取效率(注意控制并发数,避免触发API限制); - 增加详细日志,记录每个环节的耗时,方便后续定位瓶颈。
内容的提问来源于stack exchange,提问作者Mushfique Khan
相关产品推荐
相关产品推荐

