You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

生成器函数处理分页API数据时迭代变慢问题排查及优化

def clean_data(response):
    #do sth
    return response

def extract_data(response):
    # do some transforms
    clean_response = clean_data(response)
    yield clean_response

def fetch_paginated_data(first_url_of_paginated_response):
    response = requests.get(first_url_of_paginated_response)
    print(response.status_code)
    response_object = response.json()
    yield from extract_data(response_object)

    if isinstance(response_object, dict) and \
            'meta' in response_object and \
            'next' in response_object['meta'] and \
    response_object['meta']['next'] is not None:    
        yield from fetch_paginated_data(response_object['meta']['next'])

api_endpoint = "www.api.com/?filtered=True"
for data in fetch_paginated_data(api_endpoint):
    insert_data_into_db()
    pass
问题解答

1. 递归调用是否是栈溢出问题?

不是典型的栈溢出。Python生成器的yield from递归不会像普通递归那样在调用栈中持续积累帧——生成器是惰性执行的,每次yield后函数会暂停,控制权交回调用方,直到下一次迭代才会继续执行。30-50层的递归远低于Python默认的1000层递归限制,不会触发栈溢出错误,但递归结构会带来额外的栈帧维护开销,可能是后期速度变慢的原因之一。

2. 排查方法

  • 跟踪递归深度与耗时:在fetch_paginated_data中添加深度参数和计时,确认速度变慢是否和递归深度正相关:
    import time
    def fetch_paginated_data(url, depth=0):
        print(f"当前递归深度: {depth}")
        start = time.time()
        response = requests.get(url)
        # 原有逻辑
        if next_url:
            yield from fetch_paginated_data(next_url, depth+1)
        print(f"深度{depth}执行耗时: {time.time()-start:.2f}s")
    
  • 拆分耗时环节:单独统计requests.get和insert_data_into_db()的耗时,排查是API请求变慢还是数据库插入拖慢了整体速度。
  • 检查栈使用情况:用sys.getrecursionlimit()查看当前递归限制,结合traceback模块打印栈帧,确认栈空间占用是否异常。

3. 流程优化建议

(1)把递归改为迭代,消除递归开销

用循环替代递归分页逻辑,避免生成器递归带来的栈帧积累,性能更稳定:

def fetch_paginated_data(first_url):
    current_url = first_url
    while current_url is not None:
        response = requests.get(current_url)
        print(response.status_code)
        response_object = response.json()
        yield from extract_data(response_object)
        
        # 获取下一页URL
        current_url = response_object.get('meta', {}).get('next')

(2)优化API请求效率

  • 添加超时与连接复用:用requests.Session()复用TCP连接,减少握手开销,同时设置超时避免请求挂起:
    session = requests.Session()
    def fetch_paginated_data(first_url):
        current_url = first_url
        while current_url is not None:
            response = session.get(current_url, timeout=10)
            # 原有逻辑
    
  • 处理API限流:如果API有速率限制,添加合理延迟(如time.sleep(0.5)),避免被限流导致请求变慢。

(3)优化数据库插入

  • 批量插入:积累一定数量的数据后批量插入,减少数据库IO次数:
    batch_size = 50
    data_batch = []
    for data in fetch_paginated_data(api_endpoint):
        data_batch.append(data)
        if len(data_batch) >= batch_size:
            insert_batch_into_db(data_batch)
            data_batch = []
    if data_batch:
        insert_batch_into_db(data_batch)
    
  • 事务优化:将批量插入放在单个事务中,减少事务提交的开销;检查表索引,避免插入时索引维护耗时过高。

(4)其他优化

  • 若API支持,用aiohttp实现异步并发请求,提升数据获取效率(注意控制并发数,避免触发API限制);
  • 增加详细日志,记录每个环节的耗时,方便后续定位瓶颈。

内容的提问来源于stack exchange,提问作者Mushfique Khan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 13:23:28