You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中高效安全地分页处理大型REST API JSON响应,避免报错与重复数据?

处理大型分页REST API数据集的Python实践方案

针对你遇到的大型分页API数据处理问题,我会逐个解答你的疑问,同时给出可落地的代码示例和最佳实践:

1. 正确遍历所有分页数据

大多数支持startIndex参数的API,会配合返回总条目数(比如totalItems)或者当没有更多数据时返回空的items列表。我们可以用循环逐步递增startIndex,直到获取不到新数据为止:

import requests

BASE_URL = "https://example.com/api"
RESULTS_PER_PAGE = 50  # 可适当调大以减少请求次数,注意不要超过API限制
start_index = 0

with requests.Session() as session:  # 使用Session复用连接,提升请求效率
    while True:
        params = {
            "startIndex": start_index,
            "resultsPerPage": RESULTS_PER_PAGE
        }
        response = session.get(BASE_URL, params=params)
        response.raise_for_status()  # 主动抛出HTTP错误(如404、500)
        
        data = response.json()
        items = data.get("items", [])
        if not items:
            break  # 没有更多数据,退出循环
        
        # 处理当前页的数据
        for item in items:
            print(item["id"])
        
        # 更新startIndex,准备下一页
        start_index += RESULTS_PER_PAGE
        
        # 可选:如果API返回totalItems,可提前判断是否到末尾
        total_items = data.get("totalItems")
        if total_items and start_index >= total_items:
            break

2. 处理JSON解析失败的情况

网络波动、API返回非JSON内容都可能导致JSONDecodeError,我们需要用try-except捕获这类异常,同时还要处理其他请求相关的异常(比如连接超时、HTTP错误):

import requests
from json.decoder import JSONDecodeError

BASE_URL = "https://example.com/api"
start_index = 0

with requests.Session() as session:
    while True:
        params = {"startIndex": start_index, "resultsPerPage": 50}
        try:
            response = session.get(BASE_URL, params=params, timeout=10)  # 设置超时时间,避免请求挂起
            response.raise_for_status()  # 处理4xx/5xx级别的HTTP错误
            
            try:
                data = response.json()
            except JSONDecodeError:
                print(f"第{start_index//50 +1}页返回非JSON内容,跳过该页")
                start_index += 50
                continue
                
            items = data.get("items", [])
            if not items:
                break
                
            # 处理数据
            for item in items:
                print(item["id"])
                
            start_index += 50
            
        except requests.exceptions.RequestException as e:
            print(f"请求出错: {str(e)}")
            # 可选:添加重试逻辑,临时错误后等待重试
            import time
            time.sleep(5)
            continue

3. 避免重复记录

常见的去重方案分两种场景:

  • 内存中临时去重:用集合存储已经处理过的id,每次处理前检查是否存在:
processed_ids = set()

# 处理每个item时
for item in items:
    item_id = item["id"]
    if item_id not in processed_ids:
        processed_ids.add(item_id)
        # 处理/存储该item
        print(item_id)
  • 持久化存储去重:
    • 如果存到数据库,将id设为主键(唯一约束),插入时用INSERT OR IGNORE(SQLite)或ON DUPLICATE KEY UPDATE(MySQL);
    • 如果存到文件(如CSV),可先读取已有文件,把所有id加载到集合中,处理新数据时过滤重复项。

4. Python处理大型API响应的最佳实践

  • 使用requests.Session:复用TCP连接,减少握手开销,大幅提升请求效率;
  • 设置超时时间:避免请求无限挂起,推荐用timeout=(3, 10)(连接超时3秒,读取超时10秒);
  • 添加重试机制:针对临时错误(如503、网络波动),可以用tenacity库实现自动重试:
    from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
    
    @retry(stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=2, max=10),
           retry=retry_if_exception_type((requests.exceptions.ConnectionError, requests.exceptions.HTTPError)))
    def fetch_page(session, url, params):
        response = session.get(url, params=params, timeout=10)
        response.raise_for_status()
        return response.json()
    
  • 流式处理+分批存储:不要把所有数据加载到内存,处理完一页就写入文件/数据库,避免内存溢出;
  • 尊重API速率限制:查看API文档的速率规则,添加请求间隔(如time.sleep(1)),或解析响应头中的Retry-After字段;
  • 日志记录:用logging模块记录请求状态、处理条数、错误信息,方便后续排查问题;
  • 使用生成器:如果需要把数据传递给其他函数,用生成器逐步yield数据,减少内存占用:
    def fetch_all_items(session, url):
        start_index = 0
        while True:
            params = {"startIndex": start_index, "resultsPerPage": 50}
            data = fetch_page(session, url, params)
            items = data.get("items", [])
            if not items:
                break
            for item in items:
                yield item
            start_index += 50
    
    # 使用生成器处理数据
    with requests.Session() as s:
        for item in fetch_all_items(s, BASE_URL):
            print(item["id"])
    

内容的提问来源于stack exchange,提问作者Lucus_sathyabama

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 09:23:11