如何在Python中高效安全地分页处理大型REST API JSON响应,避免报错与重复数据?
处理大型分页REST API数据集的Python实践方案
针对你遇到的大型分页API数据处理问题,我会逐个解答你的疑问,同时给出可落地的代码示例和最佳实践:
1. 正确遍历所有分页数据
大多数支持startIndex参数的API,会配合返回总条目数(比如totalItems)或者当没有更多数据时返回空的items列表。我们可以用循环逐步递增startIndex,直到获取不到新数据为止:
import requests BASE_URL = "https://example.com/api" RESULTS_PER_PAGE = 50 # 可适当调大以减少请求次数,注意不要超过API限制 start_index = 0 with requests.Session() as session: # 使用Session复用连接,提升请求效率 while True: params = { "startIndex": start_index, "resultsPerPage": RESULTS_PER_PAGE } response = session.get(BASE_URL, params=params) response.raise_for_status() # 主动抛出HTTP错误(如404、500) data = response.json() items = data.get("items", []) if not items: break # 没有更多数据,退出循环 # 处理当前页的数据 for item in items: print(item["id"]) # 更新startIndex,准备下一页 start_index += RESULTS_PER_PAGE # 可选:如果API返回totalItems,可提前判断是否到末尾 total_items = data.get("totalItems") if total_items and start_index >= total_items: break
2. 处理JSON解析失败的情况
网络波动、API返回非JSON内容都可能导致JSONDecodeError,我们需要用try-except捕获这类异常,同时还要处理其他请求相关的异常(比如连接超时、HTTP错误):
import requests from json.decoder import JSONDecodeError BASE_URL = "https://example.com/api" start_index = 0 with requests.Session() as session: while True: params = {"startIndex": start_index, "resultsPerPage": 50} try: response = session.get(BASE_URL, params=params, timeout=10) # 设置超时时间,避免请求挂起 response.raise_for_status() # 处理4xx/5xx级别的HTTP错误 try: data = response.json() except JSONDecodeError: print(f"第{start_index//50 +1}页返回非JSON内容,跳过该页") start_index += 50 continue items = data.get("items", []) if not items: break # 处理数据 for item in items: print(item["id"]) start_index += 50 except requests.exceptions.RequestException as e: print(f"请求出错: {str(e)}") # 可选:添加重试逻辑,临时错误后等待重试 import time time.sleep(5) continue
3. 避免重复记录
常见的去重方案分两种场景:
- 内存中临时去重:用集合存储已经处理过的
id,每次处理前检查是否存在:
processed_ids = set() # 处理每个item时 for item in items: item_id = item["id"] if item_id not in processed_ids: processed_ids.add(item_id) # 处理/存储该item print(item_id)
- 持久化存储去重:
- 如果存到数据库,将
id设为主键(唯一约束),插入时用INSERT OR IGNORE(SQLite)或ON DUPLICATE KEY UPDATE(MySQL); - 如果存到文件(如CSV),可先读取已有文件,把所有
id加载到集合中,处理新数据时过滤重复项。
- 如果存到数据库,将
4. Python处理大型API响应的最佳实践
- 使用
requests.Session:复用TCP连接,减少握手开销,大幅提升请求效率; - 设置超时时间:避免请求无限挂起,推荐用
timeout=(3, 10)(连接超时3秒,读取超时10秒); - 添加重试机制:针对临时错误(如503、网络波动),可以用
tenacity库实现自动重试:from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type @retry(stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=2, max=10), retry=retry_if_exception_type((requests.exceptions.ConnectionError, requests.exceptions.HTTPError))) def fetch_page(session, url, params): response = session.get(url, params=params, timeout=10) response.raise_for_status() return response.json() - 流式处理+分批存储:不要把所有数据加载到内存,处理完一页就写入文件/数据库,避免内存溢出;
- 尊重API速率限制:查看API文档的速率规则,添加请求间隔(如
time.sleep(1)),或解析响应头中的Retry-After字段; - 日志记录:用
logging模块记录请求状态、处理条数、错误信息,方便后续排查问题; - 使用生成器:如果需要把数据传递给其他函数,用生成器逐步yield数据,减少内存占用:
def fetch_all_items(session, url): start_index = 0 while True: params = {"startIndex": start_index, "resultsPerPage": 50} data = fetch_page(session, url, params) items = data.get("items", []) if not items: break for item in items: yield item start_index += 50 # 使用生成器处理数据 with requests.Session() as s: for item in fetch_all_items(s, BASE_URL): print(item["id"])
内容的提问来源于stack exchange,提问作者Lucus_sathyabama
相关产品推荐
相关产品推荐

