You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中安全高效地分页处理大型JSON API响应,规避错误与重复数据?

Handling Large Paginated JSON Datasets with Python's requests Library

Great question—handling large paginated datasets is super common, and there are straightforward ways to make this efficient and safe. Let’s break down each of your concerns with practical code and best practices:

First, here's your original context for reference:

I'm working with a REST API that returns a large JSON dataset with hundreds of thousands of records, which supports pagination via parameters like startIndex and resultsPerPage. I'm using Python's requests library for development, and can correctly fetch and handle small amounts of data, but I'm unsure how to efficiently and safely process the full dataset. Here's a simplified version of my current code:

import requests
url = "https://example.com/api?resultsPerPage=5"
response = requests.get(url)
if response.status_code == 200:
    data = response.json()
    for item in data["items"]:
        print(item["id"])
else:
    print("Error:", response.status_code)

My questions are:

  1. How do I loop through all paginated data using parameters like startIndex?
  2. How do I handle cases where response.json() fails (e.g., JSONDecodeError)?
  3. What storage solution can effectively avoid duplicate records?
  4. What are the best practices for handling large API responses in Python?

1. Looping Through All Paginated Data with startIndex

Most APIs either return a total record count (e.g., in a totalItems field) or let you keep fetching until you get an empty list of items. The latter is more flexible (some APIs don’t expose total counts), so let’s use that approach. We’ll also use params instead of hardcoding the URL—this is cleaner and avoids URL encoding issues.

import requests
import time

base_url = "https://example.com/api"
results_per_page = 100  # Pick a reasonable value (check API limits!)
start_index = 0

while True:
    # Build request parameters
    params = {
        "startIndex": start_index,
        "resultsPerPage": results_per_page
    }
    
    response = requests.get(base_url, params=params)
    
    # We'll add robust error handling next, but let's assume basic success for now
    if response.status_code != 200:
        print(f"Request failed: Status code {response.status_code}")
        break
    
    try:
        data = response.json()
    except Exception as e:
        print(f"Failed to parse JSON: {str(e)}")
        break
    
    items = data.get("items", [])
    if not items:
        # No more data to fetch—exit loop
        break
    
    # Process the current page's items (we'll cover storage later)
    for item in items:
        print(item["id"])
    
    # Update start index for the next page
    start_index += results_per_page
    
    # Add a small delay to avoid hitting rate limits
    time.sleep(1)

If the API returns a totalItems field, you can calculate the total number of pages upfront and loop through that instead—just replace the while True with:

total_items = data.get("totalItems", 0)
total_pages = (total_items + results_per_page - 1) // results_per_page  # Round up

for page in range(total_pages):
    start_index = page * results_per_page
    # Rest of the request logic here

2. Handling JSONDecodeError and Response Failures

Never assume response.json() will work—network glitches, API errors, or malformed JSON can all break it. Let’s add robust error handling to cover these cases:

import json
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

# Set up a session with retries for transient errors
session = requests.Session()
retry_strategy = Retry(
    total=3,
    backoff_factor=1,  # Wait 1s, 2s, 4s between retries
    status_forcelist=[500, 502, 503, 504]  # Retry on server errors
)
adapter = HTTPAdapter(max_retries=retry_strategy)
session.mount("https://", adapter)
session.mount("http://", adapter)

# Inside your loop:
response = session.get(base_url, params=params)

# First check for HTTP errors
if response.status_code >= 400:
    print(f"API Error: {response.status_code} - {response.text[:500]}")  # Print first 500 chars
    if response.status_code == 429:
        # Handle rate limiting—maybe wait longer or exit
        print("Rate limited! Waiting 5 seconds...")
        time.sleep(5)
    continue

# Verify the response is actually JSON
content_type = response.headers.get("Content-Type", "")
if "application/json" not in content_type:
    print(f"Unexpected content type: {content_type}")
    print(f"Response snippet: {response.text[:500]}")
    continue

# Try parsing JSON with specific error handling
try:
    data = response.json()
except json.JSONDecodeError as e:
    print(f"JSON Parse Error: Position {e.pos} - {e.msg}")
    print(f"Problematic snippet: {response.text[e.pos-20:e.pos+20]}")
    continue

This code handles transient errors with retries, checks for non-JSON responses, and gives you useful debug info when things go wrong.


3. Avoiding Duplicate Records

The key here is tracking unique identifiers (like the id field in your items). The solution depends on where you’re storing the data:

In-Memory (Smaller Datasets)

Use a set to track seen IDs, so you only keep unique items:

seen_ids = set()
unique_items = []

for item in items:
    item_id = item["id"]
    if item_id not in seen_ids:
        seen_ids.add(item_id)
        unique_items.append(item)

Database (Large Datasets)

This is the best approach for hundreds of thousands of records:

  • SQL Databases: Add a unique constraint to the id column. When inserting, use:
    • MySQL: INSERT INTO items (id, ...) VALUES (...) ON DUPLICATE KEY UPDATE id=id; (does nothing if duplicate)
    • PostgreSQL: INSERT INTO items (id, ...) VALUES (...) ON CONFLICT (id) DO NOTHING;
  • MongoDB: Create a unique index on the id field, then use insert_one() with ordered=False (skips duplicates) or update_one({"id": item_id}, {"$set": item}, upsert=False) (only updates if exists, skips otherwise).

File Storage

If you must save to a file, load existing IDs into a set first, then filter new items before appending. For large files, avoid loading everything into memory—use a database instead, or process the file line-by-line (if using newline-delimited JSON).


4. Best Practices for Large API Responses

  • Batch Process, Don’t Hoard: Instead of storing all items in a list, process each page immediately (e.g., insert into the database) to avoid memory overload.
  • Respect Rate Limits: Check the API’s X-RateLimit-Limit and X-RateLimit-Remaining headers. If you hit a 429 error, wait the time specified in the Retry-After header.
  • Log Everything: Use Python’s logging module instead of print statements. Log request URLs, status codes, number of items processed, and errors—this makes debugging way easier.
  • Validate Data: Use libraries like pydantic to validate item structures. This ensures you don’t store invalid data and catches API schema changes early.
  • Stream When Possible: If the API supports streaming JSON (e.g., newline-delimited), use response.iter_lines() to process items line-by-line without loading the entire response into memory.
  • Use Async for Higher Throughput: If the API allows concurrent requests, use aiohttp instead of requests to fetch multiple pages at once (just don’t overload the API!).

内容的提问来源于stack exchange,提问作者Lucus_sathyabama

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 06:40:16