如何在Python中安全高效地分页处理大型JSON API响应,规避错误与重复数据?
requests Library Great question—handling large paginated datasets is super common, and there are straightforward ways to make this efficient and safe. Let’s break down each of your concerns with practical code and best practices:
First, here's your original context for reference:
I'm working with a REST API that returns a large JSON dataset with hundreds of thousands of records, which supports pagination via parameters like
startIndexandresultsPerPage. I'm using Python'srequestslibrary for development, and can correctly fetch and handle small amounts of data, but I'm unsure how to efficiently and safely process the full dataset. Here's a simplified version of my current code:import requests url = "https://example.com/api?resultsPerPage=5" response = requests.get(url) if response.status_code == 200: data = response.json() for item in data["items"]: print(item["id"]) else: print("Error:", response.status_code)My questions are:
- How do I loop through all paginated data using parameters like
startIndex?- How do I handle cases where
response.json()fails (e.g., JSONDecodeError)?- What storage solution can effectively avoid duplicate records?
- What are the best practices for handling large API responses in Python?
1. Looping Through All Paginated Data with startIndex
Most APIs either return a total record count (e.g., in a totalItems field) or let you keep fetching until you get an empty list of items. The latter is more flexible (some APIs don’t expose total counts), so let’s use that approach. We’ll also use params instead of hardcoding the URL—this is cleaner and avoids URL encoding issues.
import requests import time base_url = "https://example.com/api" results_per_page = 100 # Pick a reasonable value (check API limits!) start_index = 0 while True: # Build request parameters params = { "startIndex": start_index, "resultsPerPage": results_per_page } response = requests.get(base_url, params=params) # We'll add robust error handling next, but let's assume basic success for now if response.status_code != 200: print(f"Request failed: Status code {response.status_code}") break try: data = response.json() except Exception as e: print(f"Failed to parse JSON: {str(e)}") break items = data.get("items", []) if not items: # No more data to fetch—exit loop break # Process the current page's items (we'll cover storage later) for item in items: print(item["id"]) # Update start index for the next page start_index += results_per_page # Add a small delay to avoid hitting rate limits time.sleep(1)
If the API returns a totalItems field, you can calculate the total number of pages upfront and loop through that instead—just replace the while True with:
total_items = data.get("totalItems", 0) total_pages = (total_items + results_per_page - 1) // results_per_page # Round up for page in range(total_pages): start_index = page * results_per_page # Rest of the request logic here
2. Handling JSONDecodeError and Response Failures
Never assume response.json() will work—network glitches, API errors, or malformed JSON can all break it. Let’s add robust error handling to cover these cases:
import json from requests.adapters import HTTPAdapter from urllib3.util.retry import Retry # Set up a session with retries for transient errors session = requests.Session() retry_strategy = Retry( total=3, backoff_factor=1, # Wait 1s, 2s, 4s between retries status_forcelist=[500, 502, 503, 504] # Retry on server errors ) adapter = HTTPAdapter(max_retries=retry_strategy) session.mount("https://", adapter) session.mount("http://", adapter) # Inside your loop: response = session.get(base_url, params=params) # First check for HTTP errors if response.status_code >= 400: print(f"API Error: {response.status_code} - {response.text[:500]}") # Print first 500 chars if response.status_code == 429: # Handle rate limiting—maybe wait longer or exit print("Rate limited! Waiting 5 seconds...") time.sleep(5) continue # Verify the response is actually JSON content_type = response.headers.get("Content-Type", "") if "application/json" not in content_type: print(f"Unexpected content type: {content_type}") print(f"Response snippet: {response.text[:500]}") continue # Try parsing JSON with specific error handling try: data = response.json() except json.JSONDecodeError as e: print(f"JSON Parse Error: Position {e.pos} - {e.msg}") print(f"Problematic snippet: {response.text[e.pos-20:e.pos+20]}") continue
This code handles transient errors with retries, checks for non-JSON responses, and gives you useful debug info when things go wrong.
3. Avoiding Duplicate Records
The key here is tracking unique identifiers (like the id field in your items). The solution depends on where you’re storing the data:
In-Memory (Smaller Datasets)
Use a set to track seen IDs, so you only keep unique items:
seen_ids = set() unique_items = [] for item in items: item_id = item["id"] if item_id not in seen_ids: seen_ids.add(item_id) unique_items.append(item)
Database (Large Datasets)
This is the best approach for hundreds of thousands of records:
- SQL Databases: Add a unique constraint to the
idcolumn. When inserting, use:- MySQL:
INSERT INTO items (id, ...) VALUES (...) ON DUPLICATE KEY UPDATE id=id;(does nothing if duplicate) - PostgreSQL:
INSERT INTO items (id, ...) VALUES (...) ON CONFLICT (id) DO NOTHING;
- MySQL:
- MongoDB: Create a unique index on the
idfield, then useinsert_one()withordered=False(skips duplicates) orupdate_one({"id": item_id}, {"$set": item}, upsert=False)(only updates if exists, skips otherwise).
File Storage
If you must save to a file, load existing IDs into a set first, then filter new items before appending. For large files, avoid loading everything into memory—use a database instead, or process the file line-by-line (if using newline-delimited JSON).
4. Best Practices for Large API Responses
- Batch Process, Don’t Hoard: Instead of storing all items in a list, process each page immediately (e.g., insert into the database) to avoid memory overload.
- Respect Rate Limits: Check the API’s
X-RateLimit-LimitandX-RateLimit-Remainingheaders. If you hit a 429 error, wait the time specified in theRetry-Afterheader. - Log Everything: Use Python’s
loggingmodule instead ofprintstatements. Log request URLs, status codes, number of items processed, and errors—this makes debugging way easier. - Validate Data: Use libraries like
pydanticto validate item structures. This ensures you don’t store invalid data and catches API schema changes early. - Stream When Possible: If the API supports streaming JSON (e.g., newline-delimited), use
response.iter_lines()to process items line-by-line without loading the entire response into memory. - Use Async for Higher Throughput: If the API allows concurrent requests, use
aiohttpinstead ofrequeststo fetch multiple pages at once (just don’t overload the API!).
内容的提问来源于stack exchange,提问作者Lucus_sathyabama

