如何用Python请求Meetup API的590页数据并导入Pandas DataFrame
Python Script to Fetch Meetup API Pages and Load into Pandas DataFrame
Here's a complete, robust script to fetch 590 pages of Meetup group data using pagination, handle potential API errors, and load the results into a Pandas DataFrame:
Step 1: Import Required Libraries
First, install any missing packages (e.g., pip install requests pandas) then import them:
import requests import pandas as pd import time from requests.exceptions import HTTPError, Timeout
Step 2: Configure API Parameters
Fill in your valid sig_id and sig values from your example URL. We'll keep most parameters fixed and only update the offset for pagination:
# Base API endpoint for Meetup v2 groups base_url = "https://api.meetup.com/2/groups" # Fixed parameters (match your example request) fixed_params = { "format": "json", "category_id": 34, "photo-host": "public", "page": 100, "radius": 200.0, "order": "id", "desc": "false", "sig_id": "243750775", # Replace with your actual sig_id "sig": "768bcf78d9c7393..." # Replace with your actual sig } # Total pages to fetch total_pages = 590 all_groups = []
Step 3: Pagination Loop with Error Handling
We'll iterate through each offset (from 0 to 589), send requests, and collect results. Adding delays helps avoid hitting Meetup's rate limits:
for offset in range(total_pages): try: # Update offset for current page params = fixed_params.copy() params["offset"] = offset # Send GET request response = requests.get(base_url, params=params, timeout=10) response.raise_for_status() # Raise exception for HTTP errors # Parse JSON and extract group results data = response.json() groups = data.get("results", []) if not groups: print(f"No results found for offset {offset}. Stopping early.") break all_groups.extend(groups) print(f"Successfully fetched page {offset + 1}/{total_pages} ({len(groups)} groups)") # Add a 1-second delay to respect API rate limits time.sleep(1) except HTTPError as e: print(f"HTTP Error on offset {offset}: {e}") # Optional: Retry on 429 (rate limit exceeded) with longer delay if response.status_code == 429: time.sleep(5) continue break except Timeout as e: print(f"Request timed out on offset {offset}: {e}") continue except Exception as e: print(f"Unexpected error on offset {offset}: {e}") break
Step 4: Convert to Pandas DataFrame
Once all data is collected, convert the list of group dictionaries into a DataFrame:
# Create DataFrame from collected groups df = pd.DataFrame(all_groups) # Optional: Preview the first 5 rows print(df.head()) # Optional: Save to CSV for backup df.to_csv("meetup_groups.csv", index=False)
Key Notes:
- Rate Limits: Meetup's API has rate limits (typically 100 requests per 10 minutes). The 1-second delay helps stay within limits, but if you hit a 429 error, increase the delay or add retry logic.
- Valid Signature: Ensure your
sigandsig_idare up-to-date—these are time-sensitive and may expire. - Data Cleaning: Depending on the API response, you might need to clean nested fields (e.g., group photos, location data) using
pd.json_normalize()instead ofpd.DataFrame()if you want to flatten nested JSON structures. For example:df = pd.json_normalize(all_groups) - Early Termination: The loop stops early if a page returns no results, which can save time if there are fewer than 590 actual pages of data.
内容的提问来源于stack exchange,提问作者snow_fall
相关产品推荐
相关产品推荐

