如何用Python遍历分页API获取500部热门电影并规避限流
Fetch 500 Popular Drama Movies (2004+) from TMDB API with Pagination & Rate Limiting
Here's a robust, user-friendly solution that handles pagination, respects TMDB's rate limits, and stops once you've collected 500 films (or hits the end of available results):
Step-by-Step Explanation
- Fix the Filter Parameter: Your original URL had an incorrect parameter for release year (
primary_release_year>=%3D2004). TMDB usesprimary_release_year.gtefor "greater than or equal to" filters. - Pagination Handling: We'll increment the
pageparameter in each request until we reach 500 films or run out of results. - Rate Limit Safety: Even though 25 requests (for 500 films) are well under the 40 calls/10 seconds limit, adding a small delay between requests prevents accidental throttling.
- Error Handling: We'll catch HTTP errors and network issues to avoid crashes and keep the script running smoothly.
Complete Code
import requests import time # Configuration - Replace with your actual API key API_KEY = 'your_api_key_here' BASE_DISCOVER_URL = 'https://api.themoviedb.org/3/discover/movie' TARGET_FILM_COUNT = 500 DRAMA_GENRE_ID = 18 MIN_RELEASE_YEAR = 2004 # Initialize variables to track progress all_popular_films = [] current_page = 1 total_fetched = 0 print(f"Fetching up to {TARGET_FILM_COUNT} popular drama movies (2004+)...\n") while total_fetched < TARGET_FILM_COUNT: # Build request parameters with current page request_params = { 'api_key': API_KEY, 'language': 'en-US', 'sort_by': 'popularity.desc', 'include_adult': 'false', 'include_video': 'false', 'primary_release_year.gte': MIN_RELEASE_YEAR, 'with_genres': DRAMA_GENRE_ID, 'page': current_page } try: # Send request and parse response response = requests.get(BASE_DISCOVER_URL, params=request_params) response.raise_for_status() # Trigger error for HTTP 4xx/5xx statuses page_data = response.json() except requests.exceptions.RequestException as e: print(f"Failed to fetch page {current_page}: {str(e)}") break # Extract results from the page page_results = page_data.get('results', []) if not page_results: print("No more results available. Stopping early.") break # Calculate how many films we need to reach the target films_needed = TARGET_FILM_COUNT - total_fetched # Only take as many as needed from the current page films_to_add = page_results[:films_needed] # Add to our main list all_popular_films.extend(films_to_add) total_fetched += len(films_to_add) print(f"Fetched page {current_page} - Total films collected: {total_fetched}") # Move to next page current_page += 1 # Add a small delay to respect rate limits (0.3s per call = ~33 calls/10s, well under 40) time.sleep(0.3) # Print the final results in the desired format print("\nFinal Results:") print("Movie_ID Movie_Title") for film in all_popular_films: print(f"{film['id']} {film['title']}")
Key Notes
- Rate Limit Compliance: The
time.sleep(0.3)ensures we stay well below the 40 calls/10 seconds limit. If you need faster fetching, you could reduce this to 0.25s (exactly 40 calls/10s), but adding a buffer is safer to avoid accidental throttling. - Early Termination: The loop stops if we run out of results before hitting 500 films, so you won't get stuck in an infinite loop.
- Error Resilience: The
try-exceptblock handles network issues, invalid API keys, or rate limit errors gracefully, giving you clear feedback instead of crashing.
内容的提问来源于stack exchange,提问作者izzy84
相关产品推荐
相关产品推荐

