如何不使用Selenium获取IMDB电影的全部评论用户名?
Hey there! I see you're stuck with only getting the first 25 usernames from IMDB reviews because the rest require clicking "Show More", and you're hitting SSL issues with Selenium. Let's fix this by leveraging IMDB's internal AJAX API that loads additional reviews—no browser automation needed.
How IMDB Loads More Reviews
When you click "Show More", IMDB sends a POST request to an AJAX endpoint (instead of reloading the entire page) to fetch the next batch of reviews. We can mimic this request in our code to get all reviews programmatically.
Step-by-Step Code Implementation
Here's a complete script that fetches all usernames by first grabbing the initial 25, then iterating through the AJAX requests until there are no more reviews left:
import requests from bs4 import BeautifulSoup from time import sleep # Base URLs for the reviews page and AJAX endpoint base_review_url = "https://www.imdb.com/title/tt0068646/reviews?ref_=tt_urv" ajax_load_url = "https://www.imdb.com/title/tt0068646/reviews/_ajax" # Mimic a browser request to avoid being blocked headers = { "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/116.0.0.0 Safari/537.36", "Referer": base_review_url } all_usernames = [] # 1. Fetch initial page and extract first 25 usernames + pagination key response = requests.get(base_review_url, headers=headers, verify=False) soup = BeautifulSoup(response.content, "html5lib") # Extract initial usernames initial_name_elements = soup.find_all('span', class_='display-name-link') for elem in initial_name_elements: all_usernames.append(elem.get_text(strip=True)) # Get the pagination key (needed to load more reviews) load_more_div = soup.find("div", class_="load-more-data") pagination_key = load_more_div["data-key"] if load_more_div else None # 2. Loop to load remaining reviews via AJAX while pagination_key: # Add a small delay to avoid overwhelming the server sleep(1) # Prepare payload for the AJAX request payload = { "ref_": "tt_urv", "paginationKey": pagination_key } ajax_response = requests.post(ajax_load_url, headers=headers, data=payload, verify=False) ajax_soup = BeautifulSoup(ajax_response.content, "html5lib") # Extract usernames from the newly loaded batch more_name_elements = ajax_soup.find_all('span', class_='display-name-link') for elem in more_name_elements: all_usernames.append(elem.get_text(strip=True)) # Update pagination key for next batch (if available) next_load_more_div = ajax_soup.find("div", class_="load-more-data") pagination_key = next_load_more_div["data-key"] if next_load_more_div else None # Print results print(f"Total usernames collected: {len(all_usernames)}") print(all_usernames[:10]) # Print first 10 as a sample
Key Notes
- Headers Matter: IMDB blocks requests that don't look like they're coming from a browser, so we set a valid
User-AgentandRefererheader. - Pagination Key: This is a unique token IMDB uses to track which batch of reviews to load next. We extract it from the "load-more-data" div in each response.
- Rate Limiting: Adding a
sleep(1)between requests helps avoid triggering IMDB's anti-scraping measures. Adjust the delay if you get blocked. - SSL Certificate: Using
verify=Falseskips SSL validation (as you did in your original code), but in a production environment, it's better to fix the SSL certificate issue instead of disabling verification.
内容的提问来源于stack exchange,提问作者cbyoda

