You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何不使用Selenium获取IMDB电影的全部评论用户名?

Solution to Scrape All IMDB Review Usernames Without Selenium

Hey there! I see you're stuck with only getting the first 25 usernames from IMDB reviews because the rest require clicking "Show More", and you're hitting SSL issues with Selenium. Let's fix this by leveraging IMDB's internal AJAX API that loads additional reviews—no browser automation needed.

How IMDB Loads More Reviews

When you click "Show More", IMDB sends a POST request to an AJAX endpoint (instead of reloading the entire page) to fetch the next batch of reviews. We can mimic this request in our code to get all reviews programmatically.

Step-by-Step Code Implementation

Here's a complete script that fetches all usernames by first grabbing the initial 25, then iterating through the AJAX requests until there are no more reviews left:

import requests
from bs4 import BeautifulSoup
from time import sleep

# Base URLs for the reviews page and AJAX endpoint
base_review_url = "https://www.imdb.com/title/tt0068646/reviews?ref_=tt_urv"
ajax_load_url = "https://www.imdb.com/title/tt0068646/reviews/_ajax"

# Mimic a browser request to avoid being blocked
headers = {
    "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/116.0.0.0 Safari/537.36",
    "Referer": base_review_url
}

all_usernames = []

# 1. Fetch initial page and extract first 25 usernames + pagination key
response = requests.get(base_review_url, headers=headers, verify=False)
soup = BeautifulSoup(response.content, "html5lib")

# Extract initial usernames
initial_name_elements = soup.find_all('span', class_='display-name-link')
for elem in initial_name_elements:
    all_usernames.append(elem.get_text(strip=True))

# Get the pagination key (needed to load more reviews)
load_more_div = soup.find("div", class_="load-more-data")
pagination_key = load_more_div["data-key"] if load_more_div else None

# 2. Loop to load remaining reviews via AJAX
while pagination_key:
    # Add a small delay to avoid overwhelming the server
    sleep(1)
    
    # Prepare payload for the AJAX request
    payload = {
        "ref_": "tt_urv",
        "paginationKey": pagination_key
    }
    
    ajax_response = requests.post(ajax_load_url, headers=headers, data=payload, verify=False)
    ajax_soup = BeautifulSoup(ajax_response.content, "html5lib")
    
    # Extract usernames from the newly loaded batch
    more_name_elements = ajax_soup.find_all('span', class_='display-name-link')
    for elem in more_name_elements:
        all_usernames.append(elem.get_text(strip=True))
    
    # Update pagination key for next batch (if available)
    next_load_more_div = ajax_soup.find("div", class_="load-more-data")
    pagination_key = next_load_more_div["data-key"] if next_load_more_div else None

# Print results
print(f"Total usernames collected: {len(all_usernames)}")
print(all_usernames[:10])  # Print first 10 as a sample

Key Notes

  • Headers Matter: IMDB blocks requests that don't look like they're coming from a browser, so we set a valid User-Agent and Referer header.
  • Pagination Key: This is a unique token IMDB uses to track which batch of reviews to load next. We extract it from the "load-more-data" div in each response.
  • Rate Limiting: Adding a sleep(1) between requests helps avoid triggering IMDB's anti-scraping measures. Adjust the delay if you get blocked.
  • SSL Certificate: Using verify=False skips SSL validation (as you did in your original code), but in a production environment, it's better to fix the SSL certificate issue instead of disabling verification.

内容的提问来源于stack exchange,提问作者cbyoda

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 17:37:44