使用Python3.6+BeautifulSoup爬取IMDB影评仅获一条结果的求助
Hey there! Let's dig into why you're only pulling one review content while grabbing all the review titles just fine. This is a super common hiccup when scraping IMDB, so let's walk through the most likely fixes:
1. Double-Check Your Selector Methods
The #1 culprit here is usually using find() instead of find_all() for your review content. find() only grabs the first matching element on the page, while find_all() fetches every instance.
For example, if your code looks like this (which would only get one review):
review = soup.find('div', class_='text show-more__control')
You need to swap it out for:
reviews = soup.find_all('div', class_='text show-more__control')
Also, make sure your CSS selector for the review content matches the actual HTML structure on IMDB. Right now, most full review text lives inside div.text.show-more__control—but IMDB does tweak their classes occasionally, so use your browser's dev tools (F12) to inspect the element and confirm the correct class names.
2. Traverse the DOM Properly (Don’t Mix Global and Local Selections)
If you’re grabbing all titles globally but only pulling one review, you might be making a mistake in how you pair titles with their corresponding reviews. Instead of fetching all titles and then all reviews separately (which can lead to mismatches or missed entries), target the parent container of each review first, then extract both title and content from within that container.
Here’s a better approach:
# Grab every individual review container review_containers = soup.find_all('div', class_='review-container') for container in review_containers: # Extract title from THIS container review_title = container.find('a', class_='title').get_text(strip=True) # Extract review content from THIS container review_content = container.find('div', class_='text show-more__control').get_text(strip=True) print(f"Title: {review_title}") print(f"Review: {review_content}\n")
This way, you’re guaranteed to get the title and review that belong together, and you won’t miss any entries because you’re looping through every single review block.
3. Handle Dynamic Content (Hidden/JS-Loaded Reviews)
IMDB hides longer reviews behind a "Show More" button, and some content might be loaded dynamically with JavaScript. BeautifulSoup can only parse static HTML—so if the full review text isn’t in the initial page source, you’ll only get the truncated version (or maybe just one visible review).
If this is the case, you have two options:
- Use a tool that renders JavaScript: Libraries like
seleniumorrequests-htmlcan simulate a browser, click the "Show More" buttons, and load all dynamic content before scraping. - Check for truncated content: Look for elements with classes like
show-more__control.collapsed—sometimes the full text is hidden in a sibling element or a data attribute that you can extract directly without rendering JS.
4. Rule Out Anti-Scraping Measures
IMDB might limit your requests if you’re hitting their servers too quickly, which could cause partial content to load. Add a delay between requests (using time.sleep()) and make sure you’re sending a proper User-Agent header to mimic a real browser:
import time headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } # After each request: time.sleep(2)
Quick Test Code
Here’s a complete snippet you can run to verify if the issue is with your selector or traversal logic:
import requests from bs4 import BeautifulSoup import time url = "https://www.imdb.com/title/tt1375666/reviews" # Example: Inception reviews headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, 'html.parser') review_containers = soup.find_all('div', class_='review-container') print(f"Found {len(review_containers)} total reviews\n") for idx, container in enumerate(review_containers, 1): title = container.find('a', class_='title').get_text(strip=True) content = container.find('div', class_='text show-more__control').get_text(strip=True) print(f"Review {idx}:") print(f"Title: {title}") print(f"Content: {content[:200]}...\n") # Print first 200 chars to save space time.sleep(1)
Give these steps a shot—odds are it’s either a selector mix-up or a traversal error. Let me know if you still run into issues!
内容的提问来源于stack exchange,提问作者Lacri Mosa

