使用Beautiful Soup爬取IMDB剧集页面时缺失证书标签的问题排查与解决
Hey there! I’ve run into this exact issue with IMDB scraping before, so let’s break down why those certificate tags are missing and how to fix it.
Why the Certificate Tags Are Disappearing
There are two main culprits here:
- Anti-scraping measures: IMDB detects default
requestsuser agents as non-browser traffic and returns truncated or modified HTML (omitting elements like the certificate span). - Dynamic content loading: Some elements on IMDB’s search pages are loaded asynchronously via JavaScript after the initial page load. The raw HTML fetched by
requestsdoesn’t include these JS-rendered elements, even though you see them in your browser.
Fix 1: Add a Browser-like User Agent to Your Request
First, let’s try bypassing the basic anti-scraping check by spoofing a real browser’s user agent. This often works for getting the full HTML content.
Modify your code to include a headers parameter:
import requests from bs4 import BeautifulSoup url = "https://www.imdb.com/search/title/?title_type=tv_episode&num_votes=600,&sort=user_rating,desc&start=1&ref_=adv_nxt" # Use a real browser user agent (update this to match your current browser's UA if needed) headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' } response = requests.get(url, headers=headers) soup = BeautifulSoup(response.content, 'html.parser') data = soup.findAll('div', attrs={'class': 'lister-item mode-advanced'}) # Now check for the certificate tags for item in data: certificate = item.find('span', class_='certificate') print(certificate.text if certificate else "No certificate found")
Fix 2: Use a Headless Browser for Dynamic Content
If adding the user agent doesn’t work, the certificate tags are likely loaded dynamically with JavaScript. In this case, you’ll need a tool that renders the page like a real browser. Selenium is a popular choice for this.
- First, install Selenium and download the appropriate browser driver (e.g., ChromeDriver for Chrome):
pip install selenium
(Follow official driver setup instructions to get the version matching your browser.)
- Then use this code to render the page and scrape the content:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup url = "https://www.imdb.com/search/title/?title_type=tv_episode&num_votes=600,&sort=user_rating,desc&start=1&ref_=adv_nxt" # Set up headless Chrome (runs in background without a window) chrome_options = Options() chrome_options.add_argument('--headless=new') chrome_options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36') driver = webdriver.Chrome(options=chrome_options) driver.get(url) # Get the fully rendered HTML soup = BeautifulSoup(driver.page_source, 'html.parser') data = soup.findAll('div', attrs={'class': 'lister-item mode-advanced'}) # Extract certificates for item in data: certificate = item.find('span', class_='certificate') print(certificate.text if certificate else "No certificate found") driver.quit()
A Quick Note on IMDB’s Terms of Service
Just a reminder: Make sure your scraping activity complies with IMDB’s terms of service. Avoid making too many rapid requests, and consider using their official API if possible for more reliable data access.
内容的提问来源于stack exchange,提问作者psmadness

