Urllib2.urlopen请求IMDB页面返回空内容问题排查与解决咨询
Hey, I’ve run into this exact issue with IMDB before! Let’s break down what’s happening and how to fix it.
Your code works for other URLs but fails with IMDB’s full credits page for two key reasons:
Dynamic Content Loading
IMDB now loads most page content asynchronously using JavaScript. When you send a request withurllib2, you only get the initial "empty shell" HTML that browsers receive first. The actual cast/crew data gets loaded later via additional API calls that yoururllib2request doesn’t trigger (since it can’t execute JavaScript). That’s why your browser shows content after a delay, but your script gets nothing.Anti-Scraping Measures
IMDB has strict anti-scraping checks. Even with a User-Agent header, your request is missing other browser-like headers (likeAccept,Referer, orAccept-Encoding) that signal you’re a real user. It’s likely detecting your script as a crawler and returning an empty response to block you.
Option 1: Improve Your urllib2 Request Headers & Handle Compression
First, try beefing up your request headers to match what a real browser sends, and handle compressed content (many sites return gzipped data by default, which urllib2 doesn’t auto-decompress).
Here’s the modified code:
import urllib2 import ssl import gzip from io import BytesIO def openConnection(URL): try: # Full set of browser-like headers headers = { 'User-Agent': "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/96.0.4664.45 Safari/537.36", 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Accept-Encoding': 'gzip, deflate, br', 'Referer': 'https://www.imdb.com/', 'Connection': 'keep-alive', 'Upgrade-Insecure-Requests': '1' } req = urllib2.Request(URL, headers=headers) context = ssl._create_unverified_context() con = urllib2.urlopen(req, context=context) # Read and decompress content if needed content = con.read() if con.info().get('Content-Encoding') == 'gzip': content = gzip.GzipFile(fileobj=BytesIO(content)).read() print(content) return content except Exception as e: print(f"Connection failed: {str(e)}. Retrying.") def readIMDB(movie): data = openConnection("https://www.imdb.com/title/" + movie + "/fullcredits") readIMDB("tt0084726")
Note: This might work if the empty response was just due to missing headers or unhandled compression, but IMDB’s anti-scraping might still block you long-term.
Option 2: Use a Browser Automation Tool (Selenium)
For reliable access to dynamically loaded content, use Selenium. It simulates a real browser, executes JavaScript, and waits for content to load—just like a human user.
Step 1: Install Dependencies
First, install Selenium and download a browser driver (e.g., ChromeDriver for Chrome):
pip install selenium
Grab ChromeDriver that matches your Chrome version from the official Chrome driver page.
Step 2: Modified Code
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By def readIMDB(movie): url = f"https://www.imdb.com/title/{movie}/fullcredits" # Configure Chrome to run in headless mode (no visible window) chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/96.0.4664.45 Safari/537.36") driver = webdriver.Chrome(options=chrome_options) driver.get(url) # Wait explicitly for content to load (better than hardcoding sleep) wait = WebDriverWait(driver, 10) wait.until(EC.presence_of_element_located((By.TAG_NAME, 'table'))) # Wait for cast table to appear page_source = driver.page_source print(page_source) driver.quit() return page_source readIMDB("tt0084726")
This will reliably fetch the full page content, as it mimics how a real browser interacts with IMDB’s site.
- Avoid making too many rapid requests to IMDB—you risk getting your IP blocked. Add delays between requests if you’re scraping multiple pages.
- If Selenium is too heavy, you could also try
requests-html(a lighter alternative that supports JS rendering), but Selenium is more robust for complex sites like IMDB.
内容的提问来源于stack exchange,提问作者John N

