You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Urllib2.urlopen请求IMDB页面返回空内容问题排查与解决咨询

Hey, I’ve run into this exact issue with IMDB before! Let’s break down what’s happening and how to fix it.

Why This Happens

Your code works for other URLs but fails with IMDB’s full credits page for two key reasons:

  1. Dynamic Content Loading
    IMDB now loads most page content asynchronously using JavaScript. When you send a request with urllib2, you only get the initial "empty shell" HTML that browsers receive first. The actual cast/crew data gets loaded later via additional API calls that your urllib2 request doesn’t trigger (since it can’t execute JavaScript). That’s why your browser shows content after a delay, but your script gets nothing.

  2. Anti-Scraping Measures
    IMDB has strict anti-scraping checks. Even with a User-Agent header, your request is missing other browser-like headers (like Accept, Referer, or Accept-Encoding) that signal you’re a real user. It’s likely detecting your script as a crawler and returning an empty response to block you.

Fixes to Try

Option 1: Improve Your urllib2 Request Headers & Handle Compression

First, try beefing up your request headers to match what a real browser sends, and handle compressed content (many sites return gzipped data by default, which urllib2 doesn’t auto-decompress).

Here’s the modified code:

import urllib2
import ssl
import gzip
from io import BytesIO

def openConnection(URL):
    try:
        # Full set of browser-like headers
        headers = {
            'User-Agent': "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/96.0.4664.45 Safari/537.36",
            'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
            'Accept-Language': 'en-US,en;q=0.5',
            'Accept-Encoding': 'gzip, deflate, br',
            'Referer': 'https://www.imdb.com/',
            'Connection': 'keep-alive',
            'Upgrade-Insecure-Requests': '1'
        }
        req = urllib2.Request(URL, headers=headers)
        context = ssl._create_unverified_context()
        con = urllib2.urlopen(req, context=context)
        
        # Read and decompress content if needed
        content = con.read()
        if con.info().get('Content-Encoding') == 'gzip':
            content = gzip.GzipFile(fileobj=BytesIO(content)).read()
        
        print(content)
        return content
    except Exception as e:
        print(f"Connection failed: {str(e)}. Retrying.")

def readIMDB(movie):
    data = openConnection("https://www.imdb.com/title/" + movie + "/fullcredits")

readIMDB("tt0084726")

Note: This might work if the empty response was just due to missing headers or unhandled compression, but IMDB’s anti-scraping might still block you long-term.

Option 2: Use a Browser Automation Tool (Selenium)

For reliable access to dynamically loaded content, use Selenium. It simulates a real browser, executes JavaScript, and waits for content to load—just like a human user.

Step 1: Install Dependencies

First, install Selenium and download a browser driver (e.g., ChromeDriver for Chrome):

pip install selenium

Grab ChromeDriver that matches your Chrome version from the official Chrome driver page.

Step 2: Modified Code

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

def readIMDB(movie):
    url = f"https://www.imdb.com/title/{movie}/fullcredits"
    
    # Configure Chrome to run in headless mode (no visible window)
    chrome_options = Options()
    chrome_options.add_argument("--headless=new")
    chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/96.0.4664.45 Safari/537.36")
    
    driver = webdriver.Chrome(options=chrome_options)
    driver.get(url)
    
    # Wait explicitly for content to load (better than hardcoding sleep)
    wait = WebDriverWait(driver, 10)
    wait.until(EC.presence_of_element_located((By.TAG_NAME, 'table')))  # Wait for cast table to appear
    
    page_source = driver.page_source
    print(page_source)
    
    driver.quit()
    return page_source

readIMDB("tt0084726")

This will reliably fetch the full page content, as it mimics how a real browser interacts with IMDB’s site.

Final Notes
  • Avoid making too many rapid requests to IMDB—you risk getting your IP blocked. Add delays between requests if you’re scraping multiple pages.
  • If Selenium is too heavy, you could also try requests-html (a lighter alternative that supports JS rendering), but Selenium is more robust for complex sites like IMDB.

内容的提问来源于stack exchange,提问作者John N

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 09:37:32