使用Python BeautifulSoup爬取IMDB Top250电影无输出问题求助
爬取IMDB Top250电影列表无输出问题排查
问题描述
我尝试用Python的BeautifulSoup库爬取IMDB官网的Top250电影列表,但运行代码后没有任何输出,请求协助排查问题:
import requests from bs4 import BeautifulSoup url = "https://www.imdb.com/chart/top/?ref_=nv_mv_250" response = requests.get(url) html_content = response.content soup = BeautifulSoup(html_content, "html.parser") movies = soup.find_all( "li", class_="ipc-metadata-list-summary-item sc-bca49391-0 eypSaE cli-parent") for x in movies: print(x.h3.text)
排查与解决方法
先确认请求是否成功
运行代码前先打印response.status_code,如果返回值不是200,说明IMDB拦截了你的请求。IMDB会识别非浏览器请求,需要添加请求头模拟浏览器:headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(url, headers=headers)修正元素选择器
你用的类名里包含动态生成的随机字符(比如sc-bca49391-0 eypSaE),这类类名会随时变化,导致定位不到元素。改用更稳定的定位方式,比如保留固定类名+层级查找:movies = soup.find_all("li", class_="ipc-metadata-list-summary-item") for movie in movies: title_tag = movie.find("h3", class_="ipc-title__text") if title_tag: print(title_tag.text.strip())更换解析器
Python自带的html.parser解析能力有限,换成lxml解析器稳定性更高。先安装依赖:pip install lxml,再修改代码:soup = BeautifulSoup(html_content, "lxml")检查返回内容
打印html_content.decode('utf-8'),看看返回的是正常页面还是反爬验证页面。如果是验证页面,可能需要暂停请求或使用代理。
内容的提问来源于stack exchange,提问作者Alireza Pirmahdi
相关产品推荐
相关产品推荐

