You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Python爬虫从IMDB榜单中正确获取导演信息?

解决IMDB榜单导演信息爬取问题

你的代码在获取导演信息时存在两处关键错误,以下是修正方案:

原代码问题分析

  1. 定位逻辑错误:IMDB页面中,导演姓名的a标签并没有class="Director:"这个属性,"Director:"是该段落内的普通文本,用来标识导演区域。
  2. 列表操作错误:findAll()返回的是结果列表,不能直接调用.text方法,需要遍历列表提取每个元素的文本。

修正后的完整代码

from bs4 import BeautifulSoup
import requests 

try:
    source = requests.get('https://www.imdb.com/search/title/?country_of_origin=NP&sort=year,asc')
    source.raise_for_status()

    soup = BeautifulSoup(source.text, 'html.parser')
    movies = soup.find('div', class_="lister-list").findAll('div', class_="lister-item mode-advanced")
    print(f"共找到 {len(movies)} 部影片")

    for movie in movies:
        # 获取影片名称
        name = movie.find('h3', class_="lister-item-header").a.text
        # 获取影片年份
        year = movie.find('span', class_="lister-item-year text-muted unbold").text.strip('()')
        # 获取影片评分(处理无评分的情况)
        ratings = movie.find('strong').text if movie.find('strong') else "无评分"
        
        # 获取导演信息
        # 定位到包含导演的段落(通常是第二个p标签)
        director_paragraph = movie.find_all('p')[1]
        # 提取所有导演姓名(通过href包含director判断)
        directors = [a.text.strip() for a in director_paragraph.find_all('a') if 'director' in a['href']]
        # 拼接导演姓名,多个导演用逗号分隔
        director = ', '.join(directors) if directors else "无导演信息"

        # 输出信息
        print(f"影片名称: {name}")
        print(f"上映年份: {year}")
        print(f"评分: {ratings}")
        print(f"导演: {director}\n")
        # 注释掉break可以遍历所有影片
        # break

except Exception as e:
    print(f"爬取出错: {e}")

关键修正说明

  • 先通过movie.find_all('p')[1]定位到包含导演信息的段落(第一个p标签是时长、分级等信息,第二个才是导演/演员区域)
  • 利用href属性包含director来筛选导演的a标签,避免误取演员信息
  • 增加了异常处理(比如无评分、无导演信息的情况),提升代码鲁棒性

内容的提问来源于stack exchange,提问作者Anish Thapa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 04:30:26