You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Selenium爬取IMDB Top250仅返回首个结果,求排查解决

问题分析与解决

核心问题

你的代码仅获取到第一部电影数据,根源在于:

  • 你用find_elements定位的ul[contains(@role, "presentation")]是整个榜单的容器,页面中只有1个该元素,循环仅执行1次,自然只能拿到首个电影的数据。
  • 未处理页面加载延迟,可能部分元素还未渲染完成就开始定位。

修正步骤

  1. 定位单个电影条目:将循环对象改为每个电影对应的li元素,这才是榜单中每部电影的独立容器。
  2. 添加显式等待:确保页面元素加载完成后再进行定位,避免因加载慢导致的元素查找失败。
  3. 调整元素定位路径:在每个li条目内精准定位标题和年份元素(你之前的年份元素类名写错了)。

修正后的代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd

url = 'https://m.imdb.com/chart/top/'
driver = webdriver.Edge()
driver.get(url)

title = []
year = []

# 等待榜单条目加载完成,定位所有电影li元素
wait = WebDriverWait(driver, 10)
movie_items = wait.until(EC.presence_of_all_elements_located(
    (By.XPATH, './/li[contains(@class, "ipc-metadata-list-summary-item")]')
))

for item in movie_items:
    try:
        # 定位每个条目内的标题
        movie_title = item.find_element(By.XPATH, './/a[contains(@class, "title")]').text
        title.append(movie_title)
        # 定位每个条目内的年份(类名是year而非title)
        movie_year = item.find_element(By.XPATH, './/span[contains(@class, "year")]').text
        year.append(movie_year)
    except Exception as e:
        print(f"处理条目时出错: {e}")
        pass

df_movie = pd.DataFrame({'title': title, 'year': year})
df_movie.to_csv(r'C:\Users\martha\OneDrive\Python\imbd_project.csv', index=False)

driver.quit()  # 关闭浏览器进程

额外说明

  • 显式等待设置10秒超时,足够页面完成渲染;
  • 最后添加driver.quit()关闭浏览器,避免资源占用;
  • 若后续出现反爬限制,可考虑添加随机延迟或更换用户代理。

内容的提问来源于stack exchange,提问作者Martha Imoh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 01:10:02