You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取IMDB榜单导出Excel仅含表头无数据问题求助

问题分析与解决

你的代码核心问题是:在循环遍历电影数据时,只执行了打印操作,没有将抓取到的数据写入Excel工作表。Excel文件里只有你一开始添加的表头,自然没有电影数据。

修复后的代码

import requests
from bs4 import BeautifulSoup
from openpyxl import Workbook

excel = Workbook() # 创建存储数据的Excel工作簿
sheet = excel.active
sheet.title = 'Top Rated Movies' # 命名工作表
sheet.append(['rank', 'title', 'year', 'rating']) # 添加列标题


try:
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/117.0.0.0 Safari/537.36'
    }
    source = requests.get('https://www.imdb.com/chart/top/', headers=headers)
    source.raise_for_status() 
    
    soup = BeautifulSoup(source.content, 'html.parser') 
    
    movies = soup.find_all('tr')
    for movie in movies[1:]:  # 跳过第一行表头
        cells = movie.find_all('td')
        # 检查cells数量是否符合预期,避免索引越界
        if len(cells) < 3:
            continue
        rank = cells[0].text.strip()
        title = cells[1].find('a').text.strip()
        year_elem = cells[1].find('span', class_='secondaryInfo')
        year = year_elem.text.strip('()') if year_elem else 'N/A'
        rating = cells[2].text.strip()
        
        # 关键:将数据添加到工作表中
        sheet.append([rank, title, year, rating])
        print(f"Rank: {rank}, Title: {title}, Year: {year}, Rating: {rating}")
        
except Exception as e:
    print(f"出错了: {e}")

excel.save('IMDB Movie Ratings.xlsx')

额外说明

  • 新增了sheet.append([rank, title, year, rating])语句,这是把每一行电影数据写入Excel的关键操作。
  • 添加了对cells长度的检查,以及year_elem的判空处理,避免因页面结构微小变化导致代码崩溃。
  • 补充了原代码遗漏的模块导入语句,确保代码可直接运行。

内容的提问来源于stack exchange,提问作者Olumide Olaoye

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 14:24:59