Python爬取IMDB榜单导出Excel仅含表头无数据问题求助
问题分析与解决
你的代码核心问题是:在循环遍历电影数据时,只执行了打印操作,没有将抓取到的数据写入Excel工作表。Excel文件里只有你一开始添加的表头,自然没有电影数据。
修复后的代码
import requests from bs4 import BeautifulSoup from openpyxl import Workbook excel = Workbook() # 创建存储数据的Excel工作簿 sheet = excel.active sheet.title = 'Top Rated Movies' # 命名工作表 sheet.append(['rank', 'title', 'year', 'rating']) # 添加列标题 try: headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/117.0.0.0 Safari/537.36' } source = requests.get('https://www.imdb.com/chart/top/', headers=headers) source.raise_for_status() soup = BeautifulSoup(source.content, 'html.parser') movies = soup.find_all('tr') for movie in movies[1:]: # 跳过第一行表头 cells = movie.find_all('td') # 检查cells数量是否符合预期,避免索引越界 if len(cells) < 3: continue rank = cells[0].text.strip() title = cells[1].find('a').text.strip() year_elem = cells[1].find('span', class_='secondaryInfo') year = year_elem.text.strip('()') if year_elem else 'N/A' rating = cells[2].text.strip() # 关键:将数据添加到工作表中 sheet.append([rank, title, year, rating]) print(f"Rank: {rank}, Title: {title}, Year: {year}, Rating: {rating}") except Exception as e: print(f"出错了: {e}") excel.save('IMDB Movie Ratings.xlsx')
额外说明
- 新增了
sheet.append([rank, title, year, rating])语句,这是把每一行电影数据写入Excel的关键操作。 - 添加了对
cells长度的检查,以及year_elem的判空处理,避免因页面结构微小变化导致代码崩溃。 - 补充了原代码遗漏的模块导入语句,确保代码可直接运行。
内容的提问来源于stack exchange,提问作者Olumide Olaoye
相关产品推荐
相关产品推荐

