You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas和BeautifulSoup将多轮爬取结果追加到同一Excel文件

问题核心原因

你的代码存在两处错误导致仅保留最后一页爬取结果:

  • 循环内每次都重新初始化df变量,上一轮爬取的内容会被直接覆盖
  • Excel导出逻辑写在循环内部,每次循环都会生成新的Excel文件覆盖之前的文件,最终仅保留最后一次循环的导出结果

修改后完整代码

from bs4 import BeautifulSoup
import requests
import pandas as pd
from time import gmtime, strftime

# 传入待爬取的URL列表
url = ["https://www.nfl.com/standings/league/2021/reg", "https://www.nfl.com/standings/league/2020/reg", "https://www.nfl.com/standings/league/2019/reg"]

# 提前初始化总表,用来存储所有页面的爬取数据
total_df = pd.DataFrame()

for site in url:
    # 加载页面HTML
    page = requests.get(site)
    soup = BeautifulSoup(page.text, 'lxml')

    # 提取表格数据
    table = soup.find('table', {'summary':'Standings - Detailed View'})

    headers = []
    for i in table.find_all('th'):
        title = i.text.strip()
        headers.append(title)

    # 初始化单页数据的DataFrame
    df = pd.DataFrame(columns = headers)

    # 逐行提取表格内容
    for row in table.find_all('tr')[1:]:
        data = row.find_all('td')
        row_data = [td.text.strip() for td in data]
        length = len(df)
        df.loc[length] = row_data
    
    # 将单页数据追加到总表中
    total_df = pd.concat([total_df, df], ignore_index=True)
    print(f'[*] 已完成站点{site}的数据爬取')

# 所有页面爬取完成后统一导出Excel
dateTime = strftime("%d%b%Y_%H%M", gmtime())
writer = pd.ExcelWriter(f"{dateTime}Z.xlsx")
total_df.to_excel(writer, index=False)
writer.save()
print('[*] 所有数据已成功写入Excel文件')

核心修改说明

  • 循环外提前声明total_df作为总存储容器,每次爬取完单页数据后,用pd.concat()将单页数据追加到总表中,不会覆盖历史数据
  • 导出逻辑移到循环外部,所有页面爬取完成后统一导出总表,避免反复覆盖文件
  • 可选优化:导出时添加index=False参数可以去掉Excel里自带的行号列,表格更整洁

内容的提问来源于stack exchange,提问作者OutlawBandit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 21:45:03