You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pandas.read_html爬取多表格合并写入CSV仅存最后一个如何解决?

问题原因

你当前代码每次调用df.to_csv()时默认使用覆盖写入模式(mode='w'),每一轮循环都会清空文件原有内容再写入当前表格,因此最终仅能保留最后一个表格的数据。

解决方案

这里提供两种常用实现方式:

方案1:先合并所有表格再统一写入(推荐)

先把抓取到的所有表格存入列表,合并为单个DataFrame后一次性写入CSV,可避免表头重复、写入异常等问题,修改后代码如下:

import requests
import pandas as pd
from bs4 import BeautifulSoup

link = 'https://www.marketwatch.com/investing/stock/mbin/financials/balance-sheet'

def get_tabular_content(s,link):
    res = s.get(link)
    soup = BeautifulSoup(res.text,"lxml")
    # 定义列表存储所有表格的DataFrame
    df_list = []
    for selector in soup.select("table.table--overflow[aria-label^='Financials']"):
        df = pd.read_html(str(selector))[0]
        df_list.append(df)
        print(df)
    # 合并所有表格
    all_df = pd.concat(df_list, ignore_index=True)
    # 一次性写入CSV
    all_df.to_csv('marketwatch.csv', header=True, index=False, encoding='utf-8-sig')

with requests.Session() as s:
    s.headers['User-Agent'] = 'Mozilla/5.0 (Windows NT 6.1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/88.0.4324.150 Safari/537.36'
    get_tabular_content(s,link)

方案2:循环时追加写入

如果表格数据量很大,不想占用额外内存合并,可以在第一次写入时保留表头,后续循环使用追加模式写入且跳过表头:

import requests
import pandas as pd
from bs4 import BeautifulSoup

link = 'https://www.marketwatch.com/investing/stock/mbin/financials/balance-sheet'

def get_tabular_content(s,link):
    res = s.get(link)
    soup = BeautifulSoup(res.text,"lxml")
    # 标记是否是第一个表格
    first_table = True
    for selector in soup.select("table.table--overflow[aria-label^='Financials']"):
        df = pd.read_html(str(selector))[0]
        # 第一个表格写表头,后续追加且不写表头
        df.to_csv('marketwatch.csv', mode='a', header=first_table, index=False, encoding='utf-8-sig')
        first_table = False
        print(df)

with requests.Session() as s:
    s.headers['User-Agent'] = 'Mozilla/5.0 (Windows NT 6.1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/88.0.4324.150 Safari/537.36'
    get_tabular_content(s,link)

注意追加模式写入前如果存在同名文件,会直接在文件末尾追加内容,如果你需要每次运行生成全新文件,需要在循环前先判断并删除旧文件,或者运行前手动删除旧文件。

内容的提问来源于stack exchange,提问作者MITHU

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 09:18:00