You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何导出指定HTML table为CSV?现有代码只能导出全部表格该怎么修改

导出指定HTML表格的修改方案

你当前代码的逻辑是抓取页面全部<table>标签逐个导出,只要把全量遍历的逻辑改成定位到目标表格即可,按需选择以下匹配方案修改:

方案1:已知目标表格的排列顺序

如果确定要导出的是页面里第N个表格(从0开始计数),直接取对应下标的元素即可,示例代码如下:

from bs4 import BeautifulSoup
import pandas as pd

# 用with上下文管理器打开文件更安全,无需手动关闭
with open(r"C:\report.html", "r", encoding="utf-16") as HTMLFileToBeOpened:
    contents = HTMLFileToBeOpened.read()

soup = BeautifulSoup(contents, 'html.parser')
tables = soup.findAll('table')
# 比如要导出第2个表格,下标为1,替换成你需要的下标即可
target_table = tables[1]
df = pd.read_html(str(target_table), skiprows=2)
df[0].to_csv('target_table.csv', encoding='utf-8-sig') # 加utf-8-sig避免中文乱码

方案2:目标表格有id属性

如果目标<table>标签有唯一id(比如<table id="monthly_report">),可以直接按id精准查找,不需要遍历所有表格:

from bs4 import BeautifulSoup
import pandas as pd

with open(r"C:\report.html", "r", encoding="utf-16") as HTMLFileToBeOpened:
    contents = HTMLFileToBeOpened.read()

soup = BeautifulSoup(contents, 'html.parser')
# 替换id参数为你目标表格的实际id值
target_table = soup.find('table', id='monthly_report')
df = pd.read_html(str(target_table), skiprows=2)
df[0].to_csv('target_table.csv', encoding='utf-8-sig')

方案3:目标表格有唯一class属性

如果目标<table>标签有唯一class(比如<table class="data_statistics">),按class查找即可:

from bs4 import BeautifulSoup
import pandas as pd

with open(r"C:\report.html", "r", encoding="utf-16") as HTMLFileToBeOpened:
    contents = HTMLFileToBeOpened.read()

soup = BeautifulSoup(contents, 'html.parser')
# 替换class_参数为你目标表格的实际class值,注意class后面加下划线避免和Python关键字冲突
target_table = soup.find('table', class_='data_statistics')
df = pd.read_html(str(target_table), skiprows=2)
df[0].to_csv('target_table.csv', encoding='utf-8-sig')

方案4:无明确属性,按表格内特征内容匹配

如果目标表格没有id、class特征,可以匹配表格内的固定文本(比如表头有"部门业绩"字样)定位:

from bs4 import BeautifulSoup
import pandas as pd

with open(r"C:\report.html", "r", encoding="utf-16") as HTMLFileToBeOpened:
    contents = HTMLFileToBeOpened.read()

soup = BeautifulSoup(contents, 'html.parser')
tables = soup.findAll('table')
for table in tables:
    # 替换判断条件里的文本为你目标表格里的独有内容
    if "部门业绩" in table.get_text():
        target_table = table
        break
df = pd.read_html(str(target_table), skiprows=2)
df[0].to_csv('target_table.csv', encoding='utf-8-sig')

内容的提问来源于stack exchange,提问作者余振暘

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 19:06:03