You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas read_html读取HTML表格时如何保留单元格原始HTML格式?

解决方法:读取HTML表格的原始单元格内容

Pandas的pd.read_html()默认会提取表格的文本内容,自动解析HTML实体和标签,所以会丢失原始的HTML结构(比如上标<sup>、特殊字符实体&times;)。要保留原始HTML内容,需要手动用BeautifulSoup解析HTML文件:

步骤1:手动解析HTML并保留原始单元格内容

直接使用BeautifulSoup遍历HTML中的表格,提取每个单元格的原始HTML字符串,再转换为DataFrame:

from bs4 import BeautifulSoup
import pandas as pd

# 读取HTML文件(注意指定正确的编码,避免乱码)
with open('report.htm', 'r', encoding='utf-8') as f:
    soup = BeautifulSoup(f.read(), 'html.parser')

# 收集所有表格的DataFrame
tables = []
for table in soup.find_all('table'):
    # 遍历表格的每一行
    row_list = []
    for tr in table.find_all('tr'):
        # 提取当前行所有单元格(包含<th>和<td>)的原始HTML
        cell_htmls = [str(cell) for cell in tr.find_all(['th', 'td'])]
        row_list.append(cell_htmls)
    # 将行数据转为DataFrame,第一行作为表头
    df = pd.DataFrame(row_list[1:], columns=row_list[0]) if len(row_list) > 1 else pd.DataFrame(row_list)
    tables.append(df)

补充说明

  • 运行上述代码后,每个单元格存储的是原始HTML字符串(比如5&amp;times;10&lt;sup&gt;-5&lt;/sup&gt;),完全保留了原有的标签和实体。
  • 如果需要将HTML实体转换为普通字符(比如&times;转成×),可以用html.unescape()方法处理单元格内容:
    import html
    df['目标列'] = df['目标列'].apply(html.unescape)
    

内容的提问来源于stack exchange,提问作者Michael

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 22:47:33