使用Pandas read_html读取HTML表格时如何保留单元格原始HTML格式?
解决方法:读取HTML表格的原始单元格内容
Pandas的pd.read_html()默认会提取表格的文本内容,自动解析HTML实体和标签,所以会丢失原始的HTML结构(比如上标<sup>、特殊字符实体×)。要保留原始HTML内容,需要手动用BeautifulSoup解析HTML文件:
步骤1:手动解析HTML并保留原始单元格内容
直接使用BeautifulSoup遍历HTML中的表格,提取每个单元格的原始HTML字符串,再转换为DataFrame:
from bs4 import BeautifulSoup import pandas as pd # 读取HTML文件(注意指定正确的编码,避免乱码) with open('report.htm', 'r', encoding='utf-8') as f: soup = BeautifulSoup(f.read(), 'html.parser') # 收集所有表格的DataFrame tables = [] for table in soup.find_all('table'): # 遍历表格的每一行 row_list = [] for tr in table.find_all('tr'): # 提取当前行所有单元格(包含<th>和<td>)的原始HTML cell_htmls = [str(cell) for cell in tr.find_all(['th', 'td'])] row_list.append(cell_htmls) # 将行数据转为DataFrame,第一行作为表头 df = pd.DataFrame(row_list[1:], columns=row_list[0]) if len(row_list) > 1 else pd.DataFrame(row_list) tables.append(df)
补充说明
- 运行上述代码后,每个单元格存储的是原始HTML字符串(比如
5&times;10<sup>-5</sup>),完全保留了原有的标签和实体。 - 如果需要将HTML实体转换为普通字符(比如
×转成×),可以用html.unescape()方法处理单元格内容:import html df['目标列'] = df['目标列'].apply(html.unescape)
内容的提问来源于stack exchange,提问作者Michael
相关产品推荐
相关产品推荐

