You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Pandas写入BeautifulSoup时HTML表格的转义问题?

解决Pandas DataFrame插入BeautifulSoup后HTML表格双重转义问题

出现<这类双重转义字符,核心原因是Pandas生成HTML时默认转义特殊字符,而BeautifulSoup处理时又再次转义,导致字符被转义两次。以下是两种直接可行的解决方法:

方法1:禁用Pandas的HTML转义(推荐)

让Pandas生成HTML时不转义特殊字符,避免后续BeautifulSoup重复转义:

from bs4 import BeautifulSoup
import pandas as pd

# 读取原始HTML文件
with open('html-in.html', 'r', encoding='utf-8') as f:
    soup = BeautifulSoup(f.read(), 'html.parser')

# 提取页面中的表格转为DataFrame
df = pd.read_html(str(soup))[0]

# 添加新行数据
df.loc[len(df)] = ["新增内容1", "新增内容2", "新增内容3"]

# 生成不转义的表格HTML,并转为BeautifulSoup标签对象
new_table = BeautifulSoup(df.to_html(escape=False), 'html.parser').find('table')

# 替换页面中原有的表格
original_table = soup.find('table')
original_table.replace_with(new_table)

# 写入新HTML文件
with open('html-out.html', 'w', encoding='utf-8') as f:
    f.write(str(soup))

方法2:解码已转义的HTML后再处理

如果需要保留Pandas的默认转义逻辑,可以先将Pandas生成的HTML解码一次,再交给BeautifulSoup处理:

from bs4 import BeautifulSoup
import pandas as pd
import html

# 读取原始HTML文件
with open('html-in.html', 'r', encoding='utf-8') as f:
    soup = BeautifulSoup(f.read(), 'html.parser')

# 提取表格转为DataFrame
df = pd.read_html(str(soup))[0]
df.loc[len(df)] = ["新增内容1", "新增内容2", "新增内容3"]

# 生成带转义的表格HTML,解码后再转为BS标签
escaped_table_html = df.to_html()
unescaped_table_html = html.unescape(escaped_table_html)
new_table = BeautifulSoup(unescaped_table_html, 'html.parser').find('table')

# 替换并写入文件
original_table = soup.find('table')
original_table.replace_with(new_table)
with open('html-out.html', 'w', encoding='utf-8') as f:
    f.write(str(soup))

注意:写入文件时直接使用str(soup)或soup.prettify(),不要使用额外的转义方法,避免再次触发转义问题。

内容的提问来源于stack exchange,提问作者Bas Mooijman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 08:46:59