如何解决Pandas写入BeautifulSoup时HTML表格的转义问题?
解决Pandas DataFrame插入BeautifulSoup后HTML表格双重转义问题
出现<这类双重转义字符,核心原因是Pandas生成HTML时默认转义特殊字符,而BeautifulSoup处理时又再次转义,导致字符被转义两次。以下是两种直接可行的解决方法:
方法1:禁用Pandas的HTML转义(推荐)
让Pandas生成HTML时不转义特殊字符,避免后续BeautifulSoup重复转义:
from bs4 import BeautifulSoup import pandas as pd # 读取原始HTML文件 with open('html-in.html', 'r', encoding='utf-8') as f: soup = BeautifulSoup(f.read(), 'html.parser') # 提取页面中的表格转为DataFrame df = pd.read_html(str(soup))[0] # 添加新行数据 df.loc[len(df)] = ["新增内容1", "新增内容2", "新增内容3"] # 生成不转义的表格HTML,并转为BeautifulSoup标签对象 new_table = BeautifulSoup(df.to_html(escape=False), 'html.parser').find('table') # 替换页面中原有的表格 original_table = soup.find('table') original_table.replace_with(new_table) # 写入新HTML文件 with open('html-out.html', 'w', encoding='utf-8') as f: f.write(str(soup))
方法2:解码已转义的HTML后再处理
如果需要保留Pandas的默认转义逻辑,可以先将Pandas生成的HTML解码一次,再交给BeautifulSoup处理:
from bs4 import BeautifulSoup import pandas as pd import html # 读取原始HTML文件 with open('html-in.html', 'r', encoding='utf-8') as f: soup = BeautifulSoup(f.read(), 'html.parser') # 提取表格转为DataFrame df = pd.read_html(str(soup))[0] df.loc[len(df)] = ["新增内容1", "新增内容2", "新增内容3"] # 生成带转义的表格HTML,解码后再转为BS标签 escaped_table_html = df.to_html() unescaped_table_html = html.unescape(escaped_table_html) new_table = BeautifulSoup(unescaped_table_html, 'html.parser').find('table') # 替换并写入文件 original_table = soup.find('table') original_table.replace_with(new_table) with open('html-out.html', 'w', encoding='utf-8') as f: f.write(str(soup))
注意:写入文件时直接使用str(soup)或soup.prettify(),不要使用额外的转义方法,避免再次触发转义问题。
内容的提问来源于stack exchange,提问作者Bas Mooijman
相关产品推荐
相关产品推荐

