如何用Python移除XML注释行中的多余换行与空格(lxml场景)
解决XML注释中的多余换行与空格问题
问题场景
使用lxml将<Necessary/>标签转为注释时,生成的注释出现换行问题,-->被拆分到下一行,需要消除注释内部的换行及多余空格,让注释保持紧凑格式。
当前使用的lxml代码
tc.getparent().replace(tc, etree.Comment(etree.tostring(tc))) print(etree.tostring(doc2).decode())
生成的异常XML格式
<List> <Item> <Price> <Amount>100</Amount> <Next_Item> <Name>Apple</Name> <!--<Necessary/> --> </Next_Item> <Next_Item> <Name>Orange</Name> <!--<Necessary/> --> </Next_Item> </Price> </Item> </List>
已尝试的无效处理(BeautifulSoup)
soup = BeautifulSoup(open('XML1.xml', 'r'), 'xml') for elem in soup.find_all(): if elem.string is not None: elem.string = elem.string.strip()
期望的紧凑XML格式
<List> <Item> <Price> <Amount>100</Amount> <Next_Item> <Name>Apple</Name> <!--<Necessary/>--> </Next_Item> <Next_Item> <Name>Orange</Name> <!--<Necessary/>--> </Next_Item> </Price> </Item> </List>
有效解决方案
方案1:从源头生成紧凑注释文本
在创建注释前,先对标签转字符串后的内容做清理,去除换行和多余空白,再传入注释对象:
# 将标签转为字符串并解码,清除所有换行和首尾空白 comment_content = etree.tostring(tc).decode().replace('\n', '').strip() # 用清理后的内容创建注释,替换原标签 tc.getparent().replace(tc, etree.Comment(comment_content)) # 输出时保持格式化,避免自动拆分注释 print(etree.tostring(doc2, encoding='unicode', pretty_print=True))
方案2:正则替换已生成的XML字符串
如果方案1仍有格式问题,可在最终XML字符串上做正则匹配替换,修复注释格式:
import re # 生成格式化后的XML字符串 xml_output = etree.tostring(doc2, encoding='unicode', pretty_print=True) # 匹配拆分的注释,替换为紧凑格式 fixed_xml = re.sub(r'<!--(.*?)\n\s*-->', r'<!--\1-->', xml_output, flags=re.DOTALL) print(fixed_xml)
说明:
- 方案1从生成注释的环节解决问题,避免lxml自动添加换行;
- 方案2针对已生成的XML内容做后处理,适合批量修复现有文件。
内容的提问来源于stack exchange,提问作者Anonymous
相关产品推荐
相关产品推荐

