使用BeautifulSoup和lxml处理XML时标签意外修改的问题求助
解决非标准XML标签被修改的问题
你的问题核心在于用HTML解析器(lxml)处理不符合标准规范的私有XML。这类私有XML的标签格式(比如<check_type:"Windows">)不是标准HTML/XML的属性写法,BeautifulSoup的HTML解析器会将其识别为格式错误,在解析和序列化过程中破坏原有结构,导致标签内容丢失。
以下是两种可行的解决方案:
方案1:用正则直接移除目标标签(推荐,不改动其他内容)
既然只需要移除包含指定内容的<custom_item>标签,完全可以跳过解析整个文档,用正则匹配并删除目标标签,这样不会影响其他非标准标签的结构。
修改后的代码示例:
import re def build_audit_file(items, file_path): """ Removes specified item from audit file. Params: items - list - passed-in items to remove file_path - str - passed-in filename of the audit file we're presently working """ # 读取文件内容,用with自动管理文件开闭 with open(file_path, 'r') as f: file_data = f.read() # 遍历需要移除的内容,匹配包含指定文本的custom_item标签(支持跨多行) for item in items: pattern = re.compile(r'<custom_item>.*?' + re.escape(item) + r'.*?</custom_item>', re.DOTALL) file_data = pattern.sub('', file_data) # 生成输出文件名,修复原代码中file.name的错误(参数是字符串路径,无name属性) output_file_name = root_directory + r"\output\XFP-" + file_path.split("\\")[-1] # 写入输出文件,用w模式覆盖而非追加 with open(output_file_name, 'w') as output_file: output_file.write(file_data)
方案2:调整解析器,尽量保留原始结构(适合需解析其他内容的场景)
如果必须用解析器处理文档,尝试使用Python内置的html.parser,并避免自动添加HTML包裹标签,但这种方式仍可能对极端非标准标签有兼容性问题:
from bs4 import BeautifulSoup def build_audit_file(items, file_path): """ Removes specified item from audit file. Params: items - list - passed-in items to remove file_path - str - passed-in filename of the audit file we're presently working """ with open(file_path, 'r') as f: file_data = f.read() # 使用内置html.parser,减少对非标准格式的修正 soup = BeautifulSoup(file_data, "html.parser") for item in items: for tag in soup.find_all("custom_item", string=lambda text: text and item in text): tag.extract() # 直接拼接原始内容,避免自动添加html/body标签 new_audit_file_data = ''.join(str(content) for content in soup.contents) output_file_name = root_directory + r"\output\XFP-" + file_path.split("\\")[-1] with open(output_file_name, 'w') as output_file: output_file.write(new_audit_file_data)
原代码的额外问题修复提示
- 原参数
file是字符串路径,但代码中误用了file.name,字符串没有该属性,需用file_path.split("\\")[-1]获取文件名 - 原代码用
"a"模式追加写入输出文件,可能导致重复内容,建议改用"w"模式覆盖写入
内容的提问来源于stack exchange,提问作者user23254098
相关产品推荐
相关产品推荐

