You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup和lxml处理XML时标签意外修改的问题求助

解决非标准XML标签被修改的问题

你的问题核心在于用HTML解析器(lxml)处理不符合标准规范的私有XML。这类私有XML的标签格式(比如<check_type:"Windows">)不是标准HTML/XML的属性写法,BeautifulSoup的HTML解析器会将其识别为格式错误,在解析和序列化过程中破坏原有结构,导致标签内容丢失。

以下是两种可行的解决方案:

方案1:用正则直接移除目标标签(推荐,不改动其他内容)

既然只需要移除包含指定内容的<custom_item>标签,完全可以跳过解析整个文档,用正则匹配并删除目标标签,这样不会影响其他非标准标签的结构。

修改后的代码示例:

import re

def build_audit_file(items, file_path):
    """
    Removes specified item from audit file.
    Params:
    items - list - passed-in items to remove
    file_path - str - passed-in filename of the audit file we're presently working
    """
    # 读取文件内容,用with自动管理文件开闭
    with open(file_path, 'r') as f:
        file_data = f.read()
    
    # 遍历需要移除的内容,匹配包含指定文本的custom_item标签(支持跨多行)
    for item in items:
        pattern = re.compile(r'<custom_item>.*?' + re.escape(item) + r'.*?</custom_item>', re.DOTALL)
        file_data = pattern.sub('', file_data)
    
    # 生成输出文件名,修复原代码中file.name的错误(参数是字符串路径,无name属性)
    output_file_name = root_directory + r"\output\XFP-" + file_path.split("\\")[-1]
    
    # 写入输出文件,用w模式覆盖而非追加
    with open(output_file_name, 'w') as output_file:
        output_file.write(file_data)

方案2:调整解析器,尽量保留原始结构(适合需解析其他内容的场景)

如果必须用解析器处理文档,尝试使用Python内置的html.parser,并避免自动添加HTML包裹标签,但这种方式仍可能对极端非标准标签有兼容性问题:

from bs4 import BeautifulSoup

def build_audit_file(items, file_path):
    """
    Removes specified item from audit file.
    Params:
    items - list - passed-in items to remove
    file_path - str - passed-in filename of the audit file we're presently working
    """
    with open(file_path, 'r') as f:
        file_data = f.read()
    
    # 使用内置html.parser,减少对非标准格式的修正
    soup = BeautifulSoup(file_data, "html.parser")
    for item in items:
        for tag in soup.find_all("custom_item", string=lambda text: text and item in text):
            tag.extract()
    
    # 直接拼接原始内容,避免自动添加html/body标签
    new_audit_file_data = ''.join(str(content) for content in soup.contents)
    
    output_file_name = root_directory + r"\output\XFP-" + file_path.split("\\")[-1]
    with open(output_file_name, 'w') as output_file:
        output_file.write(new_audit_file_data)

原代码的额外问题修复提示

  • 原参数file是字符串路径,但代码中误用了file.name,字符串没有该属性,需用file_path.split("\\")[-1]获取文件名
  • 原代码用"a"模式追加写入输出文件,可能导致重复内容,建议改用"w"模式覆盖写入

内容的提问来源于stack exchange,提问作者user23254098

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 11:27:23