You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup提取XML文本并保留指定标签

用BeautifulSoup保留指定标签提取文本的解决方案

完全可以用BeautifulSoup实现你的需求,核心思路是递归遍历文档树节点,只保留你指定的<w:ins>和<w:del>标签,其余标签仅提取内部文本。具体实现步骤如下:

实现代码

首先定义一个递归处理函数,用来遍历节点并构建目标输出:

from bs4 import BeautifulSoup

def extract_text_with_target_tags(soup, target_tags):
    """
    提取文本并保留指定的带命名空间标签
    :param soup: BeautifulSoup对象
    :param target_tags: 字典,键为命名空间前缀,值为要保留的标签列表,如{'w': ['ins', 'del']}
    :return: 包含指定标签的文本字符串
    """
    result = []
    for elem in soup.descendants:
        if isinstance(elem, str):
            # 处理文本节点,保留有效文本(过滤纯空白)
            text = elem.strip()
            if text:
                result.append(text)
        elif elem.prefix in target_tags and elem.name in target_tags[elem.prefix]:
            # 匹配到要保留的标签,拼接开始标签
            start_tag = f"<{elem.prefix}:{elem.name}>"
            result.append(start_tag)
            # 递归处理标签内部内容
            result.append(extract_text_with_target_tags(elem, target_tags))
            # 拼接结束标签
            end_tag = f"</{elem.prefix}:{elem.name}>"
            result.append(end_tag)
        else:
            # 其他标签,仅递归提取内部文本,不保留标签本身
            result.append(extract_text_with_target_tags(elem, target_tags))
    return ''.join(result)

测试示例

假设你的docx XML片段如下:

<w:document xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main">
  <w:body>
    <w:p>
      Here is some text 
      <w:ins>inserted</w:ins> 
      into a document
      <w:del>e</w:del>
    </w:p>
  </w:body>
</w:document>

调用函数处理:

# 解析XML
xml_content = 上面的XML字符串
soup = BeautifulSoup(xml_content, 'xml')

# 指定要保留的标签(w命名空间下的ins和del)
target_tags = {'w': ['ins', 'del']}

# 生成结果
output = extract_text_with_target_tags(soup, target_tags)
print(output)

输出结果将完全符合你的预期:

Here is some text<w:ins>inserted</w:ins>into a document<w:del>e</w:del>

关键说明

  1. 命名空间处理:docx的XML使用w命名空间,BeautifulSoup会将标签的命名空间前缀和标签名分开存储,因此通过elem.prefix和elem.name可以精准匹配目标标签。
  2. 递归遍历:函数会深入每个节点内部,确保嵌套在其他标签(比如<w:t>)里的文本也能被正确提取,同时保留指定标签的结构。
  3. 空白过滤:对文本节点做了strip()处理,避免输出中出现大量冗余空白,你可以根据需求调整这部分逻辑。

内容的提问来源于stack exchange,提问作者Jordan Smith

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 04:43:32