如何使用BeautifulSoup提取XML文本并保留指定标签
用BeautifulSoup保留指定标签提取文本的解决方案
完全可以用BeautifulSoup实现你的需求,核心思路是递归遍历文档树节点,只保留你指定的<w:ins>和<w:del>标签,其余标签仅提取内部文本。具体实现步骤如下:
实现代码
首先定义一个递归处理函数,用来遍历节点并构建目标输出:
from bs4 import BeautifulSoup def extract_text_with_target_tags(soup, target_tags): """ 提取文本并保留指定的带命名空间标签 :param soup: BeautifulSoup对象 :param target_tags: 字典,键为命名空间前缀,值为要保留的标签列表,如{'w': ['ins', 'del']} :return: 包含指定标签的文本字符串 """ result = [] for elem in soup.descendants: if isinstance(elem, str): # 处理文本节点,保留有效文本(过滤纯空白) text = elem.strip() if text: result.append(text) elif elem.prefix in target_tags and elem.name in target_tags[elem.prefix]: # 匹配到要保留的标签,拼接开始标签 start_tag = f"<{elem.prefix}:{elem.name}>" result.append(start_tag) # 递归处理标签内部内容 result.append(extract_text_with_target_tags(elem, target_tags)) # 拼接结束标签 end_tag = f"</{elem.prefix}:{elem.name}>" result.append(end_tag) else: # 其他标签,仅递归提取内部文本,不保留标签本身 result.append(extract_text_with_target_tags(elem, target_tags)) return ''.join(result)
测试示例
假设你的docx XML片段如下:
<w:document xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main"> <w:body> <w:p> Here is some text <w:ins>inserted</w:ins> into a document <w:del>e</w:del> </w:p> </w:body> </w:document>
调用函数处理:
# 解析XML xml_content = 上面的XML字符串 soup = BeautifulSoup(xml_content, 'xml') # 指定要保留的标签(w命名空间下的ins和del) target_tags = {'w': ['ins', 'del']} # 生成结果 output = extract_text_with_target_tags(soup, target_tags) print(output)
输出结果将完全符合你的预期:
Here is some text<w:ins>inserted</w:ins>into a document<w:del>e</w:del>
关键说明
- 命名空间处理:docx的XML使用
w命名空间,BeautifulSoup会将标签的命名空间前缀和标签名分开存储,因此通过elem.prefix和elem.name可以精准匹配目标标签。 - 递归遍历:函数会深入每个节点内部,确保嵌套在其他标签(比如
<w:t>)里的文本也能被正确提取,同时保留指定标签的结构。 - 空白过滤:对文本节点做了
strip()处理,避免输出中出现大量冗余空白,你可以根据需求调整这部分逻辑。
内容的提问来源于stack exchange,提问作者Jordan Smith
相关产品推荐
相关产品推荐

