You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取文本加标签,修复XML嵌套结构问题

问题1:为指定文本批量添加标签

原始XML片段

<p> 
    lorem ipsum dolor <hi>a</hi>met. 
    Thing_A, <thing>Thing_B</thing>, Thing_C 
    </p>

期望结果

<p> lorem ipsum dolor <hi>a</hi>met. 
<thing>Thing_A</thing>, <thing>Thing_B</thing>, <thing>Thing_C</thing>
</p>

当前代码及问题

已创建目标文本列表:

things = ['Thing_A', 'Thing_B', 'Thing_C']

尝试用BeautifulSoup查找文本并添加标签,但find(text)返回None,代码如下:

text = things[0]
print(text)
stelle = soup.find(text)
print(stelle)
tag = soup.new_tag('rs')
tag.append(text)
print(tag)

输出:

Thing_A
None
<rs>Thing_A</rs>

问题2:修正XML结构

当前错误结构

<body><tei xml:id="pa029" xmlns="http://www.tei-c.org/ns/1.0">
    <teiheader>
<!-- The TEI-header from before-->
    </teiheader>
    <!--  <standOff/>-->
    <text>
    <pb n="29v"></pb>
    <p>
 {some content}
</p>
</text>
</tei>
</body>

期望正确结构

<text><body>{some content} </body></text></tei>

解决思路

针对问题1:批量包裹指定文本为标签

  • 拆分文本节点并匹配替换
    目标文本可能和标点、空格处于同一个文本节点,直接用find(text=...)无法匹配,需要遍历文本节点并拆分处理:
    import re
    from bs4 import BeautifulSoup
    
    soup = BeautifulSoup(your_xml_content, 'xml')
    things = ['Thing_A', 'Thing_B', 'Thing_C']
    
    for p_tag in soup.find_all('p'):
        # 遍历p标签下的所有子节点
        for child in list(p_tag.contents):
            if child.string:
                original_str = child.string
                # 按目标文本拆分字符串
                parts = re.split(f'({"|".join(things)})', original_str)
                new_content = []
                for part in parts:
                    if part in things:
                        # 创建thing标签包裹匹配项
                        thing_tag = soup.new_tag('thing')
                        thing_tag.string = part
                        new_content.append(thing_tag)
                    elif part.strip():
                        new_content.append(part)
                # 替换原文本节点
                child.replace_with(*new_content)
    
    该方法会自动处理已被标签包裹的内容(比如Thing_B),不会重复包裹。

针对问题2:调整XML结构层级

  • 移动节点修正层级关系
    通过BeautifulSoup的节点操作方法,把内容从错误的外层body移到text内部的body中:
    soup = BeautifulSoup(your_xml_content, 'xml')
    
    # 获取核心节点
    tei_tag = soup.find('tei')
    text_tag = tei_tag.find('text')
    outer_body = soup.find('body')
    
    if outer_body and text_tag:
        # 将外层body的内容移到text标签下
        for child in outer_body.contents:
            if child != tei_tag:
                text_tag.append(child)
        # 删除原外层body标签
        outer_body.decompose()
        # 在text内部创建新的body标签,并迁移内容
        new_body = soup.new_tag('body')
        # 遍历text的子节点,将需要的内容移入新body(可根据需求调整过滤规则)
        for child in list(text_tag.contents):
            if child.name not in ['teiheader', 'pb']:
                new_body.append(child)
        text_tag.append(new_body)
    
    注意使用xml解析器处理带命名空间的TEI文件,避免解析异常。

内容的提问来源于stack exchange,提问作者KWunsch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 08:05:33