如何用BeautifulSoup提取文本加标签,修复XML嵌套结构问题
问题1:为指定文本批量添加标签
原始XML片段
<p> lorem ipsum dolor <hi>a</hi>met. Thing_A, <thing>Thing_B</thing>, Thing_C </p>
期望结果
<p> lorem ipsum dolor <hi>a</hi>met. <thing>Thing_A</thing>, <thing>Thing_B</thing>, <thing>Thing_C</thing> </p>
当前代码及问题
已创建目标文本列表:
things = ['Thing_A', 'Thing_B', 'Thing_C']
尝试用BeautifulSoup查找文本并添加标签,但find(text)返回None,代码如下:
text = things[0] print(text) stelle = soup.find(text) print(stelle) tag = soup.new_tag('rs') tag.append(text) print(tag)
输出:
Thing_A None <rs>Thing_A</rs>
问题2:修正XML结构
当前错误结构
<body><tei xml:id="pa029" xmlns="http://www.tei-c.org/ns/1.0"> <teiheader> <!-- The TEI-header from before--> </teiheader> <!-- <standOff/>--> <text> <pb n="29v"></pb> <p> {some content} </p> </text> </tei> </body>
期望正确结构
<text><body>{some content} </body></text></tei>
解决思路
针对问题1:批量包裹指定文本为标签
- 拆分文本节点并匹配替换
目标文本可能和标点、空格处于同一个文本节点,直接用find(text=...)无法匹配,需要遍历文本节点并拆分处理:
该方法会自动处理已被标签包裹的内容(比如Thing_B),不会重复包裹。import re from bs4 import BeautifulSoup soup = BeautifulSoup(your_xml_content, 'xml') things = ['Thing_A', 'Thing_B', 'Thing_C'] for p_tag in soup.find_all('p'): # 遍历p标签下的所有子节点 for child in list(p_tag.contents): if child.string: original_str = child.string # 按目标文本拆分字符串 parts = re.split(f'({"|".join(things)})', original_str) new_content = [] for part in parts: if part in things: # 创建thing标签包裹匹配项 thing_tag = soup.new_tag('thing') thing_tag.string = part new_content.append(thing_tag) elif part.strip(): new_content.append(part) # 替换原文本节点 child.replace_with(*new_content)
针对问题2:调整XML结构层级
- 移动节点修正层级关系
通过BeautifulSoup的节点操作方法,把内容从错误的外层body移到text内部的body中:
注意使用soup = BeautifulSoup(your_xml_content, 'xml') # 获取核心节点 tei_tag = soup.find('tei') text_tag = tei_tag.find('text') outer_body = soup.find('body') if outer_body and text_tag: # 将外层body的内容移到text标签下 for child in outer_body.contents: if child != tei_tag: text_tag.append(child) # 删除原外层body标签 outer_body.decompose() # 在text内部创建新的body标签,并迁移内容 new_body = soup.new_tag('body') # 遍历text的子节点,将需要的内容移入新body(可根据需求调整过滤规则) for child in list(text_tag.contents): if child.name not in ['teiheader', 'pb']: new_body.append(child) text_tag.append(new_body)xml解析器处理带命名空间的TEI文件,避免解析异常。
内容的提问来源于stack exchange,提问作者KWunsch
相关产品推荐
相关产品推荐

