You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Beautiful Soup从GROBID的TEI-XML中提取参考文献标题

解决TEI-XML参考文献标题纯文本提取问题

首先明确你使用的XML解析库(常用为lxml或BeautifulSoup),以下是对应改进方案:

情况1:使用lxml解析TEI-XML

若代码基于lxml.etree实现,原reference_titles返回的是Element对象列表,可通过itertext()拼接嵌套文本,或用XPath的string()函数直接提取:

改进代码示例

from lxml import etree

class TEIFile:
    def __init__(self, tei_path):
        self.tree = etree.parse(tei_path)
        # GROBID输出的TEI默认带此命名空间
        self.tei_ns = {'tei': 'http://www.tei-c.org/ns/1.0'}

    @property
    def reference_titles(self):
        # 定位所有参考文献的<title>标签
        title_elements = self.tree.xpath('//tei:biblStruct/tei:title', namespaces=self.tei_ns)
        # 提取每个标签下的纯文本,兼容嵌套格式化标签
        return [''.join(elem.itertext()).strip() for elem in title_elements]

也可直接用XPath表达式一次性返回文本列表:

@property
def reference_titles(self):
    return self.tree.xpath('//tei:biblStruct/tei:title/string()', namespaces=self.tei_ns)

情况2:使用BeautifulSoup解析TEI-XML

若用BeautifulSoup,利用get_text()可直接提取标签及子标签的所有文本,还能自动去除首尾空白:

改进代码示例

from bs4 import BeautifulSoup

class TEIFile:
    def __init__(self, tei_path):
        with open(tei_path, 'r', encoding='utf-8') as f:
            # 用XML专用解析器处理TEI内容
            self.soup = BeautifulSoup(f.read(), 'lxml-xml')

    @property
    def reference_titles(self):
        # 遍历所有参考文献节点,提取对应标题文本
        bibl_nodes = self.soup.find_all('biblStruct', recursive=True)
        return [node.find('title').get_text(strip=True) for node in bibl_nodes if node.find('title')]

关键说明

  • GROBID生成的TEI-XML中,参考文献标题通常嵌套在<biblStruct>节点下,需准确定位层级,避免误提取正文标题或其他无关内容。
  • itertext()或get_text()可处理内可能存在的格式化子标签(如<code><hi></code>、<code><ref></code>),确保提取完整标题文本。</li> </ul> <p>内容的提问来源于stack exchange,提问作者keeran_q789</p>
相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 20:35:16