You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解析XML以获取xref ref-type="bibr"标签的前置语句

问题描述

需要提取所有位于<xref ref-type="bibr">标签之前的句子,输入示例如下:

xml = "This could be due to two nonmutually exclusive possibilities: first, the HB3 DNA sequence for these genes may be substantially rearranged or completely deleted relative to the reference strain, 3D7; second, only a few of these genes may be selectively expressed, as has been proposed (&lt;xref rid=&quot;pbio-0000005-Deitsch1&quot; ref-type=&quot;bibr&quot;&gt;Deitsch et al. 2001&lt;/xref&gt;). To identify regions of genomic variability between 3D7 and HB3, we performed microarray-based comparative genomic hybridization (CGH) analysis. Array-based CGH has been performed with human cDNA and bacterial artificial chromosome-based microarrays to characterize DNA copy-number changes associated with tumorigenesis (&lt;xref rid=&quot;pbio-0000005-Gray1&quot; ref-type=&quot;bibr&quot;&gt;Gray and Collins 2000&lt;/xref&gt;; &lt;xref rid=&quot;pbio-0000005-Pollack2&quot; ref-type=&quot;bibr&quot;&gt;Pollack et al. 2002&lt;/xref&gt;)."

预期输出:

{
Array-based CGH has been performed with human cDNA and bacterial artificial chromosome-based microarrays to characterize DNA copy-number changes associated with tumorigenesis :
(&lt;xref rid=&quot;pbio-0000005-Gray1&quot; ref-type=&quot;bibr&quot;&gt;Gray and Collins 2000&lt;/xref&gt;; &lt;xref rid=&quot;pbio-0000005-Pollack2&quot; ref-type=&quot;bibr&quot;&gt;Pollack et al. 2002&lt;/xref&gt;)
}
{
This could be due to two nonmutually exclusive possibilities: first, the HB3 DNA sequence for these genes may be substantially rearranged or completely deleted relative to the reference strain, 3D7; second, only a few of these genes may be selectively expressed, as has been proposed :(&lt;xref rid=&quot;pbio-0000005-Deitsch1&quot; ref-type=&quot;bibr&quot;&gt;Deitsch et al. 2001&lt;/xref&gt;)
}

尝试了以下代码,只能获取第一个结果:

soup = bs(XML)
soup.find("xref", attrs={'ref-type': 'bibr'}).previous_element

用find_all能拿到所有引用标签,但无法直接获取每个标签对应的前置句子,尤其不知道如何处理多个引用同属一个句子的情况。

解决方案

可以通过遍历find_all返回的所有<xref>标签,结合previous_element和next_sibling处理,同时避免重复处理同一组引用的句子。具体代码如下:

from bs4 import BeautifulSoup

xml = "This could be due to two nonmutually exclusive possibilities: first, the HB3 DNA sequence for these genes may be substantially rearranged or completely deleted relative to the reference strain, 3D7; second, only a few of these genes may be selectively expressed, as has been proposed (&lt;xref rid=&quot;pbio-0000005-Deitsch1&quot; ref-type=&quot;bibr&quot;&gt;Deitsch et al. 2001&lt;/xref&gt;). To identify regions of genomic variability between 3D7 and HB3, we performed microarray-based comparative genomic hybridization (CGH) analysis. Array-based CGH has been performed with human cDNA and bacterial artificial chromosome-based microarrays to characterize DNA copy-number changes associated with tumorigenesis (&lt;xref rid=&quot;pbio-0000005-Gray1&quot; ref-type=&quot;bibr&quot;&gt;Gray and Collins 2000&lt;/xref&gt;; &lt;xref rid=&quot;pbio-0000005-Pollack2&quot; ref-type=&quot;bibr&quot;&gt;Pollack et al. 2002&lt;/xref&gt;)."

soup = BeautifulSoup(xml, 'html.parser')
# 存储已处理的句子,避免重复输出
processed_sentences = set()

for xref in soup.find_all("xref", attrs={'ref-type': 'bibr'}):
    # 获取引用标签的前一个文本节点
    pre_text = xref.previous_element.strip()
    # 处理括号开头的情况,往前找完整句子
    if pre_text.startswith('('):
        # 去掉左括号,获取前面的句子内容
        sentence_part = xref.previous_element.previous_element.strip()
        # 收集当前引用及后续同组的引用(分号分隔的情况)
        refs = [str(xref)]
        next_sib = xref.next_sibling
        while next_sib and next_sib.strip() in [';', ' ']:
            if next_sib.next_sibling and next_sib.next_sib.name == 'xref' and next_sib.next_sib.get('ref-type') == 'bibr':
                refs.append(str(next_sib.next_sibling))
                next_sib = next_sib.next_sibling.next_sibling
            else:
                break
        # 拼接引用部分
        ref_str = '; '.join(refs)
        # 组合句子和引用
        full_entry = f"{sentence_part} :\n({ref_str})"
        if full_entry not in processed_sentences:
            processed_sentences.add(full_entry)
            print("{\n" + full_entry + "\n}")
    else:
        # 单个引用的情况
        full_entry = f"{pre_text} :({str(xref)})"
        if full_entry not in processed_sentences:
            processed_sentences.add(full_entry)
            print("{\n" + full_entry + "\n}")

代码说明

  1. 用find_all获取所有ref-type="bibr"的<xref>标签,遍历每个标签。
  2. 判断引用前的文本是否以左括号开头,区分单个引用和多引用组的情况。
  3. 对于多引用组,通过next_sibling收集后续同属一组的引用标签,避免重复处理同一句子。
  4. 用processed_sentences集合存储已输出的内容,防止重复打印同一句子。

内容的提问来源于stack exchange,提问作者Loïc Borcard

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 05:35:07