You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用多进程从PubMed Central大型XML文件提取PMID

大型PMC XML文件多进程提取PMID方案

需求背景

从PubMed Central(PMC)下载的大型XML文件中,提取所有标签为article-id且属性pub-id-type值为pmid的文本内容,目标提取约300万条PMID,要求用多进程实现以提升处理效率。

实现思路

直接单进程流式处理大XML文件效率较低,利用多进程并行处理需先将大文件分割为多个结构完整的XML小片段(每个片段包含若干独立的<article>节点),再通过进程池并行解析每个小文件提取PMID,最后合并结果。

代码实现

第一步:分割大XML文件

将原始大文件按指定数量的<article>节点分割为多个小XML文件,确保每个小文件结构完整可独立解析:

def split_pmc_xml(input_file, output_prefix, chunk_size=1000):
    """将PMC XML大文件分割为多个包含指定数量article的小文件"""
    with open(input_file, 'r', encoding='utf-8') as f:
        # 读取根节点起始标签
        root_start = next(f).strip()
        article_count = 0
        file_index = 0
        current_content = [root_start]
        
        for line in f:
            stripped_line = line.strip()
            if stripped_line.startswith('<article'):
                article_count += 1
                # 达到指定文章数量时写入文件
                if article_count > chunk_size:
                    current_content.append('</pmc-articleset>')
                    with open(f'{output_prefix}_part_{file_index}.xml', 'w', encoding='utf-8') as out_f:
                        out_f.write('\n'.join(current_content))
                    current_content = [root_start]
                    article_count = 1
                    file_index += 1
            current_content.append(stripped_line)
            
            if stripped_line == '</pmc-articleset>':
                break
        
        # 写入最后一批内容
        if current_content:
            if current_content[-1] != '</pmc-articleset>':
                current_content.append('</pmc-articleset>')
            with open(f'{output_prefix}_part_{file_index}.xml', 'w', encoding='utf-8') as out_f:
                out_f.write('\n'.join(current_content))

第二步:多进程提取PMID

使用进程池并行处理所有分割后的小文件,提取符合条件的PMID并合并结果:

import xml.etree.ElementTree as ET
from multiprocessing import Pool
import os

def extract_pmids_from_file(xml_file):
    """从单个XML文件中提取符合条件的PMID"""
    pmids = []
    iterator = ET.iterparse(xml_file, events=("end",))
    for event, elem in iterator:
        # 匹配目标标签和属性
        if elem.tag == 'article-id' and elem.attrib.get('pub-id-type') == 'pmid':
            pmid_text = elem.text.strip() if elem.text else None
            if pmid_text:
                pmids.append(pmid_text)
        # 清理元素和父节点引用,释放内存
        elem.clear()
        for ancestor in elem.iterancestors():
            ancestor.remove(elem)
    return pmids

def main():
    input_xml = 'my_file.xml'
    split_prefix = 'pmc_split'
    
    # 检查是否已分割文件,未分割则执行分割
    if not any(f.startswith(split_prefix) for f in os.listdir('.')):
        split_pmc_xml(input_xml, split_prefix, chunk_size=1000)
    
    # 获取所有分割后的小文件
    split_files = [f for f in os.listdir('.') if f.startswith(split_prefix) and f.endswith('.xml')]
    
    # 启动进程池并行处理
    with Pool(processes=os.cpu_count()) as pool:
        results = pool.map(extract_pmids_from_file, split_files)
    
    # 合并所有结果,可选去重
    all_pmids = []
    for res in results:
        all_pmids.extend(res)
    all_pmids = list(set(all_pmids))  # 若确认无重复可去掉此步骤
    
    # 保存结果到文本文件
    with open('extracted_pmids.txt', 'w', encoding='utf-8') as f:
        f.write('\n'.join(all_pmids))
    
    print(f"共提取 {len(all_pmids)} 条PMID,已保存至extracted_pmids.txt")

if __name__ == '__main__':
    main()

关键优化说明

  • 分割粒度调整:chunk_size参数可根据内存情况调整,内存充足时可增大数值(如5000),减少进程创建开销;内存紧张时缩小数值,避免单进程内存占用过高。
  • 内存泄漏防护:解析时主动清理元素及父节点引用,避免流式解析过程中内存持续增长。
  • 进程池效率:使用os.cpu_count()自动匹配CPU核心数,最大化利用多核资源提升处理速度。

内容的提问来源于stack exchange,提问作者Mathew

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 11:35:28