如何用多进程从PubMed Central大型XML文件提取PMID
大型PMC XML文件多进程提取PMID方案
需求背景
从PubMed Central(PMC)下载的大型XML文件中,提取所有标签为article-id且属性pub-id-type值为pmid的文本内容,目标提取约300万条PMID,要求用多进程实现以提升处理效率。
实现思路
直接单进程流式处理大XML文件效率较低,利用多进程并行处理需先将大文件分割为多个结构完整的XML小片段(每个片段包含若干独立的<article>节点),再通过进程池并行解析每个小文件提取PMID,最后合并结果。
代码实现
第一步:分割大XML文件
将原始大文件按指定数量的<article>节点分割为多个小XML文件,确保每个小文件结构完整可独立解析:
def split_pmc_xml(input_file, output_prefix, chunk_size=1000): """将PMC XML大文件分割为多个包含指定数量article的小文件""" with open(input_file, 'r', encoding='utf-8') as f: # 读取根节点起始标签 root_start = next(f).strip() article_count = 0 file_index = 0 current_content = [root_start] for line in f: stripped_line = line.strip() if stripped_line.startswith('<article'): article_count += 1 # 达到指定文章数量时写入文件 if article_count > chunk_size: current_content.append('</pmc-articleset>') with open(f'{output_prefix}_part_{file_index}.xml', 'w', encoding='utf-8') as out_f: out_f.write('\n'.join(current_content)) current_content = [root_start] article_count = 1 file_index += 1 current_content.append(stripped_line) if stripped_line == '</pmc-articleset>': break # 写入最后一批内容 if current_content: if current_content[-1] != '</pmc-articleset>': current_content.append('</pmc-articleset>') with open(f'{output_prefix}_part_{file_index}.xml', 'w', encoding='utf-8') as out_f: out_f.write('\n'.join(current_content))
第二步:多进程提取PMID
使用进程池并行处理所有分割后的小文件,提取符合条件的PMID并合并结果:
import xml.etree.ElementTree as ET from multiprocessing import Pool import os def extract_pmids_from_file(xml_file): """从单个XML文件中提取符合条件的PMID""" pmids = [] iterator = ET.iterparse(xml_file, events=("end",)) for event, elem in iterator: # 匹配目标标签和属性 if elem.tag == 'article-id' and elem.attrib.get('pub-id-type') == 'pmid': pmid_text = elem.text.strip() if elem.text else None if pmid_text: pmids.append(pmid_text) # 清理元素和父节点引用,释放内存 elem.clear() for ancestor in elem.iterancestors(): ancestor.remove(elem) return pmids def main(): input_xml = 'my_file.xml' split_prefix = 'pmc_split' # 检查是否已分割文件,未分割则执行分割 if not any(f.startswith(split_prefix) for f in os.listdir('.')): split_pmc_xml(input_xml, split_prefix, chunk_size=1000) # 获取所有分割后的小文件 split_files = [f for f in os.listdir('.') if f.startswith(split_prefix) and f.endswith('.xml')] # 启动进程池并行处理 with Pool(processes=os.cpu_count()) as pool: results = pool.map(extract_pmids_from_file, split_files) # 合并所有结果,可选去重 all_pmids = [] for res in results: all_pmids.extend(res) all_pmids = list(set(all_pmids)) # 若确认无重复可去掉此步骤 # 保存结果到文本文件 with open('extracted_pmids.txt', 'w', encoding='utf-8') as f: f.write('\n'.join(all_pmids)) print(f"共提取 {len(all_pmids)} 条PMID,已保存至extracted_pmids.txt") if __name__ == '__main__': main()
关键优化说明
- 分割粒度调整:
chunk_size参数可根据内存情况调整,内存充足时可增大数值(如5000),减少进程创建开销;内存紧张时缩小数值,避免单进程内存占用过高。 - 内存泄漏防护:解析时主动清理元素及父节点引用,避免流式解析过程中内存持续增长。
- 进程池效率:使用
os.cpu_count()自动匹配CPU核心数,最大化利用多核资源提升处理速度。
内容的提问来源于stack exchange,提问作者Mathew
相关产品推荐
相关产品推荐

