You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将PubMed Central文章提取至Python DataFrame并构建指定列

处理PubMed Central文章生成指定列的DataFrame

依赖准备

先安装必要的Python库:

pip install pandas lxml

如果需要在线获取PMC文章,额外安装biopython:

pip install biopython

本地XML文件解析方案

单篇XML解析函数

针对PMC标准XML格式(带JATS命名空间),编写字段提取函数:

import pandas as pd
from lxml import etree
import os

def parse_pmc_xml(xml_path):
    # 加载XML并指定命名空间
    tree = etree.parse(xml_path)
    ns = {'pmc': 'http://www.ncbi.nlm.nih.gov/JATS1'}
    
    # 提取PMC ID
    pmc_id_elem = tree.find('.//pmc:article-id[@pub-id-type="pmc"]', ns)
    pmc_id = pmc_id_elem.text if pmc_id_elem else 'N/A'
    
    # 提取标题
    title_elem = tree.find('.//pmc:article-title', ns)
    title = title_elem.text.strip() if title_elem else 'N/A'
    
    # 提取摘要(处理分段)
    abstract = ''
    abstract_node = tree.find('.//pmc:abstract', ns)
    if abstract_node:
        abstract_paras = abstract_node.findall('.//pmc:p', ns)
        abstract = '\n'.join([p.text.strip() for p in abstract_paras if p.text])
    
    # 提取全文(递归查找所有正文段落)
    full_text = ''
    body_paras = tree.findall('.//pmc:body//pmc:p', ns)
    full_text = '\n'.join([p.text.strip() for p in body_paras if p.text])
    
    # 提取作者(格式统一为"LastName, FirstName")
    authors = []
    author_nodes = tree.findall('.//pmc:contrib[@contrib-type="author"]', ns)
    for node in author_nodes:
        surname = node.find('.//pmc:surname', ns)
        given_names = node.find('.//pmc:given-names', ns)
        if surname and given_names:
            authors.append(f"{surname.text.strip()}, {given_names.text.strip()}")
    authors = '; '.join(authors) if authors else 'N/A'
    
    return {
        'pmc id': pmc_id,
        'title': title,
        'abstract': abstract,
        'full-text': full_text,
        'authors': authors
    }

批量处理生成DataFrame

遍历指定文件夹下的所有PMC XML文件,生成目标DataFrame:

def batch_process_pmc(folder_path):
    article_data_list = []
    for filename in os.listdir(folder_path):
        if filename.lower().endswith('.xml'):
            file_path = os.path.join(folder_path, filename)
            try:
                article_data = parse_pmc_xml(file_path)
                article_data_list.append(article_data)
            except Exception as e:
                print(f"处理文件 {filename} 失败: {str(e)}")
                continue
    return pd.DataFrame(article_data_list)

# 调用示例
df = batch_process_pmc('./pmc_articles_folder')
# 查看结果
print(df.head())
# 保存为CSV
df.to_csv('pmc_articles.csv', index=False)

在线获取PMC文章方案

如果无需本地文件,直接通过PMC ID在线获取:

from Bio import PMC
import pandas as pd

def fetch_pmc_by_id(pmc_id):
    try:
        article = PMC.get_article(pmc_id)
        # 提取标题
        title = article.title if hasattr(article, 'title') else 'N/A'
        # 提取摘要
        abstract = article.abstract if hasattr(article, 'abstract') else ''
        # 提取全文
        full_text = '\n'.join([para.text for para in article.body_paragraphs]) if hasattr(article, 'body_paragraphs') else ''
        # 提取作者
        authors = []
        if hasattr(article, 'authors'):
            for author in article.authors:
                if hasattr(author, 'last_name') and hasattr(author, 'first_name'):
                    authors.append(f"{author.last_name}, {author.first_name}")
        authors = '; '.join(authors) if authors else 'N/A'
        
        return {
            'pmc id': pmc_id,
            'title': title,
            'abstract': abstract,
            'full-text': full_text,
            'authors': authors
        }
    except Exception as e:
        print(f"获取PMC {pmc_id} 失败: {str(e)}")
        return None

# 示例:批量获取多个PMC ID
pmc_ids = ['PMC123456', 'PMC789012']
data = [fetch_pmc_by_id(pid) for pid in pmc_ids if fetch_pmc_by_id(pid)]
df = pd.DataFrame(data)

常见问题修正

  1. 命名空间遗漏:PMC XML默认使用JATS命名空间,必须在XPath中指定ns参数,否则会找不到任何节点。
  2. 字段缺失处理:部分文章可能无摘要、作者信息,代码中已添加if判断避免抛出异常。
  3. 全文提取不全:使用.//pmc:body//pmc:p递归查找所有正文段落,覆盖嵌套在<sec>标签内的内容。

内容的提问来源于stack exchange,提问作者Learning

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 12:18:27