如何将PubMed Central文章提取至Python DataFrame并构建指定列
处理PubMed Central文章生成指定列的DataFrame
依赖准备
先安装必要的Python库:
pip install pandas lxml
如果需要在线获取PMC文章,额外安装biopython:
pip install biopython
本地XML文件解析方案
单篇XML解析函数
针对PMC标准XML格式(带JATS命名空间),编写字段提取函数:
import pandas as pd from lxml import etree import os def parse_pmc_xml(xml_path): # 加载XML并指定命名空间 tree = etree.parse(xml_path) ns = {'pmc': 'http://www.ncbi.nlm.nih.gov/JATS1'} # 提取PMC ID pmc_id_elem = tree.find('.//pmc:article-id[@pub-id-type="pmc"]', ns) pmc_id = pmc_id_elem.text if pmc_id_elem else 'N/A' # 提取标题 title_elem = tree.find('.//pmc:article-title', ns) title = title_elem.text.strip() if title_elem else 'N/A' # 提取摘要(处理分段) abstract = '' abstract_node = tree.find('.//pmc:abstract', ns) if abstract_node: abstract_paras = abstract_node.findall('.//pmc:p', ns) abstract = '\n'.join([p.text.strip() for p in abstract_paras if p.text]) # 提取全文(递归查找所有正文段落) full_text = '' body_paras = tree.findall('.//pmc:body//pmc:p', ns) full_text = '\n'.join([p.text.strip() for p in body_paras if p.text]) # 提取作者(格式统一为"LastName, FirstName") authors = [] author_nodes = tree.findall('.//pmc:contrib[@contrib-type="author"]', ns) for node in author_nodes: surname = node.find('.//pmc:surname', ns) given_names = node.find('.//pmc:given-names', ns) if surname and given_names: authors.append(f"{surname.text.strip()}, {given_names.text.strip()}") authors = '; '.join(authors) if authors else 'N/A' return { 'pmc id': pmc_id, 'title': title, 'abstract': abstract, 'full-text': full_text, 'authors': authors }
批量处理生成DataFrame
遍历指定文件夹下的所有PMC XML文件,生成目标DataFrame:
def batch_process_pmc(folder_path): article_data_list = [] for filename in os.listdir(folder_path): if filename.lower().endswith('.xml'): file_path = os.path.join(folder_path, filename) try: article_data = parse_pmc_xml(file_path) article_data_list.append(article_data) except Exception as e: print(f"处理文件 {filename} 失败: {str(e)}") continue return pd.DataFrame(article_data_list) # 调用示例 df = batch_process_pmc('./pmc_articles_folder') # 查看结果 print(df.head()) # 保存为CSV df.to_csv('pmc_articles.csv', index=False)
在线获取PMC文章方案
如果无需本地文件,直接通过PMC ID在线获取:
from Bio import PMC import pandas as pd def fetch_pmc_by_id(pmc_id): try: article = PMC.get_article(pmc_id) # 提取标题 title = article.title if hasattr(article, 'title') else 'N/A' # 提取摘要 abstract = article.abstract if hasattr(article, 'abstract') else '' # 提取全文 full_text = '\n'.join([para.text for para in article.body_paragraphs]) if hasattr(article, 'body_paragraphs') else '' # 提取作者 authors = [] if hasattr(article, 'authors'): for author in article.authors: if hasattr(author, 'last_name') and hasattr(author, 'first_name'): authors.append(f"{author.last_name}, {author.first_name}") authors = '; '.join(authors) if authors else 'N/A' return { 'pmc id': pmc_id, 'title': title, 'abstract': abstract, 'full-text': full_text, 'authors': authors } except Exception as e: print(f"获取PMC {pmc_id} 失败: {str(e)}") return None # 示例:批量获取多个PMC ID pmc_ids = ['PMC123456', 'PMC789012'] data = [fetch_pmc_by_id(pid) for pid in pmc_ids if fetch_pmc_by_id(pid)] df = pd.DataFrame(data)
常见问题修正
- 命名空间遗漏:PMC XML默认使用JATS命名空间,必须在XPath中指定
ns参数,否则会找不到任何节点。 - 字段缺失处理:部分文章可能无摘要、作者信息,代码中已添加
if判断避免抛出异常。 - 全文提取不全:使用
.//pmc:body//pmc:p递归查找所有正文段落,覆盖嵌套在<sec>标签内的内容。
内容的提问来源于stack exchange,提问作者Learning
相关产品推荐
相关产品推荐

