You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Biopython的Entrez从PubMed获取摘要时,实际获取数量远低于搜索结果数的原因排查

使用Biopython的Entrez从PubMed获取摘要时,实际获取数量远低于搜索结果数的原因排查

从你的代码和结果来看,5万多条搜索结果最后只拿到不到1万条,主要是两个核心问题导致的,我来一步步拆解:

1. 单次efetch请求仅获取了部分记录(最主要原因)

PubMed的Entrez API对单次efetch请求有返回数量限制(通常单次最多返回10,000条记录),而你的搜索结果有55,574条,但代码里只做了一次efetch调用,没有分批次获取,所以实际只下载了第一批约1万条记录,剩下的4万多条根本没下载到——你最后输出的Fetched 9946 available abstracts刚好接近API单次返回上限,也验证了这一点。

2. 日期提取逻辑导致有效记录被误删

你的代码依赖PHST字段提取年月日,但这个字段是PubMed的出版历史字段(比如包含接受日期、在线发表日期等),并不是所有文献都有这个字段,很多文献的日期信息存在DP(出版日期)或PY(出版年份)字段里。

当遇到没有PHST的记录时,你把年/月/日设为None,最后又用df.dropna()把这些记录全部删掉了——哪怕这些记录是有摘要的有效数据,这进一步减少了最终的记录数。


修复方案

首先,修改efetch代码,分批次下载所有记录

把原来一次性下载的代码改成循环分批次获取,每次取1000条(或10000条,不超API限制即可):

# 替换原来从handle = Entrez.efetch(...)到写文件的所有代码
batch_size = 1000  # 每次下载1000条,可按需调整为10000
with open('./abstracts_medline.txt', 'w', encoding="utf-8") as f:
    for start in range(0, count, batch_size):
        end = min(start + batch_size, count)
        print(f"Downloading records {start+1} to {end}")
        handle = Entrez.efetch(
            db="pubmed",
            retmode="text",
            rettype='medline',
            retstart=start,
            retmax=batch_size,
            webenv=webEnv,
            query_key=queryKey
        )
        data = handle.read()
        f.write(data)
        handle.close()

然后,修复日期提取逻辑,避免误删有效记录

改用更通用的DP(出版日期)字段提取日期,同时兼容不同的日期格式,并且只删除真正无效的记录:

from datetime import datetime

def parse_pub_date(dp_str):
    # 处理PubMed DP字段的常见格式:"2023 Dec 15"、"2023 Dec"、"2023"等
    formats = ["%Y %b %d", "%Y %b", "%Y"]
    for fmt in formats:
        try:
            dt = datetime.strptime(dp_str, fmt)
            return (dt.year, dt.month if hasattr(dt, 'month') else None, 
                    dt.day if hasattr(dt, 'day') else None)
        except ValueError:
            continue
    # 解析失败时尝试提取年份
    if dp_str.split()[0].isdigit():
        return (int(dp_str.split()[0]), None, None)
    return (None, None, None)

# 替换原来的日期提取循环
pmids = []
years = []
months = []
days = []
titles = []
abstracts = []

with open('./abstracts_medline.txt', 'r', encoding="utf-8") as f:
    medline_rec = Medline.parse(f)
    for record in medline_rec:
        # 搜索时已用has abstract[FILT],所有记录理论上都有摘要
        pmids.append(record.get('PMID', None))
        # 解析出版日期
        dp = record.get('DP', '')
        year, month, day = parse_pub_date(dp)
        years.append(year)
        months.append(month)
        days.append(day)
        titles.append(record.get('TI', None))
        abstracts.append(record.get('AB', ''))

df = pd.DataFrame({
    'PMID': pmids,
    'Year': years,
    'Month': months,
    'Day': days,
    'Title': titles,
    'Abstract': abstracts
})

# 仅删除无PMID或无摘要的无效记录,保留日期不全的有效记录
df = df.dropna(subset=['PMID', 'Abstract'])

最后验证结果

修改后重新运行,你会看到:

  • 下载时会分批次打印进度,直到所有5万多条记录完成下载
  • 最终的Fetched X available abstracts会接近55574,和搜索结果数基本一致

额外小提示

  • 因为你已经用了has abstract[FILT]过滤器,PubMed会保证所有搜索结果都有摘要,所以不需要再做if('AB' in record.keys())的判断
  • 如果确实需要完整的年月日,可以单独筛选日期齐全的记录,但不要直接丢弃所有日期不全的有效数据

备注:内容来源于stack exchange,提问作者Denis Sechin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 17:24:28