You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Shell脚本批量获取PubMed摘要时curl报错问题求助

解决批量提取PubMed摘要时的curl格式错误问题

问题背景

使用Shell脚本从PubMed批量提取摘要时,处理50-60个PMID能正常运行,但处理更多ID时出现curl: (3) URL using bad/illegal format or missing URL错误。原脚本及PMID示例如下:

原脚本

#!/bin/bash

filename=$1
count=0
while read -r line;
do (echo -n 'PMID:'"$line")
   ruby -e 'print "\t"'
   ruby -e 'print "\n"'
   curl "http://togows.dbcls.jp/entry/ncbi-pubmed/$line/abstract"
   ruby -e 'print "\n"'
count=$((count+1))
done < "$filename"

#total no. of abstracts extracted#
echo $count

期望输出格式

PMID:15531828
text text text text text text
text text text text text text

PMID:11791095
text text text text text text
text text text text text text

问题原因

  1. 无效输入行:输入文件可能包含空行、非数字字符,或Windows格式的换行符(\r),导致拼接后的URL包含无效字符
  2. 速率限制:短时间内大量请求触发服务器速率限制,返回非预期内容,引发curl解析错误

修改后的脚本

#!/bin/bash

if [ $# -ne 1 ]; then
    echo "Usage: $0 <pmid-file>"
    exit 1
fi

filename=$1
count=0
failed=0

while read -r line; do
    # 过滤空行和非数字PMID
    [[ -z "$line" || ! "$line" =~ ^[0-9]+$ ]] && continue
    # 移除Windows换行符\r
    line=${line%$'\r'}
    
    echo "PMID:$line"
    # 用bash内置命令替代冗余ruby调用
    echo ""
    
    # 发起请求,添加延迟避免速率限制
    if curl -s "http://togows.dbcls.jp/entry/ncbi-pubmed/$line/abstract"; then
        ((count++))
    else
        echo "[ERROR] Failed to fetch abstract for PMID:$line"
        ((failed++))
    fi
    
    # 添加空行分隔不同PMID的内容
    echo -e "\n"
    # 延迟1秒,避免触发服务器限流
    sleep 1
done < "$filename"

echo "Total abstracts extracted: $count"
echo "Total failed requests: $failed"

关键改动说明

  • 输入校验:过滤空行和非数字行,确保只有有效PMID被处理
  • 换行符处理:移除Windows格式的\r字符,避免URL拼接错误
  • 简化输出:用bash内置的echo替代多次冗余的ruby调用,提升效率
  • 速率控制:添加sleep 1控制请求频率,避免触发服务器限流
  • 错误处理:捕获curl请求失败的情况,统计失败数量,方便排查问题

使用说明

将PMID保存为纯文本文件(每行一个),然后运行脚本:

bash fetch_pubmed_abstracts.sh pmids.txt

内容的提问来源于stack exchange,提问作者Thulasi R

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 03:54:27