Shell脚本批量获取PubMed摘要时curl报错问题求助
解决批量提取PubMed摘要时的curl格式错误问题
问题背景
使用Shell脚本从PubMed批量提取摘要时,处理50-60个PMID能正常运行,但处理更多ID时出现curl: (3) URL using bad/illegal format or missing URL错误。原脚本及PMID示例如下:
原脚本
#!/bin/bash filename=$1 count=0 while read -r line; do (echo -n 'PMID:'"$line") ruby -e 'print "\t"' ruby -e 'print "\n"' curl "http://togows.dbcls.jp/entry/ncbi-pubmed/$line/abstract" ruby -e 'print "\n"' count=$((count+1)) done < "$filename" #total no. of abstracts extracted# echo $count
期望输出格式
PMID:15531828 text text text text text text text text text text text text PMID:11791095 text text text text text text text text text text text text
问题原因
- 无效输入行:输入文件可能包含空行、非数字字符,或Windows格式的换行符(
\r),导致拼接后的URL包含无效字符 - 速率限制:短时间内大量请求触发服务器速率限制,返回非预期内容,引发curl解析错误
修改后的脚本
#!/bin/bash if [ $# -ne 1 ]; then echo "Usage: $0 <pmid-file>" exit 1 fi filename=$1 count=0 failed=0 while read -r line; do # 过滤空行和非数字PMID [[ -z "$line" || ! "$line" =~ ^[0-9]+$ ]] && continue # 移除Windows换行符\r line=${line%$'\r'} echo "PMID:$line" # 用bash内置命令替代冗余ruby调用 echo "" # 发起请求,添加延迟避免速率限制 if curl -s "http://togows.dbcls.jp/entry/ncbi-pubmed/$line/abstract"; then ((count++)) else echo "[ERROR] Failed to fetch abstract for PMID:$line" ((failed++)) fi # 添加空行分隔不同PMID的内容 echo -e "\n" # 延迟1秒,避免触发服务器限流 sleep 1 done < "$filename" echo "Total abstracts extracted: $count" echo "Total failed requests: $failed"
关键改动说明
- 输入校验:过滤空行和非数字行,确保只有有效PMID被处理
- 换行符处理:移除Windows格式的
\r字符,避免URL拼接错误 - 简化输出:用bash内置的
echo替代多次冗余的ruby调用,提升效率 - 速率控制:添加
sleep 1控制请求频率,避免触发服务器限流 - 错误处理:捕获curl请求失败的情况,统计失败数量,方便排查问题
使用说明
将PMID保存为纯文本文件(每行一个),然后运行脚本:
bash fetch_pubmed_abstracts.sh pmids.txt
内容的提问来源于stack exchange,提问作者Thulasi R
相关产品推荐
相关产品推荐

