使用BioPython调用KEGG REST接口遇HTTP 403 Forbidden错误的解决方法
解决KEGG API请求403 Forbidden的问题
问题描述
尝试使用BioPython的Bio.KEGG.REST模块,通过KEGG化合物CID(如C0001代表水、C00123代表亮氨酸)查询化合物名称与分子式,代码如下:
from Bio.KEGG import REST from Bio.KEGG import Compound def cpd_decoder(cid): #gets the compound name and formula from KEGG if "C" in cid: cid="cpd:"+cid kegg_entry=REST.kegg_get(cid) for record in Compound.parse(kegg_entry): cid_name=record.name[0] cid_formula=record.formula return cid_name,cid_formula cid="C00123" #example CID; this one's for leucine if cpd_decoder(cid) !=None: compound,formula=cpd_decoder(cid)
处理大量CID列表时,频繁遇到403错误:
if cpd_decoder(cid) !=None: File "/media/tessa/Storage/Alien_Earths/Network_expansion/network expansion test 2.py", line 27, in cpd_decoder kegg_entry=REST.kegg_get(cid) File "/home/tessa/.local/lib/python3.10/site-packages/Bio/KEGG/REST.py", line 208, in kegg_get resp = _q("get", dbentries) File "/home/tessa/.local/lib/python3.10/site-packages/Bio/KEGG/REST.py", line 44, in _q resp = urlopen(URL % (args)) File "/usr/lib/python3.10/urllib/request.py", line 216, in urlopen return opener.open(url, data, timeout) File "/usr/lib/python3.10/urllib/request.py", line 525, in open response = meth(req, response) File "/usr/lib/python3.10/urllib/request.py", line 634, in http_response response = self.parent.error( File "/usr/lib/python3.10/urllib/request.py", line 563, in error return self._call_chain(*args) File "/usr/lib/python3.10/urllib/request.py", line 496, in _call_chain result = func(*args) File "/usr/lib/python3.10/urllib/request.py", line 643, in http_error_default raise HTTPError(req.full_url, code, msg, hdrs, fp) urllib.error.HTTPError: HTTP Error 403: Forbidden
怀疑请求被KEGG识别为机器人拦截,以下是可行的解决方法:
解决方法
1. 添加模拟浏览器的请求头
KEGG服务器可能拦截默认Python urllib的请求头,修改请求头模拟浏览器访问:
from Bio.KEGG import REST from Bio.KEGG import Compound import urllib.request # 自定义请求头,模拟浏览器 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36' } # 修改REST模块的opener opener = urllib.request.build_opener() opener.addheaders = [(k, v) for k, v in headers.items()] urllib.request.install_opener(opener) def cpd_decoder(cid): if "C" in cid: cid = "cpd:" + cid kegg_entry = REST.kegg_get(cid) for record in Compound.parse(kegg_entry): return record.name[0], record.formula
2. 控制请求频率,添加延迟
KEGG官方建议请求频率不超过每秒1次,批量查询时添加延迟避免触发反爬:
from Bio.KEGG import REST from Bio.KEGG import Compound import time def cpd_decoder(cid): if "C" in cid: cid = "cpd:" + cid # 请求前添加延迟 time.sleep(2) kegg_entry = REST.kegg_get(cid) for record in Compound.parse(kegg_entry): return record.name[0], record.formula # 批量处理示例 cid_list = ["C0001", "C00123", "C0002"] results = {} for cid in cid_list: res = cpd_decoder(cid) if res: results[cid] = res
3. 使用批量查询接口减少请求次数
KEGG支持一次请求多个化合物ID,用逗号分隔,大幅减少请求数量:
from Bio.KEGG import REST from Bio.KEGG import Compound def batch_cpd_decoder(cid_list): # 格式化ID列表为cpd:Cxxx,cpd:Cxxx格式 formatted_ids = ",".join([f"cpd:{cid}" for cid in cid_list if "C" in cid]) kegg_entries = REST.kegg_get(formatted_ids) results = {} for record in Compound.parse(kegg_entries): # 提取原始CID(去掉cpd:前缀) cid = record.id.split(":")[-1] results[cid] = (record.name[0], record.formula) return results # 批量查询示例 cid_list = ["C0001", "C00123"] results = batch_cpd_decoder(cid_list)
4. 缓存查询结果避免重复请求
将已查询的结果存储在本地,再次遇到相同CID时直接读取缓存,减少请求量:
from Bio.KEGG import REST from Bio.KEGG import Compound import json import os CACHE_FILE = "kegg_cache.json" # 加载缓存 def load_cache(): if os.path.exists(CACHE_FILE): with open(CACHE_FILE, "r") as f: return json.load(f) return {} # 保存缓存 def save_cache(cache): with open(CACHE_FILE, "w") as f: json.dump(cache, f) cache = load_cache() def cpd_decoder(cid): if cid in cache: return cache[cid] if "C" in cid: cid_formatted = "cpd:" + cid kegg_entry = REST.kegg_get(cid_formatted) for record in Compound.parse(kegg_entry): result = (record.name[0], record.formula) cache[cid] = result save_cache(cache) return result
内容的提问来源于stack exchange,提问作者Tessa
相关产品推荐
相关产品推荐

