You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BioPython调用KEGG REST接口遇HTTP 403 Forbidden错误的解决方法

解决KEGG API请求403 Forbidden的问题

问题描述

尝试使用BioPython的Bio.KEGG.REST模块,通过KEGG化合物CID(如C0001代表水、C00123代表亮氨酸)查询化合物名称与分子式,代码如下:

from Bio.KEGG import REST
from Bio.KEGG import Compound


def cpd_decoder(cid): #gets the compound name and formula from KEGG
    if "C" in cid:
        cid="cpd:"+cid
        kegg_entry=REST.kegg_get(cid)
        for record in Compound.parse(kegg_entry):
            cid_name=record.name[0]
            cid_formula=record.formula 
            return cid_name,cid_formula

cid="C00123" #example CID; this one's for leucine
if cpd_decoder(cid) !=None:
    compound,formula=cpd_decoder(cid)

处理大量CID列表时,频繁遇到403错误:

if cpd_decoder(cid) !=None:
  File "/media/tessa/Storage/Alien_Earths/Network_expansion/network expansion test 2.py", line 27, in cpd_decoder
    kegg_entry=REST.kegg_get(cid)
  File "/home/tessa/.local/lib/python3.10/site-packages/Bio/KEGG/REST.py", line 208, in kegg_get
    resp = _q("get", dbentries)
  File "/home/tessa/.local/lib/python3.10/site-packages/Bio/KEGG/REST.py", line 44, in _q
    resp = urlopen(URL % (args))
  File "/usr/lib/python3.10/urllib/request.py", line 216, in urlopen
    return opener.open(url, data, timeout)
  File "/usr/lib/python3.10/urllib/request.py", line 525, in open
    response = meth(req, response)
  File "/usr/lib/python3.10/urllib/request.py", line 634, in http_response
    response = self.parent.error(
  File "/usr/lib/python3.10/urllib/request.py", line 563, in error
    return self._call_chain(*args)
  File "/usr/lib/python3.10/urllib/request.py", line 496, in _call_chain
    result = func(*args)
  File "/usr/lib/python3.10/urllib/request.py", line 643, in http_error_default
    raise HTTPError(req.full_url, code, msg, hdrs, fp)
urllib.error.HTTPError: HTTP Error 403: Forbidden 

怀疑请求被KEGG识别为机器人拦截,以下是可行的解决方法:

解决方法

1. 添加模拟浏览器的请求头

KEGG服务器可能拦截默认Python urllib的请求头,修改请求头模拟浏览器访问:

from Bio.KEGG import REST
from Bio.KEGG import Compound
import urllib.request

# 自定义请求头,模拟浏览器
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'
}
# 修改REST模块的opener
opener = urllib.request.build_opener()
opener.addheaders = [(k, v) for k, v in headers.items()]
urllib.request.install_opener(opener)

def cpd_decoder(cid):
    if "C" in cid:
        cid = "cpd:" + cid
        kegg_entry = REST.kegg_get(cid)
        for record in Compound.parse(kegg_entry):
            return record.name[0], record.formula

2. 控制请求频率,添加延迟

KEGG官方建议请求频率不超过每秒1次,批量查询时添加延迟避免触发反爬:

from Bio.KEGG import REST
from Bio.KEGG import Compound
import time

def cpd_decoder(cid):
    if "C" in cid:
        cid = "cpd:" + cid
        # 请求前添加延迟
        time.sleep(2)
        kegg_entry = REST.kegg_get(cid)
        for record in Compound.parse(kegg_entry):
            return record.name[0], record.formula

# 批量处理示例
cid_list = ["C0001", "C00123", "C0002"]
results = {}
for cid in cid_list:
    res = cpd_decoder(cid)
    if res:
        results[cid] = res

3. 使用批量查询接口减少请求次数

KEGG支持一次请求多个化合物ID,用逗号分隔,大幅减少请求数量:

from Bio.KEGG import REST
from Bio.KEGG import Compound

def batch_cpd_decoder(cid_list):
    # 格式化ID列表为cpd:Cxxx,cpd:Cxxx格式
    formatted_ids = ",".join([f"cpd:{cid}" for cid in cid_list if "C" in cid])
    kegg_entries = REST.kegg_get(formatted_ids)
    results = {}
    for record in Compound.parse(kegg_entries):
        # 提取原始CID(去掉cpd:前缀)
        cid = record.id.split(":")[-1]
        results[cid] = (record.name[0], record.formula)
    return results

# 批量查询示例
cid_list = ["C0001", "C00123"]
results = batch_cpd_decoder(cid_list)

4. 缓存查询结果避免重复请求

将已查询的结果存储在本地,再次遇到相同CID时直接读取缓存,减少请求量:

from Bio.KEGG import REST
from Bio.KEGG import Compound
import json
import os

CACHE_FILE = "kegg_cache.json"

# 加载缓存
def load_cache():
    if os.path.exists(CACHE_FILE):
        with open(CACHE_FILE, "r") as f:
            return json.load(f)
    return {}

# 保存缓存
def save_cache(cache):
    with open(CACHE_FILE, "w") as f:
        json.dump(cache, f)

cache = load_cache()

def cpd_decoder(cid):
    if cid in cache:
        return cache[cid]
    if "C" in cid:
        cid_formatted = "cpd:" + cid
        kegg_entry = REST.kegg_get(cid_formatted)
        for record in Compound.parse(kegg_entry):
            result = (record.name[0], record.formula)
            cache[cid] = result
            save_cache(cache)
            return result

内容的提问来源于stack exchange,提问作者Tessa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 12:55:36