You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python将基因名转换为Uniprot ID?解决mygene库缺失问题

将基因名转换为Uniprot ID的Python脚本方案

方法1:直接调用Uniprot REST API(推荐批量处理)

Uniprot的REST接口支持精准批量查询,能有效减少基因匹配丢失的情况。以下是完整脚本,支持从文本文件读取基因列表,输出匹配到的Uniprot ID及对应信息:

import requests
import pandas as pd

def gene_to_uniprot(gene_list, organism="human"):
    # 构造查询参数,限定物种(可修改为mouse、rat等)
    params = {
        "query": " OR ".join([f"gene:{gene}" for gene in gene_list]),
        "format": "tsv",
        "columns": "id,gene_names,organism_name",
        "limit": len(gene_list) * 2  # 预留冗余避免漏匹配
    }
    url = "https://rest.uniprot.org/uniprotkb/search"
    
    response = requests.get(url, params=params)
    response.raise_for_status()  # 捕获请求错误
    
    # 解析TSV格式的返回数据
    lines = response.text.strip().split("\n")
    headers = lines[0].split("\t")
    data = [line.split("\t") for line in lines[1:]]
    
    df = pd.DataFrame(data, columns=headers)
    # 去重,保留每个基因的首个匹配结果(可按需调整逻辑)
    df = df.drop_duplicates(subset="Gene Names", keep="first")
    return df

# 使用示例
if __name__ == "__main__":
    # 从文本文件读取基因名(每行一个基因)
    with open("gene_list.txt", "r") as f:
        gene_names = [line.strip() for line in f if line.strip()]
    
    # 执行转换
    result_df = gene_to_uniprot(gene_names, organism="human")
    
    # 保存结果到CSV
    result_df.to_csv("gene_to_uniprot.csv", index=False)
    print("转换完成,结果已保存到gene_to_uniprot.csv")

方法2:使用Biopython的ExPASy模块(适合单/少量基因)

如果需要灵活处理单个基因查询,Biopython的ExPASy工具可直接对接Uniprot服务:

from Bio import ExPASy
from Bio.SwissProt import read

def get_uniprot_id(gene_name, organism="Homo sapiens"):
    try:
        handle = ExPASy.get_sprot_raw(f"gene:{gene_name} AND organism:{organism}")
        record = read(handle)
        return record.accessions[0]  # 返回首个匹配的Uniprot ID
    except Exception as e:
        print(f"未找到基因 {gene_name} 的匹配结果: {e}")
        return None

# 使用示例
gene_list = ["TP53", "BRCA1", "INVALID_GENE"]
for gene in gene_list:
    uniprot_id = get_uniprot_id(gene)
    if uniprot_id:
        print(f"{gene} -> {uniprot_id}")

关键注意事项

  • 物种限定:这是基因匹配丢失的核心原因之一!mygene若未指定物种,可能返回跨物种结果或漏匹配,上述两种方法都明确限定了物种,可根据研究对象修改。
  • 基因名格式:确保输入的是标准HGNC符号(如人类基因用大写字母),避免使用别名或错误拼写。
  • 多匹配处理:部分基因对应多个Uniprot ID(如不同转录本),脚本默认保留首个匹配,可按需调整去重逻辑或保留所有结果。

内容的提问来源于stack exchange,提问作者anita

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 00:18:24