You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过NCBI解析HTML获取蛋白质序列?<pre>标签解析返回None问题

解决NCBI蛋白质序列爬取时无法找到
标签的问题 
  • 问题根源:你构造的URL请求后,NCBI返回的是纯文本格式的FASTA数据,而非HTML页面,所以用BeautifulSoup解析HTML结构自然找不到

    标签,返回None。

  • 直接获取序列的简单方法:不需要用BeautifulSoup,直接读取响应的文本内容即可,代码修改如下:

import requests

cds_id = "NP_001339842.1" 
fasta_url = f"https://www.ncbi.nlm.nih.gov/protein/{cds_id}/?report=fasta"

fasta_response = requests.get(fasta_url)
fasta_response.raise_for_status()         

# 直接打印FASTA格式的序列内容
print(fasta_response.text)
  • 更稳定的官方方法:使用NCBI官方的Entrez API(需要先安装biopython库:pip install biopython),这种方式不受页面结构变化影响,是NCBI推荐的合法访问方式:
from Bio import Entrez, SeqIO

# 必须填写你的邮箱,NCBI用来追踪API使用情况
Entrez.email = "your_email@example.com"  
cds_id = "NP_001339842.1"

# 获取FASTA格式的序列记录
handle = Entrez.efetch(db="protein", id=cds_id, rettype="fasta", retmode="text")
record = SeqIO.read(handle, "fasta")
handle.close()

# 输出序列相关信息
print(f"序列ID: {record.id}")
print(f"序列描述: {record.description}")
print(f"蛋白质序列: {record.seq}")

内容的提问来源于stack exchange,提问作者심지운

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 15:55:14