如何通过NCBI解析HTML获取蛋白质序列?<pre>标签解析返回None问题
解决NCBI蛋白质序列爬取时无法找到
标签的问题
问题根源:你构造的URL请求后,NCBI返回的是纯文本格式的FASTA数据,而非HTML页面,所以用BeautifulSoup解析HTML结构自然找不到
标签,返回None。
直接获取序列的简单方法:不需要用BeautifulSoup,直接读取响应的文本内容即可,代码修改如下:
import requests cds_id = "NP_001339842.1" fasta_url = f"https://www.ncbi.nlm.nih.gov/protein/{cds_id}/?report=fasta" fasta_response = requests.get(fasta_url) fasta_response.raise_for_status() # 直接打印FASTA格式的序列内容 print(fasta_response.text)
- 更稳定的官方方法:使用NCBI官方的Entrez API(需要先安装biopython库:
pip install biopython),这种方式不受页面结构变化影响,是NCBI推荐的合法访问方式:
from Bio import Entrez, SeqIO # 必须填写你的邮箱,NCBI用来追踪API使用情况 Entrez.email = "your_email@example.com" cds_id = "NP_001339842.1" # 获取FASTA格式的序列记录 handle = Entrez.efetch(db="protein", id=cds_id, rettype="fasta", retmode="text") record = SeqIO.read(handle, "fasta") handle.close() # 输出序列相关信息 print(f"序列ID: {record.id}") print(f"序列描述: {record.description}") print(f"蛋白质序列: {record.seq}")
内容的提问来源于stack exchange,提问作者심지운
相关产品推荐
相关产品推荐

