Python爬取NCBI特定病毒科病毒名称:BeautifulSoup优化需求
精准提取NCBI病毒科下所有病毒名称的通用方案
核心思路
NCBI Taxonomy页面的病毒分类树被封装在固定容器内,直接定位这个容器而非全局爬取,就能避开Entrez、PubMed等无关内容。无需依赖易失效的正则,通过页面结构定位即可实现通用适配。
实现步骤与代码
- 安装依赖
pip install requests beautifulsoup4
- 通用爬取脚本
import requests from bs4 import BeautifulSoup def get_viruses_in_family(taxonomy_url): # 模拟浏览器请求,避免被NCBI拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' } resp = requests.get(taxonomy_url, headers=headers) resp.raise_for_status() soup = BeautifulSoup(resp.text, 'html.parser') # 定位分类树核心容器,NCBI当前版本用id=ncbiTaxonTree存储完整子分类 taxon_container = soup.find('div', id='ncbiTaxonTree') if not taxon_container: raise ValueError("Failed to locate taxonomy tree on the page") virus_names = [] # 遍历所有分类条目(li.taxon是NCBI标记分类项的固定类) for taxon_item in taxon_container.find_all('li', class_='taxon'): name_link = taxon_item.find('a') if name_link: raw_name = name_link.get_text(strip=True) # 过滤导航类无关条目 if not any(kw in raw_name.lower() for kw in ['descendants', 'all taxa']): virus_names.append(raw_name) return virus_names # 示例:提取冠状病毒科下的所有病毒 if __name__ == '__main__': # 替换为目标病毒科的NCBI分类页面URL(可通过NCBI Taxonomy搜索科名获取) target_family_url = 'https://www.ncbi.nlm.nih.gov/Taxonomy/Browser/wwwtax.cgi?id=11118' virus_list = get_viruses_in_family(target_family_url) print(f"Viruses in target family:") for idx, virus in enumerate(virus_list, 1): print(f"{idx}. {virus}")
关键说明
- 范围锁定:仅处理
ncbiTaxonTree容器内的内容,直接排除页面侧边栏、底部的Entrez/PubMed链接等无关区域 - 通用性:只要传入对应病毒科的NCBI分类页面URL,即可自动提取该科下所有子分类的病毒名称
- 稳定性:依赖NCBI官方标记分类项的
li.taxon类,比正则更抗页面结构小幅度变更 - 反拦截:添加标准User-Agent请求头,避免被NCBI的反爬机制拒绝访问
内容的提问来源于stack exchange,提问作者Programmer_101
相关产品推荐
相关产品推荐

