能否爬取ebi.ac.uk/interpro网站?Python爬取蛋白质数据报错求助
解决InterPro蛋白质表格爬取的AttributeError问题
错误原因分析
你遇到的AttributeError: 'NoneType' object has no attribute 'find_all'本质是:soup.find("table", class_ = 'xxx')返回了None——也就是代码没找到指定class的表格,大概率是以下三种情况:
- 页面表格是动态加载的:InterPro的大量蛋白质数据通常通过AJAX异步渲染,直接用
requests.get只能拿到静态框架HTML,看不到实际表格内容 - Class名称错误:审查元素时可能误选了父容器的class,或者网站的表格class是动态生成的临时值
- 请求头缺失:网站会验证请求的
User-Agent等标识,无标识请求可能返回不完整内容
解决方案1:修正静态爬取逻辑(适用于表格静态加载场景)
先验证请求有效性,再正确定位表格:
import requests from bs4 import BeautifulSoup # 替换为你的目标条目URL url = "https://ebi.ac.uk/interpro/[你的具体条目路径]" # 添加模拟浏览器的请求头,避免被拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(url, headers=headers) # 先确认请求是否成功 if response.status_code != 200: print(f"请求失败,状态码:{response.status_code}") exit() soup = BeautifulSoup(response.text, "html.parser") # 先打印页面所有表格的class,确认目标表格的真实class all_tables = soup.find_all("table") for idx, tbl in enumerate(all_tables): print(f"表格{idx}的class:{tbl.get('class')}") # 替换为上面打印出的真实表格class table = soup.find("table", class_="[替换为实际class]") if table is None: print("未找到目标表格,请检查class是否正确") else: table_data = [] # 跳过表头,从第二行开始提取数据 for row in table.find_all('tr')[1:]: cells = row.find_all("td") row_data = [cell.get_text(strip=True) for cell in cells] table_data.append(row_data) # 打印前5条数据验证 print(table_data[:5])
解决方案2:处理动态加载数据(更适配InterPro场景)
如果表格是动态渲染的,用selenium模拟浏览器加载页面:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options import time url = "https://ebi.ac.uk/interpro/[你的具体条目路径]" # 配置无头浏览器,不弹出窗口 chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("--disable-gpu") chrome_options.add_argument("User-Agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") driver = webdriver.Chrome(options=chrome_options) driver.get(url) # 等待页面加载完成(根据实际加载速度调整时间) time.sleep(5) # 打印所有表格的class,确认目标表格 all_tables = driver.find_elements(By.TAG_NAME, "table") for idx, tbl in enumerate(all_tables): print(f"表格{idx}的class:{tbl.get_attribute('class')}") # 替换为真实表格class table = driver.find_element(By.CLASS_NAME, "[替换为实际class]") rows = table.find_elements(By.TAG_NAME, "tr") table_data = [] for row in rows[1:]: cells = row.find_elements(By.TAG_NAME, "td") row_data = [cell.text.strip() for cell in cells] table_data.append(row_data) # 打印验证 print(table_data[:5]) driver.quit()
额外提示
- 爬取大量数据时,记得添加请求间隔(比如
time.sleep(1)),避免触发网站反爬机制 - InterPro提供官方REST API,直接调用API获取数据比页面爬取更稳定、高效,可优先查看网站的API文档实现
内容的提问来源于stack exchange,提问作者stde
相关产品推荐
相关产品推荐

