You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否爬取ebi.ac.uk/interpro网站?Python爬取蛋白质数据报错求助

解决InterPro蛋白质表格爬取的AttributeError问题

错误原因分析

你遇到的AttributeError: 'NoneType' object has no attribute 'find_all'本质是:soup.find("table", class_ = 'xxx')返回了None——也就是代码没找到指定class的表格,大概率是以下三种情况:

  • 页面表格是动态加载的:InterPro的大量蛋白质数据通常通过AJAX异步渲染,直接用requests.get只能拿到静态框架HTML,看不到实际表格内容
  • Class名称错误:审查元素时可能误选了父容器的class,或者网站的表格class是动态生成的临时值
  • 请求头缺失:网站会验证请求的User-Agent等标识,无标识请求可能返回不完整内容

解决方案1:修正静态爬取逻辑(适用于表格静态加载场景)

先验证请求有效性,再正确定位表格:

import requests
from bs4 import BeautifulSoup

# 替换为你的目标条目URL
url = "https://ebi.ac.uk/interpro/[你的具体条目路径]"

# 添加模拟浏览器的请求头,避免被拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

response = requests.get(url, headers=headers)
# 先确认请求是否成功
if response.status_code != 200:
    print(f"请求失败,状态码:{response.status_code}")
    exit()

soup = BeautifulSoup(response.text, "html.parser")

# 先打印页面所有表格的class,确认目标表格的真实class
all_tables = soup.find_all("table")
for idx, tbl in enumerate(all_tables):
    print(f"表格{idx}的class:{tbl.get('class')}")

# 替换为上面打印出的真实表格class
table = soup.find("table", class_="[替换为实际class]")

if table is None:
    print("未找到目标表格,请检查class是否正确")
else:
    table_data = []
    # 跳过表头,从第二行开始提取数据
    for row in table.find_all('tr')[1:]:
        cells = row.find_all("td")
        row_data = [cell.get_text(strip=True) for cell in cells]
        table_data.append(row_data)
    # 打印前5条数据验证
    print(table_data[:5])

解决方案2:处理动态加载数据(更适配InterPro场景)

如果表格是动态渲染的,用selenium模拟浏览器加载页面:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
import time

url = "https://ebi.ac.uk/interpro/[你的具体条目路径]"

# 配置无头浏览器,不弹出窗口
chrome_options = Options()
chrome_options.add_argument("--headless=new")
chrome_options.add_argument("--disable-gpu")
chrome_options.add_argument("User-Agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")

driver = webdriver.Chrome(options=chrome_options)
driver.get(url)

# 等待页面加载完成(根据实际加载速度调整时间)
time.sleep(5)

# 打印所有表格的class,确认目标表格
all_tables = driver.find_elements(By.TAG_NAME, "table")
for idx, tbl in enumerate(all_tables):
    print(f"表格{idx}的class:{tbl.get_attribute('class')}")

# 替换为真实表格class
table = driver.find_element(By.CLASS_NAME, "[替换为实际class]")
rows = table.find_elements(By.TAG_NAME, "tr")

table_data = []
for row in rows[1:]:
    cells = row.find_elements(By.TAG_NAME, "td")
    row_data = [cell.text.strip() for cell in cells]
    table_data.append(row_data)

# 打印验证
print(table_data[:5])
driver.quit()

额外提示

  • 爬取大量数据时,记得添加请求间隔(比如time.sleep(1)),避免触发网站反爬机制
  • InterPro提供官方REST API,直接调用API获取数据比页面爬取更稳定、高效,可优先查看网站的API文档实现

内容的提问来源于stack exchange,提问作者stde

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 18:40:27