You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BS4爬虫在Google Colab运行报错,求完整数据抓取解决方案

问题解决与完整爬虫实现方案

一、解决现有报错

1. TypeError: 'NoneType' object is not subscriptable 报错处理

这个报错是因为你通过find()或find_all()获取id为paginationPagesNum的input元素时返回了None,随后尝试取['value']导致报错。核心原因是该元素由JavaScript动态生成,用requests直接请求页面无法获取到渲染后的元素。

解决方法:

  • 方法一:改用Selenium模拟浏览器加载页面,等待元素渲染完成后提取:
    from selenium import webdriver
    from selenium.webdriver.common.by import By
    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.support import expected_conditions as EC
    
    # Colab需先配置ChromeDriver,可使用!pip install colab-selenium初始化
    driver = webdriver.Chrome()
    driver.get("https://s3platform.jrc.ec.europa.eu/digital-innovation-hubs-tool")
    # 等待分页元素加载完成
    page_num_input = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.ID, "paginationPagesNum"))
    )
    total_pages = int(page_num_input.get_attribute("value"))
    
  • 方法二:绕过分页组件,直接计算总页数。比如页面默认每页显示20条,总条目数700,总页数为700//20 + (1 if 700%20 else 0),直接遍历1到35页即可。

2. NameError 报错处理

这些错误是因为缺少必要的库导入和对象初始化:

  • 未定义pd:需要导入pandas库
  • 未定义BeautifulSoup:需要导入bs4库
  • 未定义df:需要先初始化DataFrame对象

在脚本开头添加以下代码:

import requests
from bs4 import BeautifulSoup
import pandas as pd

二、完整数据抓取方案(获取全部700条详情页数据)

整体流程

  1. 遍历列表页,收集所有枢纽的详情页链接
  2. 逐个访问详情页,提取Hub Information、Description、Contact Data等所有字段
  3. 整理数据为DataFrame,导出为CSV文件

代码实现(适配Google Colab)

import requests
from bs4 import BeautifulSoup
import pandas as pd
from time import sleep

# 基础配置
base_url = "https://s3platform.jrc.ec.europa.eu"
list_url_template = "https://s3platform.jrc.ec.europa.eu/digital-innovation-hubs-tool?page={}"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

# 存储所有枢纽数据的列表
hubs_full_data = []

# 第一步:抓取所有详情页链接
print("开始抓取列表页链接...")
for page in range(1, 36):
    response = requests.get(list_url_template.format(page), headers=headers)
    soup = BeautifulSoup(response.text, "html.parser")
    
    # 提取当前页所有枢纽的详情链接
    hub_links = soup.select("div.hub-item a.btn-primary")
    for link in hub_links:
        hubs_full_data.append({
            "detail_url": base_url + link["href"]
        })
    
    sleep(1)  # 控制请求频率,避免被拦截

# 第二步:抓取每个详情页的完整数据
print("开始抓取详情页数据...")
for idx, hub in enumerate(hubs_full_data):
    print(f"处理第 {idx+1}/{len(hubs_full_data)} 条数据")
    response = requests.get(hub["detail_url"], headers=headers)
    soup = BeautifulSoup(response.text, "html.parser")
    
    # 提取Hub Information字段
    info_sections = soup.select("div.hub-info-section")
    for section in info_sections:
        label = section.select_one("div.hub-info-label").text.strip()
        value = section.select_one("div.hub-info-value").text.strip()
        hub[label] = value
    
    # 提取Description
    description_elem = soup.select_one("div.hub-description")
    hub["Description"] = description_elem.text.strip() if description_elem else ""
    
    # 提取Contact Data字段
    contact_sections = soup.select("div.contact-data-section")
    for section in contact_sections:
        label = section.select_one("div.contact-data-label").text.strip()
        value = section.select_one("div.contact-data-value").text.strip()
        hub[label] = value
    
    sleep(1)

# 第三步:导出为CSV
df = pd.DataFrame(hubs_full_data)
df.to_csv("eu_digital_hubs_full_data.csv", index=False, encoding="utf-8-sig")
print("数据导出完成,文件为 eu_digital_hubs_full_data.csv")

注意事项

  • 若遇到请求被拦截,可适当延长sleep()的时间,或添加更多请求头字段模拟真实浏览器
  • 若页面结构发生变化,需根据实际HTML调整select()或select_one()的选择器参数

内容的提问来源于stack exchange,提问作者zero

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 12:23:18