BS4爬虫在Google Colab运行报错,求完整数据抓取解决方案
问题解决与完整爬虫实现方案
一、解决现有报错
1. TypeError: 'NoneType' object is not subscriptable 报错处理
这个报错是因为你通过find()或find_all()获取id为paginationPagesNum的input元素时返回了None,随后尝试取['value']导致报错。核心原因是该元素由JavaScript动态生成,用requests直接请求页面无法获取到渲染后的元素。
解决方法:
- 方法一:改用Selenium模拟浏览器加载页面,等待元素渲染完成后提取:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Colab需先配置ChromeDriver,可使用!pip install colab-selenium初始化 driver = webdriver.Chrome() driver.get("https://s3platform.jrc.ec.europa.eu/digital-innovation-hubs-tool") # 等待分页元素加载完成 page_num_input = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, "paginationPagesNum")) ) total_pages = int(page_num_input.get_attribute("value")) - 方法二:绕过分页组件,直接计算总页数。比如页面默认每页显示20条,总条目数700,总页数为
700//20 + (1 if 700%20 else 0),直接遍历1到35页即可。
2. NameError 报错处理
这些错误是因为缺少必要的库导入和对象初始化:
- 未定义
pd:需要导入pandas库 - 未定义
BeautifulSoup:需要导入bs4库 - 未定义
df:需要先初始化DataFrame对象
在脚本开头添加以下代码:
import requests from bs4 import BeautifulSoup import pandas as pd
二、完整数据抓取方案(获取全部700条详情页数据)
整体流程
- 遍历列表页,收集所有枢纽的详情页链接
- 逐个访问详情页,提取
Hub Information、Description、Contact Data等所有字段 - 整理数据为DataFrame,导出为CSV文件
代码实现(适配Google Colab)
import requests from bs4 import BeautifulSoup import pandas as pd from time import sleep # 基础配置 base_url = "https://s3platform.jrc.ec.europa.eu" list_url_template = "https://s3platform.jrc.ec.europa.eu/digital-innovation-hubs-tool?page={}" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } # 存储所有枢纽数据的列表 hubs_full_data = [] # 第一步:抓取所有详情页链接 print("开始抓取列表页链接...") for page in range(1, 36): response = requests.get(list_url_template.format(page), headers=headers) soup = BeautifulSoup(response.text, "html.parser") # 提取当前页所有枢纽的详情链接 hub_links = soup.select("div.hub-item a.btn-primary") for link in hub_links: hubs_full_data.append({ "detail_url": base_url + link["href"] }) sleep(1) # 控制请求频率,避免被拦截 # 第二步:抓取每个详情页的完整数据 print("开始抓取详情页数据...") for idx, hub in enumerate(hubs_full_data): print(f"处理第 {idx+1}/{len(hubs_full_data)} 条数据") response = requests.get(hub["detail_url"], headers=headers) soup = BeautifulSoup(response.text, "html.parser") # 提取Hub Information字段 info_sections = soup.select("div.hub-info-section") for section in info_sections: label = section.select_one("div.hub-info-label").text.strip() value = section.select_one("div.hub-info-value").text.strip() hub[label] = value # 提取Description description_elem = soup.select_one("div.hub-description") hub["Description"] = description_elem.text.strip() if description_elem else "" # 提取Contact Data字段 contact_sections = soup.select("div.contact-data-section") for section in contact_sections: label = section.select_one("div.contact-data-label").text.strip() value = section.select_one("div.contact-data-value").text.strip() hub[label] = value sleep(1) # 第三步:导出为CSV df = pd.DataFrame(hubs_full_data) df.to_csv("eu_digital_hubs_full_data.csv", index=False, encoding="utf-8-sig") print("数据导出完成,文件为 eu_digital_hubs_full_data.csv")
注意事项
- 若遇到请求被拦截,可适当延长
sleep()的时间,或添加更多请求头字段模拟真实浏览器 - 若页面结构发生变化,需根据实际HTML调整
select()或select_one()的选择器参数
内容的提问来源于stack exchange,提问作者zero
相关产品推荐
相关产品推荐

