如何使用Selenium爬取指定页面的Floor price history图表数据
基于Selenium爬取目标站点Floor price history图表的可落地方案
这个站点的底价历史图表是前端Canvas渲染的,不需要做OCR识别,也不用模拟鼠标悬停逐点读数值,站点本身是Next.js架构,所有图表渲染用的结构化数据都提前随页面下发,直接提取结构化数据准确率100%,实现成本极低。
前置依赖
- 安装所需Python包:
pip install selenium webdriver-manager - 不需要手动下载匹配版本的ChromeDriver,webdriver-manager会自动适配本地Chrome版本,避免版本不兼容报错
实现逻辑
- 初始化Chrome实例时加基础反爬参数,移除Selenium自动化特征,避免被站点拦截
必须加
--disable-blink-features=AutomationControlled启动参数,同时覆盖navigator.webdriver属性,否则页面会触发反爬校验,返回空数据
- 访问目标页面后,用显式等待等待页面加载完成,不要写固定sleep硬等
- 直接提取页面中id为
__NEXT_DATA__的script标签内容,这是Next.js框架默认挂载的全量页面数据容器,所有Floor price history的时间戳、对应底价、交易量数据全在这个JSON里,不需要解析图表元素 - 如果后续站点改版移除了
__NEXT_DATA__节点,可以切换为CDP监听网络响应的方式,直接拦截加载历史价格的接口返回值,稳定性更高 - 如果需求是获取图表本身的图片而非原始数值,直接定位图表对应的Canvas元素,调用元素自带的
screenshot()方法即可导出高清图表截图
核心可运行代码
from selenium import webdriver from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import json import time # 初始化浏览器配置 options = webdriver.ChromeOptions() options.add_argument("--disable-blink-features=AutomationControlled") options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('useAutomationExtension', False) driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options) driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})") try: driver.get("https://nftpricefloor.com/it/bored-ape-yacht-club") # 等待数据容器加载完成 wait = WebDriverWait(driver, 15) data_script = wait.until(EC.presence_of_element_located((By.ID, "__NEXT_DATA__"))) # 解析结构化数据 page_data = json.loads(data_script.get_attribute("textContent")) # 顺着层级找底价历史数据,站点字段调整时打印page_data顺着props路径找即可 floor_history = page_data["props"]["pageProps"]["collection"]["floorHistory"] # 验证数据 print("底价历史前10条数据:", floor_history[:10]) # 如需导出图表截图,打开下方注释即可 # chart_canvas = driver.find_element(By.CSS_SELECTOR, "canvas[role='img']") # chart_canvas.screenshot("floor_price_history.png") finally: time.sleep(2) driver.quit()
常见踩坑点
- 不要逐点模拟鼠标悬停读图表tooltip:这种方式速度极慢,还容易因为页面元素偏移读错数值,完全没必要
- 如果提取到的历史数据长度不对,适当延长显式等待时间,等前端把全量历史数据拉取完成再读取标签内容
- 解析JSON时如果找不到对应字段,先把
page_data整体打印出来,顺着props->pageProps->collection路径找,站点字段调整的话顺着层级找1分钟就能定位到
内容的提问来源于stack exchange,提问作者GKM__
相关产品推荐
相关产品推荐

