You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

OpenSea交易数据爬取异常:滚动后仅获19条innerHTML值求解决

解决OpenSea Selenium爬取交易数据仅获取19条的问题

我之前也踩过动态加载页面爬取的坑,你的代码核心问题出在滚动策略和等待逻辑设置不合理,导致页面还没加载完所有交易数据就停止滚动了,再加上元素收集时机单一,自然只能拿到少量数据。下面是具体的分析和改进方案:

问题根源拆解

  1. 滚动过于激进:window.scrollTo(0, 100*document.body.scrollHeight)直接跳到了远超当前页面的位置,很多网站(包括OpenSea)是检测到滚动接近底部时才触发新内容加载,这种跳转会跳过加载触发点。
  2. 超时时间太短:max_run_time = 1只给了1秒等待时间,页面加载新交易数据需要网络请求和渲染时间,尤其是网络波动时,还没加载完就停止滚动了。
  3. 元素收集时机单一:只在滚动结束后一次性查找元素,可能有些新加载的元素还没完全渲染,或者滚动过程中加载的内容没被捕获。

改进后的完整代码

import time
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

options = Options()
options.add_experimental_option("detach", True)
options.add_argument("start-maximized")
options.add_experimental_option('excludeSwitches', ['enable-logging'])
# 模拟真实浏览器UA,降低反爬识别概率
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
# 禁用自动化特征检测
options.add_argument("--disable-blink-features=AutomationControlled")

driver = webdriver.Chrome(options=options)

target_url = "https://opensea.io/activity?search[collections][0]=fvckrender&search[collections][1]=artifex-fvckrender&search[collections][2]=fvck-limited&search[collections][3]=unidentified-contract-kg9mf80eue&search[eventTypes][0]=AUCTION_SUCCESSFUL"
driver.get(target_url)
# 显式等待核心元素加载完成,避免过早操作
WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.XPATH, "//div[@class='sc-fe5f9c83-0 mGAUR Price--fiat-amount']"))
)

# 优化滚动加载逻辑
pre_scroll_height = driver.execute_script('return document.body.scrollHeight;')
run_time, max_run_time = 0, 12  # 延长超时时间到12秒,给足加载缓冲
last_item_count = 0

while True:
    iteration_start = time.time()
    
    # 改为滚动到当前页面底部,触发OpenSea的加载逻辑
    driver.execute_script('window.scrollTo(0, document.body.scrollHeight);')
    
    # 等待新内容渲染,可根据网络情况调整时长
    time.sleep(2.5)
    # 等待页面加载状态完成
    WebDriverWait(driver, 5).until(
        lambda d: d.execute_script('return document.readyState') == 'complete'
    )
    
    post_scroll_height = driver.execute_script('return document.body.scrollHeight;')
    # 实时检查已加载的交易数量
    current_items = driver.find_elements(By.XPATH, "//div[@class='sc-fe5f9c83-0 mGAUR Price--fiat-amount']")
    current_item_count = len(current_items)
    
    scrolled = post_scroll_height != pre_scroll_height
    new_items_loaded = current_item_count > last_item_count
    timed_out = run_time >= max_run_time
    
    if scrolled or new_items_loaded:
        # 有新内容加载,重置超时计时器
        run_time = 0
        pre_scroll_height = post_scroll_height
        last_item_count = current_item_count
        print(f"已加载 {current_item_count} 条交易数据...")
    elif not scrolled and not new_items_loaded and not timed_out:
        # 无新内容时累计等待时间
        run_time += time.time() - iteration_start
    else:
        # 超时且无新内容,停止滚动
        print("已加载所有可获取的交易数据,停止滚动")
        break

# 最终收集所有数据
price_of_pieces_fiat = driver.find_elements(By.XPATH, "//div[@class='sc-fe5f9c83-0 mGAUR Price--fiat-amount']")
print(f"\n总共获取到 {len(price_of_pieces_fiat)} 条数据:")
for idx, price_of_piece in enumerate(price_of_pieces_fiat, 1):
    print(f"{idx}. {price_of_piece.get_attribute('innerHTML')}")

关键改进点说明

  • 滚动策略优化:改成逐步滚动到当前页面底部,完美匹配OpenSea的滚动加载触发逻辑,确保每次滚动都能触发新内容加载。
  • 等待逻辑升级:延长超时时间,增加滚动后的固定等待和页面状态检查,确保新元素完全渲染后再进行判断。
  • 数据跟踪机制:通过对比前后交易元素的数量,确认是否有新内容加载,避免因页面高度不变但元素仍在加载的情况导致提前停止。
  • 反爬适配:添加真实用户代理和禁用自动化检测,降低被OpenSea反爬机制识别的概率。

额外小贴士

  • 如果还是出现数据不全,可以适当延长time.sleep的时长,或者增加滚动循环的最大次数。
  • 不要过于频繁地执行爬取操作,避免触发OpenSea的IP封禁机制。
  • 可以尝试用ActionChains模拟鼠标滚轮滚动,更贴近真实用户的操作行为。

内容的提问来源于stack exchange,提问作者omno

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 10:22:54