You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium爬取懒加载页面前如何滚动至底部获取全部数据

代码问题定位

你原有代码滚动失败的核心原因有3个:

  1. 高度获取方式不匹配目标站点,document.documentElement.scrollHeight无法拿到真实的页面总高度
  2. 滚动逻辑刚好卡在懒加载触发阈值外,没法触发后端加载新内容
  3. 爬取逻辑错误嵌套在滚动循环中,不仅会重复写入数据,还可能干扰滚动加载的判断
修正后可直接运行的代码
from selenium import webdriver
from selenium.webdriver.common.by import By
import time
import csv

# 初始化Chrome浏览器,Selenium 4.6+无需手动指定chromedriver路径
driver = webdriver.Chrome()
# 如果你是旧版本Selenium,就保留你原来的路径写法:
# driver = webdriver.Chrome("C:.............chromedriver_win32/chromedriver.exe")

# 打开目标页面
driver.get('https://icodrops.com/category/ended-ico/')
# 给页面初始加载留时间
time.sleep(3)

# 滚动到底部加载全部内容逻辑
last_height = 0
# 连续相同高度计数,避免网络波动误判
same_height_count = 0
while same_height_count < 3:
    # 取两种高度的最大值,适配不同页面规则
    current_height = max(
        driver.execute_script("return document.body.scrollHeight"),
        driver.execute_script("return document.documentElement.scrollHeight")
    )
    # 滚动到比当前高度多100像素的位置,触发懒加载
    driver.execute_script(f"window.scrollTo(0, {current_height + 100});")
    time.sleep(5) # 可根据你的网络情况调整等待时长
    # 计算新高度
    new_height = max(
        driver.execute_script("return document.body.scrollHeight"),
        driver.execute_script("return document.documentElement.scrollHeight")
    )
    print(f"当前页面高度:{new_height}")
    if new_height == current_height:
        same_height_count += 1
    else:
        same_height_count = 0
        last_height = new_height

print("已滚动至页面最底部,开始采集数据")

# 数据写入逻辑
csv_file = open('icodrops_ended_icos.csv', 'w', encoding='utf-8', newline='')
writer = csv.writer(csv_file)
writer.writerow(['Project_Name', 'Interest', 'Category', 'Received', 'Goal', 'End_Date', 'Ticker'])

try:
    # 提取所有行数据
    rows = driver.find_elements(By.XPATH, '//div[@class="col-md-12 col-12 a_ico"]') 
    print(f"共获取到{len(rows)}条项目数据")
    for row in rows:
        project_name = row.find_element(By.XPATH, './/div[@class="ico-row"]/div[2]/h3/a').text
        interest = row.find_element(By.XPATH, './/div[@class="interest"]').text
        category = row.find_element(By.XPATH, './/div[@class="categ_type"]').text
        received = row.find_element(By.XPATH, './/div[@id="new_column_categ_invisted"]/span').text
        goal = row.find_element(By.XPATH, './/div[@id="categ_desctop"]').text
        end_date = row.find_element(By.XPATH, './/div[@class="date"]').text
        ticker = row.find_element(By.XPATH, './/div[@id="t_tikcer"]').text
        writer.writerow([project_name, interest, category, received, goal, end_date, ticker])
except Exception as e:
    print(f"采集过程出错:{e}")
finally:
    # 保证资源一定会释放
    csv_file.close()
    driver.quit()
关键修改说明
  • 高度获取逻辑适配:同时取body和documentElement的高度最大值,适配目标站点的高度计算规则
  • 滚动触发优化:每次滚动超出当前高度100像素,确保触发页面的懒加载机制,不会卡在加载阈值上
  • 底部判定逻辑优化:连续3次检测高度无变化才判定滚动到底,避免网络卡顿、加载慢导致的提前终止
  • 代码结构调整:滚动加载和数据采集完全拆分,先加载完所有内容再统一采集,避免重复写入数据,也不会干扰滚动逻辑的判断
  • 语法适配:替换了Selenium旧版本已废弃的find_elements_by_xpath写法,兼容Selenium 4.x版本
  • 编码优化:csv文件新增utf-8编码和newline参数,避免Windows下打开csv出现乱码、空行问题
  • 资源释放优化:用finally块保证不管采集是否成功,都会关闭csv文件和浏览器,不会出现资源泄漏

内容的提问来源于stack exchange,提问作者Diop Chopra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 02:54:04