You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Selenium Python实现网页缓慢滚动?解决动态内容爬取加载问题

解决Selenium无限滚动加载不全的问题

嘿,我来帮你搞定这个无限滚动的问题!你遇到的情况很常见——像Twitter、Facebook这类动态加载内容的网站,对滚动的节奏要求比较高,太快或太机械的滚动都会导致内容加载不出来,甚至误判已经滚到页面底部。

先分析你现有代码的问题

你写的第二段代码里,一次性循环滚动100次(每次加600像素)后才去检查页面高度,这就导致了两个核心问题:

  1. 加载延迟未被处理:在这100次滚动过程中,页面可能还没来得及请求并渲染新内容,此时获取的new_height和初始last_height一致,直接触发了停止逻辑。
  2. 滚动逻辑太生硬:固定步长的滚动可能无法触发某些网站的加载机制,部分平台需要更“自然”的滚动交互才会加载新内容。

解决方案1:改进缓慢滚动,加入重试机制

这个方案会每次滚动到底部,等待内容加载,若高度未变化则重试几次(避免短暂加载延迟导致误判),甚至可以轻微回滚再滚动,触发平台的加载逻辑:

import time
from selenium import webdriver

driver = webdriver.Chrome()
driver.get("你的目标页面URL")

SCROLL_PAUSE_TIME = 1.5  # 根据页面加载速度调整
MAX_RETRIES = 3  # 连续几次高度不变后停止

# 注意:如果是Twitter/Facebook,可能需要找内部滚动容器,而不是body
# 比如Twitter的滚动容器是 div[data-testid="primaryColumn"]
last_height = driver.execute_script("return document.body.scrollHeight")
retry_count = 0

while retry_count < MAX_RETRIES:
    # 滚动到当前页面底部
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(SCROLL_PAUSE_TIME)
    
    new_height = driver.execute_script("return document.body.scrollHeight")
    
    if new_height == last_height:
        retry_count += 1
        print(f"暂未加载新内容,重试 {retry_count}/{MAX_RETRIES}...")
        # 轻微回滚再滚动,触发加载(部分平台需要这个交互)
        driver.execute_script("window.scrollBy(0, -150);")
        time.sleep(0.5)
        driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        time.sleep(SCROLL_PAUSE_TIME)
        # 再次检查高度
        new_height = driver.execute_script("return document.body.scrollHeight")
        if new_height != last_height:
            retry_count = 0
            last_height = new_height
    else:
        retry_count = 0
        last_height = new_height
        print(f"已加载新内容,当前页面高度:{new_height}")

print("滚动已停止,已到达页面底部或多次重试无新内容")

解决方案2:用显式等待替代固定sleep(更优雅)

这种方法利用Selenium的WebDriverWait显式等待,直到页面高度变化再继续滚动,比固定sleep更高效,也更贴合动态页面的加载节奏:

from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

driver = webdriver.Chrome()
driver.get("你的目标页面URL")

MAX_WAIT_TIME = 12  # 等待高度变化的最长时间(秒)
# 替换为对应平台的滚动容器,比如Twitter的容器:'return document.querySelector("div[data-testid=\'primaryColumn\']").scrollHeight'
get_scroll_height = "return document.body.scrollHeight"
scroll_to_bottom = "window.scrollTo(0, document.body.scrollHeight);"

last_height = driver.execute_script(get_scroll_height)

while True:
    driver.execute_script(scroll_to_bottom)
    try:
        # 等待页面高度变化,超时则认为无新内容
        WebDriverWait(driver, MAX_WAIT_TIME).until(
            lambda d: d.execute_script(get_scroll_height) != last_height
        )
        last_height = driver.execute_script(get_scroll_height)
        print(f"新内容加载完成,当前高度:{last_height}")
    except:
        print("无新内容加载,停止滚动")
        break

针对Twitter/Facebook的额外提示

这类平台的滚动容器通常不是整个body,而是某个内部元素:

  • Twitter:滚动容器是div[data-testid="primaryColumn"],你需要把代码中的document.body替换成这个选择器
  • Facebook:滚动容器一般是div[role="main"]或者类似的元素

替换示例:

# Twitter的滚动高度获取和滚动代码
get_scroll_height = 'return document.querySelector("div[data-testid=\'primaryColumn\']").scrollHeight'
scroll_to_bottom = 'document.querySelector("div[data-testid=\'primaryColumn\']").scrollTo(0, document.querySelector("div[data-testid=\'primaryColumn\']").scrollHeight);'

这样就能精准触发内容加载,避免因为滚动错误的容器导致的加载失败~

内容的提问来源于stack exchange,提问作者alex-uarent-alex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 16:52:32