You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Selenium实现网页无限滚动加载全部元素失败问题求助

问题描述

尝试用Python的Selenium实现网页滚动直至所有元素加载完成,但运行一段时间后脚本失败。调整sleep时长后,滚动一段时间元素加载过慢,Selenium会提前触达页面底部,希望实现等待下一批元素加载完成后再继续滚动。

用户提供的原始代码:

import time
import selenium
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.edge.options import Options


# Specify the path to the ChromeDriver executable
chrome_driver_path = r'C:\Users\ansar\edgedriver\msedgedriver.exe'

# Create a new instance of the Chrome driver
driver = webdriver.Edge(options=edge_options)
driver.maximize_window()
# Navigate to the desired webpage
url = 'https://www.bigbasket.com/ps/?q=rice'
driver.get(url)
# Approach 1

while True:
    driver.execute_script("window.scrollBy(0, 500);")
    time.sleep(0.5)

    #getting at the last of the page
    try:
        last = driver.find_element(By.XPATH,'//div[@class="w-full text-black text-center mt-10"]').text
        if last == '- Thats all folks -':
            print('we reached end of the page')
            time.sleep(120)
            break
    except:
        pass


# Approach 2 
SCROLL_PAUSE_TIME = 10

last_height = driver.execute_script("return document.body.scrollHeight")
while True: # Scroll down to bottom 
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
# Wait to load page
    time.sleep(SCROLL_PAUSE_TIME)

# Calculate new scroll height and compare with last scroll height
    new_height = driver.execute_script("return document.body.scrollHeight")
    if new_height == last_height:
        break
    last_height = new_height

优化解决方案

核心思路是放弃固定等待时长,改为等待页面元素数量增加或确认页面底部标识后再继续操作,用Selenium的显式等待替代time.sleep,精准匹配元素加载状态。

优化后的代码

import time
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.edge.options import Options

# 初始化Edge浏览器配置
edge_options = Options()
chrome_driver_path = r'C:\Users\ansar\edgedriver\msedgedriver.exe'
driver = webdriver.Edge(executable_path=chrome_driver_path, options=edge_options)
driver.maximize_window()

# 访问目标页面
url = 'https://www.bigbasket.com/ps/?q=rice'
driver.get(url)

# 初始化等待器,最长等待10秒
wait = WebDriverWait(driver, 10)
# 等待初始产品列表加载完成
wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "div[qa='product']")))

# 记录初始产品数量
previous_product_count = len(driver.find_elements(By.CSS_SELECTOR, "div[qa='product']"))
end_page_indicator = "- Thats all folks -"

while True:
    # 滚动到当前页面底部
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    
    try:
        # 等待新元素加载:直到产品数量增加,超时则进入下一步判断
        wait.until(lambda d: len(d.find_elements(By.CSS_SELECTOR, "div[qa='product']")) > previous_product_count)
        # 更新已加载产品数量
        previous_product_count = len(driver.find_elements(By.CSS_SELECTOR, "div[qa='product']"))
        print(f"当前已加载产品数: {previous_product_count}")
    except:
        # 检查是否到达页面底部
        try:
            page_bottom_text = driver.find_element(By.XPATH, '//div[@class="w-full text-black text-center mt-10"]').text
            if page_bottom_text == end_page_indicator:
                print("所有元素加载完成,已到达页面底部")
                break
        except:
            # 既无新元素也未到底部,短暂等待后重试
            time.sleep(2)
            continue

# 关闭浏览器
driver.quit()

关键优化点

  • 显式等待替代固定sleep:通过监听产品元素数量变化,确保新内容加载完成后再继续滚动,避免因加载延迟导致的提前触底。
  • 双重判断机制:当无法加载新元素时,先验证是否真的到达页面底部,避免误判为加载完成。
  • 动态跟踪元素状态:通过对比滚动前后的产品数量,精准判断是否有新内容加载,提升脚本稳定性。

内容的提问来源于stack exchange,提问作者Salman Ansari

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 01:24:59