Python Selenium实现网页无限滚动加载全部元素失败问题求助
问题描述
尝试用Python的Selenium实现网页滚动直至所有元素加载完成,但运行一段时间后脚本失败。调整sleep时长后,滚动一段时间元素加载过慢,Selenium会提前触达页面底部,希望实现等待下一批元素加载完成后再继续滚动。
用户提供的原始代码:
import time import selenium from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.edge.options import Options # Specify the path to the ChromeDriver executable chrome_driver_path = r'C:\Users\ansar\edgedriver\msedgedriver.exe' # Create a new instance of the Chrome driver driver = webdriver.Edge(options=edge_options) driver.maximize_window() # Navigate to the desired webpage url = 'https://www.bigbasket.com/ps/?q=rice' driver.get(url) # Approach 1 while True: driver.execute_script("window.scrollBy(0, 500);") time.sleep(0.5) #getting at the last of the page try: last = driver.find_element(By.XPATH,'//div[@class="w-full text-black text-center mt-10"]').text if last == '- Thats all folks -': print('we reached end of the page') time.sleep(120) break except: pass # Approach 2 SCROLL_PAUSE_TIME = 10 last_height = driver.execute_script("return document.body.scrollHeight") while True: # Scroll down to bottom driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # Wait to load page time.sleep(SCROLL_PAUSE_TIME) # Calculate new scroll height and compare with last scroll height new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height
优化解决方案
核心思路是放弃固定等待时长,改为等待页面元素数量增加或确认页面底部标识后再继续操作,用Selenium的显式等待替代time.sleep,精准匹配元素加载状态。
优化后的代码
import time from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.edge.options import Options # 初始化Edge浏览器配置 edge_options = Options() chrome_driver_path = r'C:\Users\ansar\edgedriver\msedgedriver.exe' driver = webdriver.Edge(executable_path=chrome_driver_path, options=edge_options) driver.maximize_window() # 访问目标页面 url = 'https://www.bigbasket.com/ps/?q=rice' driver.get(url) # 初始化等待器,最长等待10秒 wait = WebDriverWait(driver, 10) # 等待初始产品列表加载完成 wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "div[qa='product']"))) # 记录初始产品数量 previous_product_count = len(driver.find_elements(By.CSS_SELECTOR, "div[qa='product']")) end_page_indicator = "- Thats all folks -" while True: # 滚动到当前页面底部 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") try: # 等待新元素加载:直到产品数量增加,超时则进入下一步判断 wait.until(lambda d: len(d.find_elements(By.CSS_SELECTOR, "div[qa='product']")) > previous_product_count) # 更新已加载产品数量 previous_product_count = len(driver.find_elements(By.CSS_SELECTOR, "div[qa='product']")) print(f"当前已加载产品数: {previous_product_count}") except: # 检查是否到达页面底部 try: page_bottom_text = driver.find_element(By.XPATH, '//div[@class="w-full text-black text-center mt-10"]').text if page_bottom_text == end_page_indicator: print("所有元素加载完成,已到达页面底部") break except: # 既无新元素也未到底部,短暂等待后重试 time.sleep(2) continue # 关闭浏览器 driver.quit()
关键优化点
- 显式等待替代固定sleep:通过监听产品元素数量变化,确保新内容加载完成后再继续滚动,避免因加载延迟导致的提前触底。
- 双重判断机制:当无法加载新元素时,先验证是否真的到达页面底部,避免误判为加载完成。
- 动态跟踪元素状态:通过对比滚动前后的产品数量,精准判断是否有新内容加载,提升脚本稳定性。
内容的提问来源于stack exchange,提问作者Salman Ansari
相关产品推荐
相关产品推荐

