如何用Python Selenium定位页面右侧滚动条并滚动加载全部商品内容
页面动态加载商品抓取故障排查与解决
问题说明
需要抓取页面https://allinone.pospal.cn/m#/categories的全部商品信息,该页面存在双滚动条,右侧滚动条滚动至底部时会触发动态加载。目前尝试多种全局滚动方法均失败,仅能提取前20条商品,实际页面包含1500+条商品,附上尝试代码如下:
尝试过的代码
import time import requests from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.common.keys import Keys from selenium.webdriver.common.action_chains import ActionChains from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Load the webpage url = 'https://allinone.pospal.cn/m#/categories' driver = webdriver.Chrome() driver.get(url) # Wait for the promotion image to load and click it promotion_image = WebDriverWait(driver, 10).until( EC.presence_of_element_located( (By.XPATH, '//img[@src="//imgw.pospal.cn/we/westroe/img/categories/discount.png"]')) ) promotion_image.click() # Get focus on the right side of the page (two scroll bar, focus on the right) # Wait for the div to load and get focus on the element items_div = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, 'div.yb-scrollable')) ) #items_div.send_keys(Keys.NULL) items_div.click() # Do something with the div, e.g. get its text content #@print(items_div.text) #- #### attempt 1 - to scroll ##Scroll to the bottom of the page # scroll_pause_time = 1 # scroll_step = 500 # scroll_height = driver.execute_script("return Math.max( document.body.scrollHeight, document.body.offsetHeight, document.documentElement.clientHeight, document.documentElement.scrollHeight, document.documentElement.offsetHeight );") # while True: # driver.execute_script(f"window.scrollTo(0, {scroll_height});") # scroll_height_new = driver.execute_script("return Math.max( document.body.scrollHeight, document.body.offsetHeight, document.documentElement.clientHeight, document.documentElement.scrollHeight, document.documentElement.offsetHeight );") # if scroll_height_new == scroll_height: # break # scroll_height = scroll_height_new # time.sleep(scroll_pause_time) # - # - #### attempt 2 # """A method for scrolling to the bottom of the page.""" # # Get scroll height. # last_height = driver.execute_script("return document.body.scrollHeight") # while True: # # Scroll down to the bottom. # driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # # Wait to load the page. # time.sleep(2) # # Calculate new scroll height and compare with last scroll height. # new_height = driver.execute_script("return document.body.scrollHeight") # if new_height == last_height: # break # last_height = new_height #### attempt 3 lenOfPage = driver.execute_script("window.scrollTo(0, document.body.scrollHeight);var lenOfPage=document.body.scrollHeight;return lenOfPage;") match=False while(match==False): lastCount = lenOfPage time.sleep(3) lenOfPage = driver.execute_script("window.scrollTo(0, document.body.scrollHeight);var lenOfPage=document.body.scrollHeight;return lenOfPage;") if lastCount==lenOfPage: match=True # Extract the content of the yb-item tags soup = BeautifulSoup(driver.page_source, 'html.parser') yb_items = soup.find_all('div', {'class': 'yb-item'}) for yb_item in yb_items: print(yb_item.text.strip()) # Close the browser window driver.quit()
解决方案
之前的所有尝试都是针对全局页面滚动,但目标内容在独立的滚动容器(div.yb-scrollable)内,必须针对该容器执行滚动操作才能触发动态加载。
核心思路:
- 定位到右侧的滚动容器元素
- 循环执行JavaScript,将容器的
scrollTop设置为其scrollHeight(即滚动到底部) - 每次滚动后等待内容加载,对比滚动前后的容器高度,直到高度不再变化(说明已加载完所有内容)
修改后的完整代码
import time from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Load the webpage url = 'https://allinone.pospal.cn/m#/categories' driver = webdriver.Chrome() driver.get(url) # Wait for the promotion image to load and click it promotion_image = WebDriverWait(driver, 10).until( EC.presence_of_element_located( (By.XPATH, '//img[@src="//imgw.pospal.cn/we/westroe/img/categories/discount.png"]')) ) promotion_image.click() # Wait for the scrollable container to load items_div = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, 'div.yb-scrollable')) ) # Scroll the container to load all content scroll_pause_time = 2 last_height = driver.execute_script("return arguments[0].scrollHeight;", items_div) while True: # Scroll to the bottom of the container driver.execute_script("arguments[0].scrollTop = arguments[0].scrollHeight;", items_div) # Wait for content to load time.sleep(scroll_pause_time) # Get new scroll height new_height = driver.execute_script("return arguments[0].scrollHeight;", items_div) # Break if height doesn't change if new_height == last_height: break last_height = new_height # Extract all items soup = BeautifulSoup(driver.page_source, 'html.parser') yb_items = soup.find_all('div', {'class': 'yb-item'}) print(f"共抓取到 {len(yb_items)} 条商品") for idx, yb_item in enumerate(yb_items, 1): print(f"--- 商品 {idx} ---") print(yb_item.text.strip()) # Close the browser driver.quit()
内容的提问来源于stack exchange,提问作者Panco
相关产品推荐
相关产品推荐

