You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python Selenium定位页面右侧滚动条并滚动加载全部商品内容

页面动态加载商品抓取故障排查与解决

问题说明

需要抓取页面https://allinone.pospal.cn/m#/categories的全部商品信息,该页面存在双滚动条,右侧滚动条滚动至底部时会触发动态加载。目前尝试多种全局滚动方法均失败,仅能提取前20条商品,实际页面包含1500+条商品,附上尝试代码如下:

尝试过的代码

import time
import requests
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.common.keys import Keys
from selenium.webdriver.common.action_chains import ActionChains
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# Load the webpage
url = 'https://allinone.pospal.cn/m#/categories'
driver = webdriver.Chrome()
driver.get(url)

# Wait for the promotion image to load and click it
promotion_image = WebDriverWait(driver, 10).until(
    EC.presence_of_element_located(
        (By.XPATH, '//img[@src="//imgw.pospal.cn/we/westroe/img/categories/discount.png"]'))
)
promotion_image.click()


# Get focus on the right side of the page (two scroll bar, focus on the right)
# Wait for the div to load and get focus on the element 
items_div = WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.CSS_SELECTOR, 'div.yb-scrollable'))
)
#items_div.send_keys(Keys.NULL)
items_div.click()

# Do something with the div, e.g. get its text content
#@print(items_div.text) 
#-

#### attempt 1 - to scroll
##Scroll to the bottom of the page
# scroll_pause_time = 1
# scroll_step = 500
# scroll_height = driver.execute_script("return Math.max( document.body.scrollHeight, document.body.offsetHeight, document.documentElement.clientHeight, document.documentElement.scrollHeight, document.documentElement.offsetHeight );")
# while True:
#     driver.execute_script(f"window.scrollTo(0, {scroll_height});")
#     scroll_height_new = driver.execute_script("return Math.max( document.body.scrollHeight, document.body.offsetHeight, document.documentElement.clientHeight, document.documentElement.scrollHeight, document.documentElement.offsetHeight );")
#     if scroll_height_new == scroll_height:
#         break
#     scroll_height = scroll_height_new
#     time.sleep(scroll_pause_time)
# -
# -

#### attempt 2
# """A method for scrolling to the bottom of the page."""
# # Get scroll height.
# last_height = driver.execute_script("return document.body.scrollHeight")
# while True:
#     # Scroll down to the bottom.
#     driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
#     # Wait to load the page.
#     time.sleep(2)

#     # Calculate new scroll height and compare with last scroll height.
#     new_height = driver.execute_script("return document.body.scrollHeight")
#     if new_height == last_height:
#         break

#     last_height = new_height

#### attempt 3
lenOfPage = driver.execute_script("window.scrollTo(0, document.body.scrollHeight);var lenOfPage=document.body.scrollHeight;return lenOfPage;")
match=False
while(match==False):
    lastCount = lenOfPage
    time.sleep(3)
    lenOfPage = driver.execute_script("window.scrollTo(0, document.body.scrollHeight);var lenOfPage=document.body.scrollHeight;return lenOfPage;")
    if lastCount==lenOfPage:
        match=True


# Extract the content of the yb-item tags
soup = BeautifulSoup(driver.page_source, 'html.parser')
yb_items = soup.find_all('div', {'class': 'yb-item'})
for yb_item in yb_items:
    print(yb_item.text.strip())

# Close the browser window
driver.quit()

解决方案

之前的所有尝试都是针对全局页面滚动,但目标内容在独立的滚动容器(div.yb-scrollable)内,必须针对该容器执行滚动操作才能触发动态加载。

核心思路:

  • 定位到右侧的滚动容器元素
  • 循环执行JavaScript,将容器的scrollTop设置为其scrollHeight(即滚动到底部)
  • 每次滚动后等待内容加载,对比滚动前后的容器高度,直到高度不再变化(说明已加载完所有内容)

修改后的完整代码

import time
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# Load the webpage
url = 'https://allinone.pospal.cn/m#/categories'
driver = webdriver.Chrome()
driver.get(url)

# Wait for the promotion image to load and click it
promotion_image = WebDriverWait(driver, 10).until(
    EC.presence_of_element_located(
        (By.XPATH, '//img[@src="//imgw.pospal.cn/we/westroe/img/categories/discount.png"]'))
)
promotion_image.click()

# Wait for the scrollable container to load
items_div = WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.CSS_SELECTOR, 'div.yb-scrollable'))
)

# Scroll the container to load all content
scroll_pause_time = 2
last_height = driver.execute_script("return arguments[0].scrollHeight;", items_div)

while True:
    # Scroll to the bottom of the container
    driver.execute_script("arguments[0].scrollTop = arguments[0].scrollHeight;", items_div)
    # Wait for content to load
    time.sleep(scroll_pause_time)
    # Get new scroll height
    new_height = driver.execute_script("return arguments[0].scrollHeight;", items_div)
    # Break if height doesn't change
    if new_height == last_height:
        break
    last_height = new_height

# Extract all items
soup = BeautifulSoup(driver.page_source, 'html.parser')
yb_items = soup.find_all('div', {'class': 'yb-item'})
print(f"共抓取到 {len(yb_items)} 条商品")
for idx, yb_item in enumerate(yb_items, 1):
    print(f"--- 商品 {idx} ---")
    print(yb_item.text.strip())

# Close the browser
driver.quit()

内容的提问来源于stack exchange,提问作者Panco

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 01:57:38