You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何爬取滚动加载而非分页展示数据的网页?(Selenium+BeautifulSoup)

解决滚动加载网页爬取数据重复问题

问题根源

滚动加载后页面DOM中保留了重复的元素节点,导致你的选择器匹配到了重复内容,最终输出重复数据。

可行解决方案

1. 基于文本内容去重

利用集合自动去重的特性过滤重复文本,若需要保留加载顺序,可使用有序去重方式:

import time
from selenium import webdriver
from bs4 import BeautifulSoup

driver = webdriver.Chrome(service=service, options=chrome_options)
driver.get(url)

scroll_pause_time = 2 
last_height = driver.execute_script("return document.body.scrollHeight")

while True:
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(scroll_pause_time)
    new_height = driver.execute_script("return document.body.scrollHeight")
    if new_height == last_height:
        break
    last_height = new_height

html = driver.page_source
bs = BeautifulSoup(html, 'html.parser')

games = bs.find_all('div', {'class': 'GVj7ae imso-medium-font qJnhT imso-ani'})
# 提取非空文本并有序去重
game_texts = [game.get_text().strip() for game in games if game.get_text().strip()]
unique_games = list(dict.fromkeys(game_texts))

for game in unique_games:
    print(game)

driver.quit()

2. 定位元素的唯一标识(更可靠)

检查目标元素的HTML结构,找到每个游戏项的唯一属性(如data-id、id等),通过唯一属性过滤重复节点:

# 修改find_all的选择器,结合实际存在的唯一属性
games = bs.find_all('div', {
    'class': 'GVj7ae imso-medium-font qJnhT imso-ani',
    'data-match-id': True  # 示例属性,需替换为页面实际存在的唯一标识
})

3. 滚动过程中逐步提取去重

每次滚动后立即提取当前页面的元素并加入集合去重,避免最后一次性提取时包含重复节点:

import time
from selenium import webdriver
from bs4 import BeautifulSoup

driver = webdriver.Chrome(service=service, options=chrome_options)
driver.get(url)

scroll_pause_time = 2 
last_height = driver.execute_script("return document.body.scrollHeight")
unique_games = set()

while True:
    # 提取当前页面元素并去重
    html = driver.page_source
    bs = BeautifulSoup(html, 'html.parser')
    for game in bs.find_all('div', {'class': 'GVj7ae imso-medium-font qJnhT imso-ani'}):
        text = game.get_text().strip()
        if text:
            unique_games.add(text)
    
    # 滚动加载
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(scroll_pause_time)
    
    new_height = driver.execute_script("return document.body.scrollHeight")
    if new_height == last_height:
        break
    last_height = new_height

# 输出结果(如需顺序可排序)
for game in sorted(unique_games):
    print(game)

driver.quit()

优化建议

  • 替换time.sleep()为Selenium的显式等待,等待新元素加载完成,提升稳定性:
    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.support import expected_conditions as EC
    from selenium.webdriver.common.by import By
    
    # 滚动后等待新元素加载
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "GVj7ae"))
    )
    
  • 检查网页的XHR请求,若滚动加载的数据来自API接口,直接调用接口获取数据会比模拟滚动更高效,且从根源避免DOM重复问题。

内容的提问来源于stack exchange,提问作者Carlos Ramirez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 23:47:00