如何爬取滚动加载而非分页展示数据的网页?(Selenium+BeautifulSoup)
解决滚动加载网页爬取数据重复问题
问题根源
滚动加载后页面DOM中保留了重复的元素节点,导致你的选择器匹配到了重复内容,最终输出重复数据。
可行解决方案
1. 基于文本内容去重
利用集合自动去重的特性过滤重复文本,若需要保留加载顺序,可使用有序去重方式:
import time from selenium import webdriver from bs4 import BeautifulSoup driver = webdriver.Chrome(service=service, options=chrome_options) driver.get(url) scroll_pause_time = 2 last_height = driver.execute_script("return document.body.scrollHeight") while True: driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(scroll_pause_time) new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height html = driver.page_source bs = BeautifulSoup(html, 'html.parser') games = bs.find_all('div', {'class': 'GVj7ae imso-medium-font qJnhT imso-ani'}) # 提取非空文本并有序去重 game_texts = [game.get_text().strip() for game in games if game.get_text().strip()] unique_games = list(dict.fromkeys(game_texts)) for game in unique_games: print(game) driver.quit()
2. 定位元素的唯一标识(更可靠)
检查目标元素的HTML结构,找到每个游戏项的唯一属性(如data-id、id等),通过唯一属性过滤重复节点:
# 修改find_all的选择器,结合实际存在的唯一属性 games = bs.find_all('div', { 'class': 'GVj7ae imso-medium-font qJnhT imso-ani', 'data-match-id': True # 示例属性,需替换为页面实际存在的唯一标识 })
3. 滚动过程中逐步提取去重
每次滚动后立即提取当前页面的元素并加入集合去重,避免最后一次性提取时包含重复节点:
import time from selenium import webdriver from bs4 import BeautifulSoup driver = webdriver.Chrome(service=service, options=chrome_options) driver.get(url) scroll_pause_time = 2 last_height = driver.execute_script("return document.body.scrollHeight") unique_games = set() while True: # 提取当前页面元素并去重 html = driver.page_source bs = BeautifulSoup(html, 'html.parser') for game in bs.find_all('div', {'class': 'GVj7ae imso-medium-font qJnhT imso-ani'}): text = game.get_text().strip() if text: unique_games.add(text) # 滚动加载 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(scroll_pause_time) new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height # 输出结果(如需顺序可排序) for game in sorted(unique_games): print(game) driver.quit()
优化建议
- 替换
time.sleep()为Selenium的显式等待,等待新元素加载完成,提升稳定性:from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By # 滚动后等待新元素加载 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "GVj7ae")) ) - 检查网页的XHR请求,若滚动加载的数据来自API接口,直接调用接口获取数据会比模拟滚动更高效,且从根源避免DOM重复问题。
内容的提问来源于stack exchange,提问作者Carlos Ramirez
相关产品推荐
相关产品推荐

