You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium与BeautifulSoup爬取Lazada翻页数据重复问题求解

问题背景

爬取目标为Lazada新加坡站SG Mart全商品列表页,同时使用BeautifulSoup(bs4)与Selenium实现爬取时,翻页到约35页后会出现页面数据重复的异常,手动正常浏览网站不会触发该问题。
原实现代码如下:

driver = webdriver.Chrome(service=chrome_driver_path)
driver.get('https://www.lazada.sg/the-sg-mart/?from=wangpu&langFlag=en&page=1&pageTypeId=2&q=All-Products&sort=pricedesc')


names = []
prices = []

soup_names = []
soup_prices = []
while True:
    for i in range(1, 41):
        path = '//*[@id="root"]/div/div[3]/div[1]/div/div[1]/div[2]/div[' + str(i) + ']/div/div/div[2]/div[2]/a'    
        name = driver.find_element(By.XPATH, path).text
        if name in names:
            print(i)
            pass

        path_price        =    '//*[@id="root"]/div/div[3]/div[1]/div/div[1]/div[2]/div[' + str(i) + ']/div/div/div[2]/div[3]/span'   
        price = driver.find_element(By.XPATH, path_price) .text
        
        names.append(name)
        prices.append(price)
     
        
        for item in soup.find_all(class_='RfADt'):
            print(item.a.get('title'))
            soup_names.append(item.a.get('title'))
        
        for item in soup.find_all(class_='aBrP0'):
            print(item.text)
            soup_prices.append(item.text)
    
   
    
    button=driver.find_element_by_xpath("//li[@title='Next Page']")
    driver.execute_script("arguments[0].click();", button)
    print('new page')
    sleep(5)
问题产生原因
  • 触发平台反爬风控:Lazada对连续翻页的请求有行为检测,原代码固定间隔5秒翻页、无任何真人操作模拟,请求特征和真人浏览差异过大,爬取到30页以上后会被风控拦截,平台不再返回新的页面数据,而是返回之前缓存过的列表页内容,导致数据重复。手动浏览时存在随机停留、鼠标滚动、不规则操作间隔等真人特征,不会触发该风控。
  • 页面跳转判断缺失:点击下一页后仅固定等待5秒,没有校验页面是否真的完成跳转、新页商品是否加载完成。受网络波动、风控拦截影响,很多时候点击下一页后页面停留在原页,代码就直接开始抓取,自然拿到重复数据。
  • bs4解析逻辑完全错误:代码中没有在每次翻页后重新获取页面源码更新soup对象,soup始终保存的是首次打开页面的第一页内容;且把全页元素解析的逻辑写在了单商品遍历的循环内,单页40个商品就会重复把第一页的商品数据存入列表40次,从代码运行开始就会产生重复数据。
  • 异常重试逻辑缺失:使用了Selenium已废弃的find_element_by_xpath方法,点击下一页后没有校验当前页码是否递增,出现跳转失败、按钮点击无效的情况时没有重试机制,会持续抓取当前页数据。
修复方案
  • 补充真人行为模拟:翻页间隔设置为2-5秒随机时长,页面加载完成后模拟随机滚动页面,降低被风控识别的概率。
  • 新增页面加载校验:点击下一页后,首先校验URL中的page参数是否为上一页页码+1,同时等待页面首个商品元素刷新完成,确认新页加载成功后再开始抓取。
  • 修正bs4解析逻辑:每次确认新页加载完成后,重新调用driver.page_source初始化soup对象,单页仅做一次全页元素解析,禁止在单商品循环内重复解析全页内容。
  • 增加失败重试兜底:每抓取完一页,对比当前页商品和上一页商品的重合度,重合度超过80%则判定为跳转失败,自动触发最多2次重试,重试失败则终止爬取。
  • 替换Selenium废弃方法,统一使用find_element(By.XPATH, 路径)的新写法。

修正后的参考代码:

import time
import random
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

driver = webdriver.Chrome(service=chrome_driver_path)
driver.get('https://www.lazada.sg/the-sg-mart/?from=wangpu&langFlag=en&page=1&pageTypeId=2&q=All-Products&sort=pricedesc')
wait = WebDriverWait(driver, 10)

names = []
prices = []
soup_names = []
soup_prices = []
current_page = 1

while True:
    # 随机等待+模拟滚动,模拟真人行为
    time.sleep(random.uniform(2,4))
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight/2)")
    time.sleep(random.uniform(0.5,1.5))
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight)")
    time.sleep(random.uniform(0.5,1))

    # 等待商品加载完成
    wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'RfADt')))
    
    # 重新初始化soup,单页只解析一次
    soup = BeautifulSoup(driver.page_source, 'html.parser')
    page_soup_names = [item.a.get('title') for item in soup.find_all(class_='RfADt')]
    page_soup_prices = [item.text for item in soup.find_all(class_='aBrP0')]
    
    # 用Selenium抓商品数据
    page_names = []
    page_prices = []
    for i in range(1, 41):
        try:
            name_path = f'//*[@id="root"]/div/div[3]/div[1]/div/div[1]/div[2]/div[{i}]/div/div/div[2]/div[2]/a'
            name = driver.find_element(By.XPATH, name_path).text
            price_path = f'//*[@id="root"]/div/div[3]/div[1]/div/div[1]/div[2]/div[{i}]/div/div/div[2]/div[3]/span'
            price = driver.find_element(By.XPATH, price_path).text
            page_names.append(name)
            page_prices.append(price)
        except:
            break
    
    # 重复校验
    if len(set(page_names) & set(names[-40:])) > 32:
        print("页面重复,重试")
        time.sleep(3)
        continue
    
    # 存入总列表
    names.extend(page_names)
    prices.extend(page_prices)
    soup_names.extend(page_soup_names)
    soup_prices.extend(page_soup_prices)
    print(f"第{current_page}页抓取完成,累计{len(names)}条商品")
    
    # 翻页
    try:
        next_btn = wait.until(EC.element_to_be_clickable((By.XPATH, "//li[@title='Next Page']")))
        # 校验下一页按钮是否可点,不可点说明到最后一页
        if 'disabled' in next_btn.get_attribute('class'):
            print("已到最后一页,爬取结束")
            break
        driver.execute_script("arguments[0].click();", next_btn)
        # 等待URL页码更新
        wait.until(EC.url_contains(f"page={current_page+1}"))
        current_page +=1
    except Exception as e:
        print(f"翻页失败: {e}")
        break

driver.quit()

内容的提问来源于stack exchange,提问作者Nuri Taş

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 15:03:10