You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取Lazada多页商品数据失败问题排查求助

爬取Lazada Guardian店铺商品遇到的问题与修复

问题描述

尝试用Python爬取Lazada马来西亚站Guardian店铺的全部商品名称和价格,该店铺共102页商品,但目前仅能提取第一页数据,点击下一页时触发报错。目标页面URL:https://www.lazada.com.my/guardian/?from=wangpu&langFlag=en&page=1&pageTypeId=2&q=All-Products

原实现代码

import time
from selenium import webdriver
from bs4 import BeautifulSoup
import pandas as pd
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

class ScrapeLazada():
    
    def scrape(self):
        url = 'https://www.lazada.com.my/guardian/?from=wangpu&langFlag=en&page=1&pageTypeId=2&q=All-Products'
        driver = webdriver.Chrome()
        driver.get(url)
        
        products=[]
        for i in range(102):
            WebDriverWait(driver, 5).until(EC.presence_of_element_located((By.CSS_SELECTOR, "#root")))
            time.sleep(2)

            soup = BeautifulSoup(driver.page_source, "html.parser")
            for item in soup.findAll('div', class_='Bm3ON'):
                product_name = item.find('div', class_='RfADt').text
                price = item.find('span', class_='ooOxS').text.replace('RM', '')
                products.append(
                    (product_name, price) 
                )

            time.sleep(2)
            driver.find_element(By.CSS_SELECTOR, ".ant-pagination-next > button").click()
            time.sleep(3)

            df = pd.DataFrame(products, columns=['Product Name', 'Price'])
            print(df)

            df.to_excel('Lazada_Guardian_Scrape.xlsx', index=False)
            print('Data saved in local disk')
    
    
        driver.close()
        
sl = ScrapeLazada()
sl.scrape()

运行报错信息

Product Name   Price
0   UPHAMOL 250 Children Suspension Delicious Oran...    7.80
1   Darlie Double Action Fresh + Clean Toothpaste ...   20.92
...
38  Avene Pre-Serum Hydrating Essence-In-Lotion 200Ml   87.30
39  Guardian Essential Lavender Refreshing Body Wa...   10.10
Traceback (most recent call last):
  File "Lazada_Guardian.py", line 43, in <module>
    sl.scrape()
  File "Lazada_Guardian.py", line 30, in scrape
    driver.find_element(By.CSS_SELECTOR, ".ant-pagination-next > button").click()
...
selenium.common.exceptions.ElementClickInterceptedException: Message: element click intercepted: Element <button class="ant-pagination-item-link" type="button" tabindex="-1">...</button> is not clickable at point (1186, 693). Other element would receive the click: <html lang="en" class=" ">...</html>

问题分析与修复方案

核心问题

ElementClickInterceptedException报错说明下一页按钮被页面元素(弹窗、可视区域外遮挡)拦截,无法直接触发点击;同时原代码存在逻辑冗余、未处理页面加载状态等问题。

修复后代码

import time
from selenium import webdriver
from bs4 import BeautifulSoup
import pandas as pd
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from selenium.webdriver.common.action_chains import ActionChains

class ScrapeLazada():
    
    def scrape(self):
        url = 'https://www.lazada.com.my/guardian/?from=wangpu&langFlag=en&page=1&pageTypeId=2&q=All-Products'
        driver = webdriver.Chrome()
        driver.maximize_window()  # 最大化窗口减少遮挡
        driver.get(url)
        
        # 处理Cookie弹窗(如果存在)
        try:
            cookie_btn = WebDriverWait(driver, 5).until(EC.element_to_be_clickable((By.CSS_SELECTOR, "button.cookie-btn-accept")))
            cookie_btn.click()
        except:
            pass
        
        products = []
        # 第一页已加载,循环101次获取剩余页面
        for i in range(101):
            # 等待所有商品元素加载完成
            WebDriverWait(driver, 10).until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.Bm3ON")))
            time.sleep(1)

            soup = BeautifulSoup(driver.page_source, "html.parser")
            for item in soup.findAll('div', class_='Bm3ON'):
                try:
                    product_name = item.find('div', class_='RfADt').text.strip()
                    price = item.find('span', class_='ooOxS').text.replace('RM', '').strip()
                    products.append((product_name, price))
                except AttributeError:
                    # 跳过解析失败的商品
                    continue

            # 处理下一页点击
            try:
                # 等待下一页按钮可点击,排除禁用状态
                next_btn = WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, ".ant-pagination-next:not(.ant-pagination-disabled) > button")))
                ActionChains(driver).move_to_element(next_btn).perform()  # 滚动到按钮可视区域
                time.sleep(1)
                next_btn.click()
                time.sleep(3)
            except:
                print("已到达最后一页或无法点击下一页")
                break

        # 循环结束后统一保存数据
        df = pd.DataFrame(products, columns=['Product Name', 'Price'])
        print(df)
        df.to_excel('Lazada_Guardian_Scrape.xlsx', index=False)
        print('数据已保存到本地')
        
        driver.close()
        
sl = ScrapeLazada()
sl.scrape()

关键修改点

  1. 增加窗口最大化,减少页面元素遮挡概率
  2. 添加Cookie弹窗处理逻辑,避免弹窗拦截后续操作
  3. 调整循环次数为101次,因第一页已提前加载
  4. 使用presence_of_all_elements_located等待所有商品加载,确保解析完整性
  5. 对商品解析增加异常捕获,跳过解析失败的条目
  6. 等待下一页按钮处于可点击状态,通过ActionChains滚动到按钮可视区域后再点击
  7. 将Excel保存操作移到循环外,提升运行效率
  8. 增加下一页点击的异常捕获,避免因最后一页按钮禁用导致程序崩溃

内容的提问来源于stack exchange,提问作者程家乐

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 06:25:04