Python爬取Lazada多页商品数据失败问题排查求助
爬取Lazada Guardian店铺商品遇到的问题与修复
问题描述
尝试用Python爬取Lazada马来西亚站Guardian店铺的全部商品名称和价格,该店铺共102页商品,但目前仅能提取第一页数据,点击下一页时触发报错。目标页面URL:https://www.lazada.com.my/guardian/?from=wangpu&langFlag=en&page=1&pageTypeId=2&q=All-Products
原实现代码
import time from selenium import webdriver from bs4 import BeautifulSoup import pandas as pd from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By class ScrapeLazada(): def scrape(self): url = 'https://www.lazada.com.my/guardian/?from=wangpu&langFlag=en&page=1&pageTypeId=2&q=All-Products' driver = webdriver.Chrome() driver.get(url) products=[] for i in range(102): WebDriverWait(driver, 5).until(EC.presence_of_element_located((By.CSS_SELECTOR, "#root"))) time.sleep(2) soup = BeautifulSoup(driver.page_source, "html.parser") for item in soup.findAll('div', class_='Bm3ON'): product_name = item.find('div', class_='RfADt').text price = item.find('span', class_='ooOxS').text.replace('RM', '') products.append( (product_name, price) ) time.sleep(2) driver.find_element(By.CSS_SELECTOR, ".ant-pagination-next > button").click() time.sleep(3) df = pd.DataFrame(products, columns=['Product Name', 'Price']) print(df) df.to_excel('Lazada_Guardian_Scrape.xlsx', index=False) print('Data saved in local disk') driver.close() sl = ScrapeLazada() sl.scrape()
运行报错信息
Product Name Price 0 UPHAMOL 250 Children Suspension Delicious Oran... 7.80 1 Darlie Double Action Fresh + Clean Toothpaste ... 20.92 ... 38 Avene Pre-Serum Hydrating Essence-In-Lotion 200Ml 87.30 39 Guardian Essential Lavender Refreshing Body Wa... 10.10 Traceback (most recent call last): File "Lazada_Guardian.py", line 43, in <module> sl.scrape() File "Lazada_Guardian.py", line 30, in scrape driver.find_element(By.CSS_SELECTOR, ".ant-pagination-next > button").click() ... selenium.common.exceptions.ElementClickInterceptedException: Message: element click intercepted: Element <button class="ant-pagination-item-link" type="button" tabindex="-1">...</button> is not clickable at point (1186, 693). Other element would receive the click: <html lang="en" class=" ">...</html>
问题分析与修复方案
核心问题
ElementClickInterceptedException报错说明下一页按钮被页面元素(弹窗、可视区域外遮挡)拦截,无法直接触发点击;同时原代码存在逻辑冗余、未处理页面加载状态等问题。
修复后代码
import time from selenium import webdriver from bs4 import BeautifulSoup import pandas as pd from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from selenium.webdriver.common.action_chains import ActionChains class ScrapeLazada(): def scrape(self): url = 'https://www.lazada.com.my/guardian/?from=wangpu&langFlag=en&page=1&pageTypeId=2&q=All-Products' driver = webdriver.Chrome() driver.maximize_window() # 最大化窗口减少遮挡 driver.get(url) # 处理Cookie弹窗(如果存在) try: cookie_btn = WebDriverWait(driver, 5).until(EC.element_to_be_clickable((By.CSS_SELECTOR, "button.cookie-btn-accept"))) cookie_btn.click() except: pass products = [] # 第一页已加载,循环101次获取剩余页面 for i in range(101): # 等待所有商品元素加载完成 WebDriverWait(driver, 10).until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.Bm3ON"))) time.sleep(1) soup = BeautifulSoup(driver.page_source, "html.parser") for item in soup.findAll('div', class_='Bm3ON'): try: product_name = item.find('div', class_='RfADt').text.strip() price = item.find('span', class_='ooOxS').text.replace('RM', '').strip() products.append((product_name, price)) except AttributeError: # 跳过解析失败的商品 continue # 处理下一页点击 try: # 等待下一页按钮可点击,排除禁用状态 next_btn = WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.CSS_SELECTOR, ".ant-pagination-next:not(.ant-pagination-disabled) > button"))) ActionChains(driver).move_to_element(next_btn).perform() # 滚动到按钮可视区域 time.sleep(1) next_btn.click() time.sleep(3) except: print("已到达最后一页或无法点击下一页") break # 循环结束后统一保存数据 df = pd.DataFrame(products, columns=['Product Name', 'Price']) print(df) df.to_excel('Lazada_Guardian_Scrape.xlsx', index=False) print('数据已保存到本地') driver.close() sl = ScrapeLazada() sl.scrape()
关键修改点
- 增加窗口最大化,减少页面元素遮挡概率
- 添加Cookie弹窗处理逻辑,避免弹窗拦截后续操作
- 调整循环次数为101次,因第一页已提前加载
- 使用
presence_of_all_elements_located等待所有商品加载,确保解析完整性 - 对商品解析增加异常捕获,跳过解析失败的条目
- 等待下一页按钮处于可点击状态,通过
ActionChains滚动到按钮可视区域后再点击 - 将Excel保存操作移到循环外,提升运行效率
- 增加下一页点击的异常捕获,避免因最后一页按钮禁用导致程序崩溃
内容的提问来源于stack exchange,提问作者程家乐
相关产品推荐
相关产品推荐

