You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Shopee商品数据抓取不稳定:字段交替为空问题求助

问题:Shopee菲律宾站美妆商品详情抓取不稳定,字段交替为空

成功获取Shopee菲律宾站美妆类商品URL,但抓取商品详情(店铺名、商品名、评分等字段)时数据不稳定,部分字段交替显示为空或正常内容。已尝试调整time.sleep()时长(1、1.5、2、5、15秒),排除网站封禁可能。

实现代码

导入模块

import bs4
import pandas as pd
import numpy as np
import random
import requests
from lxml import etree
import time
from tqdm.notebook import tqdm

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import ElementNotInteractableException
from time import sleep
from webdriver_manager.chrome import ChromeDriverManager

获取商品URL

driver = webdriver.Chrome(ChromeDriverManager().install())

for page in tqdm(range(5, 10)):
    driver.get("https://shopee.ph/Makeup-Fragrances-cat.11021036?facet=100664&page="+str(page)+"&sortBy=pop")
    
    skincare = driver.find_elements(By.XPATH, '//div[@class="col-xs-2-4 shopee-search-item-result__item"]//a[@data-sqe="link"]')

    for _skincare in tqdm(skincare):
        urls.append({"url":_skincare.get_attribute('href')})
driver.quit()

抓取商品详情

data_final = pd.DataFrame(urls)

driver = webdriver.Chrome(ChromeDriverManager().install())
skincares = []

for product in tqdm(data_final["url"]):
    driver.get(product)
    try:
        company = driver.find_element(By.XPATH,"//div[@class='CKGyuW']//div[@class='_1Yaflp page-product__shop']//div[@class='_1YY3XU']//div[@class='zYQ1eS']//div[@class='_3LoNDM']").text
    except:
        company = 'none'
    try:
        product_name = driver.find_element(By.XPATH,"//div[@class='flex flex-auto eTjGTe']//div[@class='flex-auto flex-column  _1Kkkb-']//div[@class='_2rQP1z']//span").text
    except:
        product_name = 'none'
    try:
        rating = driver.find_element(By.XPATH,"//div[@class='flex _3tkSsu']//div[@class='flex _3T9OoL']//div[@class='_3y5XOB _14izon']").text
    except:
        rating = 'none'
    try:
        number_of_ratings = driver.find_element(By.XPATH,"//div[@class='flex _3tkSsu']//div[@class='flex _3T9OoL']//div[@class='_3y5XOB']").text
    except:
        number_of_ratings = 'none'
    try:
        sold = driver.find_element(By.XPATH,"//div[@class='flex _3tkSsu']//div[@class='flex _3EOMd6']//div[@class='HmRxgn']").text
    except:
        sold = 'none'
    try:
        price = driver.find_element(By.XPATH,"//div[@class='_2Shl1j']").text
    except:
        price = 'none'
    try:
        description = driver.find_element(By.XPATH,"//div[@class='_1MqcWX']//p[@class='_2jrvqA']").text
    except:
        description = 'none'
        
    skincares.append({
        "url": product,
        "company": company,
        "product name": product_name,
        "rating": rating,
        "number of ratings": number_of_ratings,
        "sold": sold,
        "price": price,
        "description": description
        })
    time.sleep(5)

问题原因分析

  • 动态class依赖问题:Shopee前端的class名称是动态生成的,每次页面加载可能会变更,导致硬编码的XPATH定位失败
  • 固定等待不可靠:time.sleep()是固定时长等待,无法适配网络波动或页面异步加载的差异,元素可能未完成渲染就执行定位操作
  • XPATH层级冗余:过长的XPATH层级增加了定位失败概率,只要中间某一层的class变化,整个路径就失效

解决建议

1. 替换动态class为稳定定位属性

优先使用Shopee页面中带有的data-sqe语义化属性(这类属性不会随前端渲染变更),或者基于元素功能特征定位,避免依赖动态class:

# 示例:用data-sqe属性定位店铺名
company = WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.XPATH, "//div[@data-sqe='shop-name']"))
).text

# 示例:用data-sqe属性定位商品名
product_name = WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.XPATH, "//span[@data-sqe='name']"))
).text

2. 用显式等待替代固定sleep

使用WebDriverWait结合expected_conditions,确保元素加载完成后再获取数据,彻底解决加载时机问题:

# 封装通用获取文本函数,带重试
def get_element_text(driver, locator, wait_time=10, retry=2):
    for _ in range(retry + 1):
        try:
            return WebDriverWait(driver, wait_time).until(
                EC.presence_of_element_located(locator)
            ).text
        except:
            continue
    return 'none'

# 使用示例
rating = get_element_text(driver, (By.XPATH, "//div[@data-sqe='rating']"))
sold = get_element_text(driver, (By.XPATH, "//div[@data-sqe='sold']"))

3. 优化XPATH写法,缩短层级

避免冗余的层级依赖,只保留关键定位特征:

  • 原评分XPATH://div[@class='flex _3tkSsu']//div[@class='flex _3T9OoL']//div[@class='_3y5XOB _14izon']
  • 优化后://div[contains(@class, '_14izon')] 或直接用data-sqe属性

4. 确保页面完全加载

添加页面就绪状态检查,覆盖异步加载场景:

# 等待页面完全加载完成
WebDriverWait(driver, 15).until(
    lambda d: d.execute_script('return document.readyState') == 'complete'
)

内容的提问来源于stack exchange,提问作者user20576555

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 12:40:24