You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Selenium爬取Shopee商品链接遇空列表及数量不足问题求助

Shopee搜索结果商品链接爬取问题排查与解决

问题背景

需要用Python Selenium遍历Shopee搜索结果页面,爬取所有Crown饼干的商品链接,但遇到两个问题:

  1. 初始代码运行后输出空列表
  2. 添加滚动加载逻辑后,仅能获取15条链接

原代码问题分析

初始代码的核心问题

  • 页面加载时间不足:time.sleep(0.5)等待时间过短,Shopee动态商品内容未加载完成,导致find_elements返回空列表
  • 循环变量冲突:外层循环用i作为页码变量,内层循环又用i遍历元素,直接覆盖页码值,导致循环逻辑混乱
  • 动态加载未处理:Shopee商品采用滚动加载机制,直接获取元素只会拿到当前可见的少量商品
  • 结果存储逻辑错误:hlink每次循环都被重置为空列表,最后仅保留最后一页内容,且去重逻辑仅作用于最后一页

修改后代码的遗留问题

  • 结果存储依旧错误:hlink仍在循环内初始化,每次页面刷新后都会清空之前的爬取结果
  • 滚动逻辑不够精准:直接滚动到页面底部可能触发不了全部商品加载,部分商品卡片尚未渲染完成
  • 等待条件不够严格:presence_of_all_elements_located仅确保元素存在于DOM中,但未必可见,可能导致获取无效链接

修正后的完整代码

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

# 初始化浏览器
driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()))
driver.maximize_window()  # 最大化窗口,避免元素因窗口过小被隐藏
driver.implicitly_wait(10)  # 全局隐式等待

text = "b%C3%A1nh%20crown"
all_links = []  # 全局存储所有商品链接,避免循环清空

# 获取最大页码
url_maxpage = f"https://shopee.vn/search?brands=3372239&keyword={text}&noCorrection=true&page=0"
driver.get(url_maxpage)

max_page_element = WebDriverWait(driver, 15).until(
    EC.visibility_of_element_located((By.CLASS_NAME, 'shopee-mini-page-controller__total'))
)
max_page = int(max_page_element.text)

# 遍历所有页码
for page_num in range(max_page):
    url = f"https://shopee.vn/search?brands=3372239&keyword={text}&noCorrection=true&page={page_num}"
    driver.get(url)
    
    # 滚动加载所有商品
    last_height = driver.execute_script("return document.body.scrollHeight")
    scroll_count = 0
    while True:
        # 分步滚动,避免一次性滚到底部触发不了加载
        driver.execute_script("window.scrollBy(0, 800);")
        time.sleep(2)  # 根据网络情况调整延迟
        
        new_height = driver.execute_script("return document.body.scrollHeight")
        scroll_count += 1
        
        # 终止条件:高度不再变化,或滚动次数达到上限(防止无限循环)
        if new_height == last_height or scroll_count >= 10:
            break
        last_height = new_height
    
    # 等待所有商品链接加载完成并可见
    elements = WebDriverWait(driver, 20).until(
        EC.visibility_of_all_elements_located((By.CSS_SELECTOR, ".col-xs-2-4 a"))
    )
    
    # 提取链接并添加到全局列表
    for element in elements:
        link = element.get_attribute('href')
        if link not in all_links:  # 实时去重,避免重复存储
            all_links.append(link)

# 最终去重(双重保险)
unique_links = list(set(all_links))
print(f"共爬取到 {len(unique_links)} 条商品链接")
print(unique_links)

driver.quit()

关键修复说明

  • 全局结果存储:将all_links初始化在循环外,确保所有页面的链接都被累积存储
  • 滚动逻辑优化:采用分步滚动(每次滚动800px)+ 多次检查的方式,确保触发所有商品加载,同时设置滚动次数上限防止无限循环
  • 等待条件升级:使用visibility_of_all_elements_located确保元素可见,避免获取未渲染的无效链接
  • 变量名修正:将循环变量改为page_num,避免与内层循环变量冲突
  • 实时去重:在添加链接时检查是否已存在,减少后续去重开销
  • 窗口最大化:避免因窗口尺寸过小导致商品卡片被隐藏,无法被定位

内容的提问来源于stack exchange,提问作者jay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 07:25:14