使用BeautifulSoup和Selenium爬取Shopee印尼站仅获42条数据,如何获取全部60条?
解决Shopee印尼站搜索结果爬取不全问题
问题原因分析
- 无头模式(headless)下默认窗口尺寸过小,Shopee会根据窗口宽度调整加载的商品列数,导致单页商品数量不足60
- 固定次数的滚动逻辑可能未触发全部商品的动态加载
- 仅检查初始元素存在的等待条件,无法确保所有商品加载完成
修改后的完整代码
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC from time import sleep import pandas as pd # 搜索链接 url = 'https://shopee.co.id/search?keyword=obat%20kanker&page=0' path = '/Applications/chromedriver' # Chrome配置 chrome_options = Options() chrome_options.add_argument('start-maximized') chrome_options.add_argument('--no-sandbox') chrome_options.add_argument('--headless=new') # 新版无头模式,更接近正常浏览器行为 chrome_options.add_argument('disable-notifications') chrome_options.add_argument('--disable-infobars') chrome_options.add_argument('--window-size=1920,1080') # 设置固定窗口尺寸,确保加载完整商品列数 # 初始化浏览器 driver = webdriver.Chrome(executable_path=path, options=chrome_options) driver.get(url) # 等待初始商品容器加载 WebDriverWait(driver, 10).until(EC.presence_of_all_elements_located((By.CLASS_NAME, "shopee-search-item-result__item"))) # 循环滚动到底部,触发所有商品动态加载 last_height = driver.execute_script("return document.body.scrollHeight") while True: driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") sleep(2) # 等待新内容加载 new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height # 等待商品数量达到预期(最多等待10秒) WebDriverWait(driver, 10).until( lambda d: len(d.find_elements(By.CLASS_NAME, "shopee-search-item-result__item")) >= 60 ) # 获取页面源码并解析 html = driver.page_source soup = BeautifulSoup(html, "html.parser") # 提取商品数据 product_names = [item.get_text(strip=True) for item in soup.find_all('div', class_="ie3A+n bM+7UW Cve6sh")] product_prices = [item.get_text(strip=True) for item in soup.find_all('span', class_='ZEgDH9')] product_solds = [item.get_text(strip=True) if item else '0' for item in soup.find_all('div', class_="r6HknA uEPGHT")] # 输出结果数量 print(f"商品名称数量: {len(product_names)}") print(f"商品价格数量: {len(product_prices)}") print(f"商品销量数量: {len(product_solds)}") # 关闭浏览器 driver.quit() # 可选:将数据保存为CSV df = pd.DataFrame({ '商品名称': product_names, '价格': product_prices, '销量': product_solds }) df.to_csv('shopee_kanker_obat.csv', index=False, encoding='utf-8-sig')
关键修改说明
- 固定窗口尺寸:添加
--window-size=1920,1080,避免Shopee因窗口过小减少加载的商品列数 - 动态滚动逻辑:循环滚动直到页面高度不再变化,确保触发所有未加载的商品
- 强化等待条件:新增等待商品数量达到60的判断,确保所有商品加载完成后再提取数据
- 新版无头模式:使用
--headless=new替代旧版无头模式,降低被网站识别为爬虫的概率 - 冗余代码清理:移除重复的URL访问和不必要的模块导入
内容的提问来源于stack exchange,提问作者puyocode
相关产品推荐
相关产品推荐

