如何用Python抓取带多规格的Shopee静态商品页面?
解决Shopee多规格商品信息抓取问题
你的现有代码仅能获取默认规格的价格,要抓取多规格商品的所有规格项及对应价格,这里提供两种可靠方案:
方法一:解析页面内置商品数据(推荐,高效稳定)
Shopee会将商品完整数据(含所有规格、价格、库存等)嵌入页面的__NEXT_DATA__脚本标签中,直接解析该JSON数据即可一次性获取所有规格信息,无需模拟点击操作。
修改后的完整代码:
!apt-get update !apt install chromium-chromedriver !cp /usr/lib/chromium-browser/chromedriver /usr/bin !pip install selenium-wire pandas numpy # 导入依赖库 from seleniumwire import webdriver from selenium.webdriver.common.by import By import pandas as pd import json from time import sleep from random import randint from datetime import datetime today = datetime.today().strftime('%Y-%m-%d') # 配置浏览器选项 options = webdriver.ChromeOptions() options.set_capability("goog:loggingPrefs", {"performance": "ALL", "browser": "ALL"}) options.add_argument('--headless') options.add_argument('--no-sandbox') options.add_argument('--disable-dev-shm-usage') driver = webdriver.Chrome('chromedriver', options=options) # 目标商品链接 shopee=['https://shopee.co.id/ACMIC-Braided-Line-Kabel-Data-Fast-Charging-for-iPhone-1-M-2-M-3-M-i.27769962.18163430950?sp_atk=5c463b34-ab0b-40da-af85-05206b95f616&xptdk=5c463b34-ab0b-40da-af85-05206b95f616'] shopeedf=pd.DataFrame() for urls in shopee: try: driver.get(urls) sleep(randint(3,5)) # 提取页面内置的商品JSON数据 script_data = driver.find_element(By.ID, "__NEXT_DATA__").get_attribute('innerHTML') product_json = json.loads(script_data) # 基础商品信息 product_name = product_json['props']['pageProps']['item']['name'] compid = urls.split(".")[4].split("?")[0] # 遍历所有规格项 variations = product_json['props']['pageProps']['item']['models'] for var in variations: spec_name = var['name'] # Shopee价格单位为分,需转换为实际金额 normal_price = str(var['price'] // 100000) discounted_price = str(var['promotion_price'] // 100000) if var.get('promotion_price') else normal_price discount = str(var['discount']) if var.get('discount') else "0" # 组装数据行 dat={ 'product_name': product_name, 'spec_name': spec_name, 'normal_price': normal_price, 'discounted_price': discounted_price, 'discount': discount, 'competitor_id': compid, 'url': urls, 'date_key': today, 'web': 'shopee' } dat=pd.DataFrame([dat]) shopeedf=pd.concat([shopeedf, dat], ignore_index=True) print(f"成功抓取{len(variations)}个规格数据") except Exception as e: print(f"{urls} error") print(e) # 输出结果 print(shopeedf)
方法二:模拟点击规格选项获取价格(适配JSON结构变动场景)
若页面JSON结构更新,可通过模拟点击每个规格选项,获取对应规格的价格。此方法需处理多规格组合(如颜色+尺寸)的逻辑,示例页面仅含单规格(长度),代码如下:
!apt-get update !apt install chromium-chromedriver !cp /usr/lib/chromium-browser/chromedriver /usr/bin !pip install selenium-wire pandas numpy from seleniumwire import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pandas as pd from time import sleep from random import randint from datetime import datetime today = datetime.today().strftime('%Y-%m-%d') options = webdriver.ChromeOptions() options.set_capability("goog:loggingPrefs", {"performance": "ALL", "browser": "ALL"}) options.add_argument('--headless') options.add_argument('--no-sandbox') options.add_argument('--disable-dev-shm-usage') driver = webdriver.Chrome('chromedriver', options=options) shopee=['https://shopee.co.id/ACMIC-Braided-Line-Kabel-Data-Fast-Charging-for-iPhone-1-M-2-M-3-M-i.27769962.18163430950?sp_atk=5c463b34-ab0b-40da-af85-05206b95f616&xptdk=5c463b34-ab0b-40da-af85-05206b95f616'] shopeedf=pd.DataFrame() for urls in shopee: try: driver.get(urls) sleep(randint(3,5)) product_name = driver.find_element(By.CSS_SELECTOR, ".YPqix5").text compid = urls.split(".")[4].split("?")[0] # 等待规格选项加载完成 spec_options = WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CSS_SELECTOR, ".product-variation .flex.items-center.justify-center")) ) # 遍历每个规格选项 for option in spec_options: option.click() sleep(1) # 获取当前规格价格 try: normal_price = driver.find_element(By.CSS_SELECTOR, ".Kg2R-S").text.replace('Rp','').replace(".","") except: normal_price = driver.find_element(By.CSS_SELECTOR, ".X0xUb5").text.replace('Rp','').replace(".","") # 获取折扣信息 try: discount = driver.find_element(By.CSS_SELECTOR, ".+1IO+x").text except: discount = "0" # 组装数据行 dat={ 'product_name': product_name, 'spec_name': option.text.strip(), 'normal_price': normal_price, 'discount': discount, 'competitor_id': compid, 'url': urls, 'date_key': today, 'web': 'shopee' } dat=pd.DataFrame([dat]) shopeedf=pd.concat([shopeedf, dat], ignore_index=True) print(f"成功抓取{len(spec_options)}个规格数据") except Exception as e: print(f"{urls} error") print(e) print(shopeedf)
注意事项
- 方法一的JSON字段路径可能随Shopee页面更新变动,需定期验证
- 抓取时添加随机延迟,避免触发反爬机制
- 若遇到验证码,需补充验证码处理逻辑或更换代理IP
内容的提问来源于stack exchange,提问作者Hal
相关产品推荐
相关产品推荐

