You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:获取沃尔玛商品页面源代码提取UPC,遭CAPTCHA拦截

沃尔玛商品页UPC提取及CAPTCHA绕过方案

问题背景

业余开发者,需从以下沃尔玛商品页面提取仅存在于源码中的UPC码:

https://www.walmart.com/ip/Rice-Krispies-Treats-Original-Chewy-Crispy-Marshmallow-Squares-Ready-to-Eat-12-4-oz-16-Count/10818666?athbdg=L1600&from=/search

使用Anaconda Jupyter Lab开发,编写的三段代码均被沃尔玛CAPTCHA(提示语:Activate and hold the button to confirm that you’re human. Thank You!)拦截,无法获取页面源码。

已尝试的代码

  • 尝试1(urllib请求谷歌测试)
#import urllib2
import urllib.request as urllib2

response = urllib2.urlopen("http://google.de")
page_source = response.read()
page_source
  • 尝试2(urllib直接请求沃尔玛)
response = urllib2.urlopen("https://www.walmart.com/ip/Rice-Krispies-Treats-Original-Chewy-Crispy-Marshmallow-Squares-Ready-to-Eat-12-4-oz-16-Count/10818666?athbdg=L1600&from=/search")
page_source = response.read()
page_source
  • 尝试3(无头Selenium)
#pip install -U selenium

from selenium import webdriver
import time
from selenium.webdriver.chrome.options import Options

chrome_options = Options()
chrome_options.add_argument("--headless")
driver = webdriver.Chrome(options=chrome_options)
driver.get("https://www.walmart.com/ip/Rice-Krispies-Treats-Original-Chewy-Crispy-Marshmallow-Squares-Ready-to-Eat-12-4-oz-16-Count/10818666?athbdg=L1600&from=/search")

time.sleep(10)

html_source = driver.page_source
html_source

可行解决方案

1. 使用undetected-chromedriver绕过检测

无头模式和普通Selenium容易被识别,改用专门规避检测的undetected-chromedriver:

# 先安装:!pip install undetected-chromedriver
import undetected_chromedriver as uc
import time
import re

driver = uc.Chrome()
driver.get("https://www.walmart.com/ip/Rice-Krispies-Treats-Original-Chewy-Crispy-Marshmallow-Squares-Ready-to-Eat-12-4-oz-16-Count/10818666?athbdg=L1600&from=/search")

# 手动完成CAPTCHA验证(首次运行需要),之后等待页面加载
time.sleep(15)

html_source = driver.page_source
# 从源码中匹配UPC
upc_match = re.search(r'"upc":"(\d+)"', html_source)
if upc_match:
    print("提取到UPC:", upc_match.group(1))

driver.quit()

2. 手动验证后复用Cookies

先启动带界面浏览器,手动过CAPTCHA,保存Cookies后用请求库复用:

# 第一步:获取并保存Cookies
from selenium import webdriver
import pickle

driver = webdriver.Chrome()
driver.get("https://www.walmart.com/ip/Rice-Krispies-Treats-Original-Chewy-Crispy-Marshmallow-Squares-Ready-to-Eat-12-4-oz-16-Count/10818666?athbdg=L1600&from=/search")
# 手动完成CAPTCHA,页面加载完成后执行保存
pickle.dump(driver.get_cookies(), open("walmart_cookies.pkl", "wb"))
driver.quit()

# 第二步:使用Cookies请求页面并提取UPC
import urllib.request as urllib2
import re

opener = urllib2.build_opener()
# 加载保存的Cookies
cookies = pickle.load(open("walmart_cookies.pkl", "rb"))
for cookie in cookies:
    opener.addheaders.append(('Cookie', f"{cookie['name']}={cookie['value']}"))
# 添加真实浏览器请求头
opener.addheaders.append(('User-Agent', 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'))

response = opener.open("https://www.walmart.com/ip/Rice-Krispies-Treats-Original-Chewy-Crispy-Marshmallow-Squares-Ready-to-Eat-12-4-oz-16-Count/10818666?athbdg=L1600&from=/search")
page_source = response.read().decode('utf-8')

upc_match = re.search(r'"upc":"(\d+)"', page_source)
if upc_match:
    print("提取到UPC:", upc_match.group(1))

3. 优化普通Selenium配置(降低检测概率)

不想用第三方库的话,给Selenium添加更多模拟真实浏览器的参数:

from selenium import webdriver
import time
import re
from selenium.webdriver.chrome.options import Options

chrome_options = Options()
# 禁用自动化检测特征
chrome_options.add_argument("--disable-blink-features=AutomationControlled")
chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"])
chrome_options.add_experimental_option('useAutomationExtension', False)
# 设置真实用户代理
chrome_options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')
# 启用JavaScript
chrome_options.add_argument("--enable-javascript")

driver = webdriver.Chrome(options=chrome_options)
# 执行脚本隐藏webdriver标识
driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})")

driver.get("https://www.walmart.com/ip/Rice-Krispies-Treats-Original-Chewy-Crispy-Marshmallow-Squares-Ready-to-Eat-12-4-oz-16-Count/10818666?athbdg=L1600&from=/search")

# 手动完成CAPTCHA(首次运行)
time.sleep(15)

html_source = driver.page_source
upc_match = re.search(r'"upc":"(\d+)"', html_source)
if upc_match:
    print("提取到UPC:", upc_match.group(1))

driver.quit()

内容的提问来源于stack exchange,提问作者Brendan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 08:06:09