求助:获取沃尔玛商品页面源代码提取UPC,遭CAPTCHA拦截
沃尔玛商品页UPC提取及CAPTCHA绕过方案
问题背景
业余开发者,需从以下沃尔玛商品页面提取仅存在于源码中的UPC码:
https://www.walmart.com/ip/Rice-Krispies-Treats-Original-Chewy-Crispy-Marshmallow-Squares-Ready-to-Eat-12-4-oz-16-Count/10818666?athbdg=L1600&from=/search
使用Anaconda Jupyter Lab开发,编写的三段代码均被沃尔玛CAPTCHA(提示语:Activate and hold the button to confirm that you’re human. Thank You!)拦截,无法获取页面源码。
已尝试的代码
- 尝试1(urllib请求谷歌测试)
#import urllib2 import urllib.request as urllib2 response = urllib2.urlopen("http://google.de") page_source = response.read() page_source
- 尝试2(urllib直接请求沃尔玛)
response = urllib2.urlopen("https://www.walmart.com/ip/Rice-Krispies-Treats-Original-Chewy-Crispy-Marshmallow-Squares-Ready-to-Eat-12-4-oz-16-Count/10818666?athbdg=L1600&from=/search") page_source = response.read() page_source
- 尝试3(无头Selenium)
#pip install -U selenium from selenium import webdriver import time from selenium.webdriver.chrome.options import Options chrome_options = Options() chrome_options.add_argument("--headless") driver = webdriver.Chrome(options=chrome_options) driver.get("https://www.walmart.com/ip/Rice-Krispies-Treats-Original-Chewy-Crispy-Marshmallow-Squares-Ready-to-Eat-12-4-oz-16-Count/10818666?athbdg=L1600&from=/search") time.sleep(10) html_source = driver.page_source html_source
可行解决方案
1. 使用undetected-chromedriver绕过检测
无头模式和普通Selenium容易被识别,改用专门规避检测的undetected-chromedriver:
# 先安装:!pip install undetected-chromedriver import undetected_chromedriver as uc import time import re driver = uc.Chrome() driver.get("https://www.walmart.com/ip/Rice-Krispies-Treats-Original-Chewy-Crispy-Marshmallow-Squares-Ready-to-Eat-12-4-oz-16-Count/10818666?athbdg=L1600&from=/search") # 手动完成CAPTCHA验证(首次运行需要),之后等待页面加载 time.sleep(15) html_source = driver.page_source # 从源码中匹配UPC upc_match = re.search(r'"upc":"(\d+)"', html_source) if upc_match: print("提取到UPC:", upc_match.group(1)) driver.quit()
2. 手动验证后复用Cookies
先启动带界面浏览器,手动过CAPTCHA,保存Cookies后用请求库复用:
# 第一步:获取并保存Cookies from selenium import webdriver import pickle driver = webdriver.Chrome() driver.get("https://www.walmart.com/ip/Rice-Krispies-Treats-Original-Chewy-Crispy-Marshmallow-Squares-Ready-to-Eat-12-4-oz-16-Count/10818666?athbdg=L1600&from=/search") # 手动完成CAPTCHA,页面加载完成后执行保存 pickle.dump(driver.get_cookies(), open("walmart_cookies.pkl", "wb")) driver.quit() # 第二步:使用Cookies请求页面并提取UPC import urllib.request as urllib2 import re opener = urllib2.build_opener() # 加载保存的Cookies cookies = pickle.load(open("walmart_cookies.pkl", "rb")) for cookie in cookies: opener.addheaders.append(('Cookie', f"{cookie['name']}={cookie['value']}")) # 添加真实浏览器请求头 opener.addheaders.append(('User-Agent', 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')) response = opener.open("https://www.walmart.com/ip/Rice-Krispies-Treats-Original-Chewy-Crispy-Marshmallow-Squares-Ready-to-Eat-12-4-oz-16-Count/10818666?athbdg=L1600&from=/search") page_source = response.read().decode('utf-8') upc_match = re.search(r'"upc":"(\d+)"', page_source) if upc_match: print("提取到UPC:", upc_match.group(1))
3. 优化普通Selenium配置(降低检测概率)
不想用第三方库的话,给Selenium添加更多模拟真实浏览器的参数:
from selenium import webdriver import time import re from selenium.webdriver.chrome.options import Options chrome_options = Options() # 禁用自动化检测特征 chrome_options.add_argument("--disable-blink-features=AutomationControlled") chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"]) chrome_options.add_experimental_option('useAutomationExtension', False) # 设置真实用户代理 chrome_options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36') # 启用JavaScript chrome_options.add_argument("--enable-javascript") driver = webdriver.Chrome(options=chrome_options) # 执行脚本隐藏webdriver标识 driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})") driver.get("https://www.walmart.com/ip/Rice-Krispies-Treats-Original-Chewy-Crispy-Marshmallow-Squares-Ready-to-Eat-12-4-oz-16-Count/10818666?athbdg=L1600&from=/search") # 手动完成CAPTCHA(首次运行) time.sleep(15) html_source = driver.page_source upc_match = re.search(r'"upc":"(\d+)"', html_source) if upc_match: print("提取到UPC:", upc_match.group(1)) driver.quit()
内容的提问来源于stack exchange,提问作者Brendan
相关产品推荐
相关产品推荐

