You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Walmart爬虫启用代理时Cookie拦截致数据缺失如何解决

Walmart爬虫Cookie失效问题解决方案

问题复现与根因

  • 业务场景:300款目标商品搭配100个邮政编码,爬取Walmart不同区域的商品售价,未启用代理时测试6组任务可正常返回结构化数据,字段包含商品ID、UPC编码、邮政编码、门店ID、售价、采集时间,样例输出:
43819800,041167412213,10003,3520,19.96,2022-07-05 14:06:47
43819800,041167412213,48104,5472,19.96,2022-07-05 14:06:47
224749468,300450206909,10003,3520,42.47,2022-07-05 14:06:49
224749468,300450206909,48104,5472,42.47,2022-07-05 14:06:50
14053317,681131187091,10003,3520,2.52,2022-07-05 14:06:51
14053317,681131187091,48104,5472,2.52,2022-07-05 14:06:52
  • 启用代理后大量数据缺失,核心问题不是Cookie本身过期,是代码逻辑存在3个硬伤:
    • 虽然初始化了requests.Session()对象,但全程使用裸requests.request发请求,会话自动维护Cookie的能力完全没生效,所有Cookie都是硬编码的固定值
    • 硬编码的_pxvid是Walmart所用PerimeterX反爬的风控标识,和IP、TLS指纹强绑定,切换代理IP后旧标识直接失效,会被风控拦截返回空数据
    • 请求头中x-o-correlation-id、traceparent、wm_page_url、referer均为固定值,和当前请求的商品、会话不匹配,额外触发风控校验;且代码中proxies写死为None,代理配置实际未生效
  • 附原始问题代码:
def main(url_list, zip_code_list, ip_list, _now, save_dict, num, csv_list, utc_tz):
    _dict = {}
    debug = False
    s = requests.Session()

    output_json_file = f'backup/{num}_' + _now.strftime("%Y%m%d_%H%M.json")
    output_csv_file = f'backup/{num}_' + _now.strftime("%Y%m%d_%H%M.csv")

    flag = True

    if debug:
        url_list = [
            'https://www.target.com/p/claritin-24-hour-non-drowsy-allergy-relief-tablets-loratadine/-/A-80354268?preselect=14351285#lnk=sametab',
            'https://www.target.com/p/genexa-dextromethorphan-kids-39-cough-and-chest-congestion-suppressant-4-fl-oz/-/A-80130848#lnk=sametab'
        ]
        zip_code_list = [
            10005,
        ]
    i = 0
    for _url in url_list:
        for zip_code in zip_code_list:
            # proxy service
            proxies = {"http": None, "https": None}

            i += 1
            _dict[i] = {}
            start_time = time.perf_counter()

            try:
                item = _url.split("/")[-1]
                url_type = 1
                page_num = item

                if '?' in item:
                    url_type = 3
                    item2 = item.split("?")
                    page_num = item2[0]
                else:
                    pass
            except Exception as e:
                end_time = time.perf_counter()
                continue

            zip_code_url = "https://www.walmart.com/orchestra/home/graphql"

            payload = json.dumps({
                "query": "......",
                "variables": {
                    "input": {
                        "postalCode": str(zip_code),
                        "accessTypes": [
                            "PICKUP_INSTORE",
                            "PICKUP_CURBSIDE",
                            "PICKUP_SPOKE",
                            "PICKUP_POPUP"
                        ],
                        "nodeTypes": [
                            "STORE",
                            "PICKUP_SPOKE",
                            "PICKUP_POPUP"
                        ],
                        "latitude": None,
                        "longitude": None,
                        "radius": None
                    },
                    "checkItemAvailability": False,
                    "checkWeeklyReservation": False,
                    "enableStoreSelectorMarketplacePickup": False
                }
            })
            headers = {
                'authority': 'www.walmart.com',
                'pragma': 'no-cache',
                'cache-control': 'no-cache',
                'x-o-segment': 'oaoh',
                'x-o-correlation-id': 'Tt33HoVZ_Pqtlie1ABII1nfekFaSEtbRQPSc',
                'device_profile_ref_id': '-f6R8qf8Vd3gwky1UOzoEwW_XoTeRKqppMfK',
                'x-latency-trace': '1',
                'wm_mp': 'true',
                'wm_page_url': 'https://www.walmart.com/ip/Allegra-Adult-24HR-Gelcaps-24-Ct-180-mg-Allergy-Relief/43819800',
                'x-o-platform-version': 'main-1.2.0-3a465c',
                'x-o-gql-query': 'query nearByNodes',
                'x-o-bu': 'WALMART-US',
                'x-apollo-operation-name': 'nearByNodes',
                'traceparent': 'Tt33HoVZ_Pqtlie1ABII1nfekFaSEtbRQPSc',
                'x-o-mart': 'B2C',
                'user-agent': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/99.0.4844.82 Safari/537.36',
                'x-o-platform': 'rweb',
                'content-type': 'application/json',
                'accept': 'application/json',
                'x-enable-server-timing': '1',
                'x-o-ccm': 'server',
                'wm_qos.correlation_id': 'Tt33HoVZ_Pqtlie1ABII1nfekFaSEtbRQPSc',
                'origin': 'https://www.walmart.com',
                'sec-fetch-site': 'same-origin',
                'sec-fetch-mode': 'cors',
                'sec-fetch-dest': 'empty',
                'referer': 'https://www.walmart.com/ip/Allegra-Adult-24HR-Gelcaps-24-Ct-180-mg-Allergy-Relief/43819800',
                'accept-language': 'en-US,en;q=0.9,zh-CN;q=0.8,zh;q=0.7',
                'cookie': '_pxvid=10811607-f238-11ec-a720-4e756c594d76; ACID=2263b9c6-4e5a-44ce-a9da-05e028a9b8c7; hasACID=true'
            }

            try:
                response = requests.request("POST", zip_code_url, headers=headers, data=payload, proxies=proxies,timeout=10)
                content = response.json()
            except Exception as e:
                end_time = time.perf_counter()
                continue

            try:
                store_id = content['data']['nearByNodes']['nodes'][0]['id']
            except Exception as e:
                end_time = time.perf_counter()
                continue

            url2 = "https://www.walmart.com/orchestra/home/graphql/ip/"+page_num
            payload2 = json.dumps({
                "query": "...",
                "variables": {
                    "channel": "WWW",
                    "pageType": "ItemPageGlobal",
                    "tenant": "WM_GLASS",
                    "version": "v1",
                    "itemId": str(page_num),
                    "layout": ["itemDesktop"],
                    "fetchBuyBoxAd": True,
                    "fetchSkyline": True,
                    "fetchIdml": True,
                    "fetchReviews": True,
                    "fetchFitment": True,
                    "fetchSEO": True,
                    "fetchP13N": True,
                    "fetchAffirm": True,
                    "fetchMarquee": True,
                    "fetchSpCarousel": True,
                    "fetchBrandBox": True,
                    "fetchDiscounts": False,
                    "enableItemIbotta": True
                }
            })
            headers2 = {
                'authority': 'www.walmart.com',
                'pragma': 'no-cache',
                'cache-control': 'no-cache',
                'x-o-correlation-id': 'EtyZZHcTSBCiVMxl4i44a6h_Gzqf1F-MkpXd',
                'x-o-item-id': str(page_num),
                'user-agent': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/99.0.4844.82 Safari/537.36',
                'content-type': 'application/json',
                'accept': 'application/json',
                'origin': 'https://www.walmart.com',
                'referer': 'https://www.walmart.com/ip/Equate-Maximum-Strength-Severe-Allergy-Plus-Sinus-Headache-Caplets-20-Count/14053317',
                'accept-language': 'en-US,en;q=0.9',
                'cookie': '_pxvid=10811607-f238-11ec-a720-4e756c594d76; ACID=2263b9c6-4e5a-44ce-a9da-05e028a9b8c7; hasACID=true'
            }
            try:
                response2 = requests.request("POST", url2, headers=headers2, data=payload2, proxies=proxies,timeout=10)
                content2 = response2.json()
            except Exception as e:
                end_time = time.perf_counter()
                continue

            _now2 = datetime.datetime.now(tz=utc_tz)

            try:
                product_price = content2['data']['product']['priceInfo']['currentPrice']['price']
            except Exception as e:
                product_price = None

            try:
                product_gtin = content2['data']['product']['upc']
            except Exception as e:
                product_gtin = None
            try:
                pruduct_time = _now2.strftime("%Y-%m-%d %H:%M:%S")
                _dict[i] = {
                    'page_num': page_num,
                    'product_gtin': product_gtin,
                    'store_id': store_id,
                    'product_price': product_price,
                    'zip_code': zip_code,
                    'pruduct_time': pruduct_time
                }

                with open(output_csv_file, "a", encoding='utf-8') as fw2:
                    new_line = f'{page_num},{product_gtin},{zip_code},{store_id},{product_price},{pruduct_time}\n'
                    fw2.write(new_line)
                    if output_csv_file not in csv_list:
                        csv_list.append(output_csv_file)
                    end_time = time.perf_counter()
            except Exception as e:
                end_time = time.perf_counter()

    with open(output_json_file, "w", encoding='utf-8') as fw3:
        fw3.write(json.dumps(_dict))
    save_dict[num] = _dict
    return _dict


if __name__ == '__main__':
    if not os.path.exists('output'):
        os.makedirs('output')
    if not os.path.exists('backup'):
        os.makedirs('backup')
    utc_tz = pytz.timezone('UTC')
    _now = datetime.datetime.now(tz=utc_tz)

    ip_list = ['45.142.28.83:8094','45.137.60.112:6640']
    url_list = [
        'https://www.walmart.com/ip/Allegra-Adult-24HR-Gelcaps-24-Ct-180-mg-Allergy-Relief/43819800',
        'https://www.walmart.com/ip/Zyrtec-24-Hour-Allergy-Relief-Tablets-with-10-mg-Cetirizine-HCl-90-ct/224749468?athbdg=L1600',
        'https://www.walmart.com/ip/Equate-Maximum-Strength-Severe-Allergy-Plus-Sinus-Headache-Caplets-20-Count/14053317'
    ]
    zip_code_list = [10003,48104]

    start_time = time.perf_counter()
    save_dict = {}

    thread_num = 1
    one_thread_url_num = 3
    pool = threadpool.ThreadPool(thread_num)

    param_list = []
    csv_list = []
    for i in range(thread_num):
        save_dict[i + 1] = {}
        start_url_num = i * one_thread_url_num
        end_url_num = start_url_num + one_thread_url_num
        if end_url_num >= len(url_list):
            param_list.append(([url_list[start_url_num:], zip_code_list, ip_list, _now, save_dict, i + 1, csv_list, utc_tz], None))
        else:
            param_list.append(([url_list[start_url_num:end_url_num], zip_code_list, ip_list, _now, save_dict, i + 1, csv_list, utc_tz], None))
    tasks = threadpool.makeRequests(main, param_list)
    [pool.putRequest(task) for task in tasks]
    pool.wait()

修复方案

1. 会话与Cookie自动维护

  • 每个代理IP绑定一个独立的requests.Session实例,禁止多个代理IP共用同一会话,也禁止同一会话中途切换IP
  • 新代理IP首次使用时,先发起Walmart首页GET请求,让Session自动接收服务端Set-Cookie返回的初始合法Cookie(包含_pxvid、ACID等必填字段),不要手动硬编码Cookie值
  • 后续所有业务请求都通过该Session实例发起,不要手动覆盖Cookie字段,Session会自动更新过期的Cookie值

2. 动态请求头修正

  • 删除所有硬编码的固定x-o-correlation-id、traceparent、wm_qos.correlation-id值,每次请求生成32位随机十六进制字符串匹配原站格式填充
  • wm_page_url、referer字段动态替换为当前爬取的商品页URL,不要所有请求复用同一个测试商品链接
  • 删除请求头中zh-CN相关的语言配置,统一使用en-US,en;q=0.9,避免被识别为非美国本地用户
  • 每个Session绑定一个独立的、版本匹配的Chrome User-Agent,不要全量请求复用同一个UA

3. 风控与频率控制

  • 单代理IP请求间隔控制在3-5秒,加入1-2秒随机等待,不要短时间高频请求接口触发Cookie强制失效
  • 优先使用美国原生住宅代理,机房代理被Walmart拦截概率超过90%,即使Cookie合法也无法返回正常数据
  • 如果请求返回PerimeterX校验页面,暂停1-3秒后重试,直到Session自动更新完校验通过的Cookie再继续业务请求

4. 核心逻辑修正示例

import uuid
import random
import requests
from time import sleep

# 每个代理IP初始化独立会话,自动获取合法Cookie
def init_session(proxy_ip):
    session = requests.Session()
    # 绑定代理
    session.proxies = {
        "http": f"http://{proxy_ip}",
        "https": f"http://{proxy
相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 03:49:12