You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Requests爬取Coach Outlet预发布环境API时持续遭遇403访问被拒错误

使用Python Requests爬取Coach Outlet预发布环境API时持续遭遇403访问被拒错误

看起来你遇到的问题核心是两个Staging站点的反爬策略差异,再加上代码里的一个URL替换Bug,导致Outlet站点的请求要么被拦截,要么请求地址完全无效。我们一步步拆解解决:

一、先修复代码里的URL替换Bug

你看输出里最后一个请求的URL变成了 https://staging1.coachoutlet.com/api/shop/api/shop-by,这明显是错的!原因是你的代码里用了 full_url.replace('/shop', '/api/shop'),这个方法默认会替换所有匹配的/shop子串。比如原URL是/shop/shop-by,会把两个/shop都替换成/api/shop,导致地址完全无效。

解决方法:只替换第一个出现的/shop,把替换代码改成:

full_url = full_url.replace('/shop', '/api/shop', 1)

这样就只会把URL开头的/shop换成/api/shop,不会影响后面路径里的shop字符串。

二、Outlet Staging站点的403拦截原因及解决思路

日本Coach Staging站点的auth-bypass=true Cookie对Outlet站点无效,而且Outlet的反爬策略更严格,需要调整请求逻辑:

1. 不要硬写Cookie,先通过Session获取站点会话Cookie

Outlet的Staging站点可能需要先访问主页,获取有效会话Cookie后再请求API,而不是直接用针对日本站点的auth-bypass=true。修改你的Session初始化部分:

session = requests.Session()
# 先访问Outlet主页,获取会话Cookie
homepage_response = session.get('https://staging1.coachoutlet.com/', headers=headers, verify=False)
# 可以打印Set-Cookie看看站点返回了哪些Cookie
# print(homepage_response.cookies)

这样Session会自动保存主页返回的Cookie,后续请求API时会带上。

2. 更新User-Agent为现代浏览器版本

你当前用的是Chrome 58的UA,太旧了,很多反爬系统会直接拦截这种过时的UA。换成一个最新的Chrome UA,比如:

'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/129.0.0.0 Safari/537.36'

3. 补充AJAX请求常见的必要Headers

很多API会检查X-Requested-With头来确认是AJAX请求,在你的headers里加上:

'X-Requested-With': 'XMLHttpRequest'

4. 验证API端点是否正确

你的URL转换逻辑是基于日本站点的规则(/shop转/api/shop),但Outlet站点的API路径可能不一样!建议你手动打开Outlet Staging的页面(比如https://staging1.coachoutlet.com/shop/black-friday-deals/view-all),打开浏览器开发者工具的「网络」标签,找到加载产品的XHR请求,确认实际的API地址是什么,再调整你的URL转换逻辑。

三、修改后的完整代码示例

结合上面的修复,调整后的代码大概是这样:

import requests
import warnings
from time import sleep

# Suppress SSL warnings
warnings.simplefilter('ignore', requests.packages.urllib3.exceptions.InsecureRequestWarning)

# List of URLs to iterate over
urls = [
    'https://staging1.coachoutlet.com/shop/black-friday-deals/view-all', 
    'https://staging1.coachoutlet.com/shop/women/view-all', 
    'https://staging1.coachoutlet.com/shop/men/view-all', 
    'https://staging1.coachoutlet.com/shop/bags/view-all', 
    'https://staging1.coachoutlet.com/shop/gifts/view-all', 
    'https://staging1.coachoutlet.com/shop/shop-by'
]

# Full header simulation of a real browser session
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/129.0.0.0 Safari/537.36',
    'Accept': 'application/json',
    'Connection': 'keep-alive',
    'Accept-Language': 'en-US,en;q=0.9',
    'Accept-Encoding': 'gzip, deflate, br',
    'Referer': 'https://staging1.coachoutlet.com/shop',
    'Upgrade-Insecure-Requests': '1',
    'TE': 'Trailers',
    'X-Requested-With': 'XMLHttpRequest'  # 新增AJAX标识
}

# Start a session to handle cookies across requests
session = requests.Session()
# 先访问主页获取有效会话Cookie
try:
    homepage_resp = session.get('https://staging1.coachoutlet.com/', headers=headers, verify=False)
    print(f"Homepage response code: {homepage_resp.status_code}")
except Exception as e:
    print(f"Failed to access homepage: {e}")

# Iterate over each URL in the list
for url in urls:
    count_of_items = 1  # Reset item count for each URL
    page = 1  # Starting page for pagination
    if '/shop' not in url:
        print(f"No products available in this PLP: {url}")
        continue
    try:
        while True:
            ct = count_of_items + 15  # Adjust count to match pagination
            if page == 1:
                full_url = f"{url}"
            else:
                full_url = f"{url}?page={page}"
            if 'api/shop' not in full_url:                     
                # 只替换第一个出现的/shop,避免重复替换
                full_url = full_url.replace('/shop', '/api/shop', 1)
            print(f"Fetching: {full_url}")  
            response = session.get(full_url, headers=headers, verify=False)
            if response.status_code == 403:
                print(f"Access denied for URL {full_url}. Trying next page...")
                # 可以尝试增加延迟,避免被频率限制
                sleep(2)
                page +=1
                if page >3: # 最多试3页,避免无限循环
                    print(f"Max retries reached for {full_url}")
                    break
                continue
            print(f"Response Status Code: {response.status_code}")
            if response.status_code != 200:
                print(f"Failed to retrieve data from {full_url}")
                break
            products = response.json().get('pageData', {}).get('products', [])       
            if not products:
                print(f"No products available in this PLP: {full_url}")
                break
            pro_count = response.json()['pageData'].get('total', 0)
            print(f"Total products found: {pro_count}")
            break
    except Exception as e:
        print(f"Exception raised for URL {url}: {e}")
        continue

额外提示

如果还是403,建议你用浏览器的「复制为cURL」功能,把浏览器里成功请求API的cURL命令复制出来,然后用requests.utils.dict_from_cookiejar()或者在线工具转换成Requests的headers和Cookie,对比你的代码里的请求参数,看看缺了什么(比如CSRF Token、特殊的Cookie字段)。

备注:内容来源于stack exchange,提问作者Annie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 15:48:05