使用Python Requests爬取Coach Outlet预发布环境API时持续遭遇403访问被拒错误
看起来你遇到的问题核心是两个Staging站点的反爬策略差异,再加上代码里的一个URL替换Bug,导致Outlet站点的请求要么被拦截,要么请求地址完全无效。我们一步步拆解解决:
一、先修复代码里的URL替换Bug
你看输出里最后一个请求的URL变成了 https://staging1.coachoutlet.com/api/shop/api/shop-by,这明显是错的!原因是你的代码里用了 full_url.replace('/shop', '/api/shop'),这个方法默认会替换所有匹配的/shop子串。比如原URL是/shop/shop-by,会把两个/shop都替换成/api/shop,导致地址完全无效。
解决方法:只替换第一个出现的/shop,把替换代码改成:
full_url = full_url.replace('/shop', '/api/shop', 1)
这样就只会把URL开头的/shop换成/api/shop,不会影响后面路径里的shop字符串。
二、Outlet Staging站点的403拦截原因及解决思路
日本Coach Staging站点的auth-bypass=true Cookie对Outlet站点无效,而且Outlet的反爬策略更严格,需要调整请求逻辑:
1. 不要硬写Cookie,先通过Session获取站点会话Cookie
Outlet的Staging站点可能需要先访问主页,获取有效会话Cookie后再请求API,而不是直接用针对日本站点的auth-bypass=true。修改你的Session初始化部分:
session = requests.Session() # 先访问Outlet主页,获取会话Cookie homepage_response = session.get('https://staging1.coachoutlet.com/', headers=headers, verify=False) # 可以打印Set-Cookie看看站点返回了哪些Cookie # print(homepage_response.cookies)
这样Session会自动保存主页返回的Cookie,后续请求API时会带上。
2. 更新User-Agent为现代浏览器版本
你当前用的是Chrome 58的UA,太旧了,很多反爬系统会直接拦截这种过时的UA。换成一个最新的Chrome UA,比如:
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/129.0.0.0 Safari/537.36'
3. 补充AJAX请求常见的必要Headers
很多API会检查X-Requested-With头来确认是AJAX请求,在你的headers里加上:
'X-Requested-With': 'XMLHttpRequest'
4. 验证API端点是否正确
你的URL转换逻辑是基于日本站点的规则(/shop转/api/shop),但Outlet站点的API路径可能不一样!建议你手动打开Outlet Staging的页面(比如https://staging1.coachoutlet.com/shop/black-friday-deals/view-all),打开浏览器开发者工具的「网络」标签,找到加载产品的XHR请求,确认实际的API地址是什么,再调整你的URL转换逻辑。
三、修改后的完整代码示例
结合上面的修复,调整后的代码大概是这样:
import requests import warnings from time import sleep # Suppress SSL warnings warnings.simplefilter('ignore', requests.packages.urllib3.exceptions.InsecureRequestWarning) # List of URLs to iterate over urls = [ 'https://staging1.coachoutlet.com/shop/black-friday-deals/view-all', 'https://staging1.coachoutlet.com/shop/women/view-all', 'https://staging1.coachoutlet.com/shop/men/view-all', 'https://staging1.coachoutlet.com/shop/bags/view-all', 'https://staging1.coachoutlet.com/shop/gifts/view-all', 'https://staging1.coachoutlet.com/shop/shop-by' ] # Full header simulation of a real browser session headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/129.0.0.0 Safari/537.36', 'Accept': 'application/json', 'Connection': 'keep-alive', 'Accept-Language': 'en-US,en;q=0.9', 'Accept-Encoding': 'gzip, deflate, br', 'Referer': 'https://staging1.coachoutlet.com/shop', 'Upgrade-Insecure-Requests': '1', 'TE': 'Trailers', 'X-Requested-With': 'XMLHttpRequest' # 新增AJAX标识 } # Start a session to handle cookies across requests session = requests.Session() # 先访问主页获取有效会话Cookie try: homepage_resp = session.get('https://staging1.coachoutlet.com/', headers=headers, verify=False) print(f"Homepage response code: {homepage_resp.status_code}") except Exception as e: print(f"Failed to access homepage: {e}") # Iterate over each URL in the list for url in urls: count_of_items = 1 # Reset item count for each URL page = 1 # Starting page for pagination if '/shop' not in url: print(f"No products available in this PLP: {url}") continue try: while True: ct = count_of_items + 15 # Adjust count to match pagination if page == 1: full_url = f"{url}" else: full_url = f"{url}?page={page}" if 'api/shop' not in full_url: # 只替换第一个出现的/shop,避免重复替换 full_url = full_url.replace('/shop', '/api/shop', 1) print(f"Fetching: {full_url}") response = session.get(full_url, headers=headers, verify=False) if response.status_code == 403: print(f"Access denied for URL {full_url}. Trying next page...") # 可以尝试增加延迟,避免被频率限制 sleep(2) page +=1 if page >3: # 最多试3页,避免无限循环 print(f"Max retries reached for {full_url}") break continue print(f"Response Status Code: {response.status_code}") if response.status_code != 200: print(f"Failed to retrieve data from {full_url}") break products = response.json().get('pageData', {}).get('products', []) if not products: print(f"No products available in this PLP: {full_url}") break pro_count = response.json()['pageData'].get('total', 0) print(f"Total products found: {pro_count}") break except Exception as e: print(f"Exception raised for URL {url}: {e}") continue
额外提示
如果还是403,建议你用浏览器的「复制为cURL」功能,把浏览器里成功请求API的cURL命令复制出来,然后用requests.utils.dict_from_cookiejar()或者在线工具转换成Requests的headers和Cookie,对比你的代码里的请求参数,看看缺了什么(比如CSRF Token、特殊的Cookie字段)。
备注:内容来源于stack exchange,提问作者Annie

