Falabella部分URL提取JSON时出现403 Error问题排查求助
解决Falabella部分URL爬取返回403错误的问题
部分URL返回403、部分正常的情况,大概率是反爬机制对不同页面的检测强度不同,你的请求模拟程度还不足以绕过检测。以下是针对性的优化方案:
1. 补充完整请求头
仅设置User-Agent不足以绕过反爬,需要添加浏览器请求中常见的其他头信息,让请求更接近真实用户的浏览器行为:
session.headers.update({ 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36', 'Referer': 'https://www.falabella.com/falabella-cl/', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'es-CL,es;q=0.8,en-US;q=0.5,en;q=0.3', 'Upgrade-Insecure-Requests': '1' })
2. 先获取站点初始Cookies
部分页面会验证会话是否带有从首页跳转的Cookies,建议先让Session访问一次Falabella首页,再请求目标URL:
# 在调用爬取函数前先访问首页获取初始Cookies session.get('https://www.falabella.com/falabella-cl/')
3. 添加请求延迟
频繁请求会触发反爬阈值,在每次请求前加入随机延迟,模拟真实用户的浏览间隔:
import time import random # 在请求前添加1-3秒的随机延迟 time.sleep(random.uniform(1, 3)) response = session.get(url)
修改后的完整代码
import requests import json from bs4 import BeautifulSoup import time import random session = requests.Session() session.headers.update({ 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36', 'Referer': 'https://www.falabella.com/falabella-cl/', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'es-CL,es;q=0.8,en-US;q=0.5,en;q=0.3', 'Upgrade-Insecure-Requests': '1' }) # 预加载首页获取Cookies session.get('https://www.falabella.com/falabella-cl/') def extract_json_from_falabella(url): try: # 添加随机延迟 time.sleep(random.uniform(1, 3)) response = session.get(url) response.raise_for_status() soup = BeautifulSoup(response.content, 'html.parser') script_tag = soup.find('script', id='__NEXT_DATA__') if script_tag: json_text = script_tag.string.strip() data = json.loads(json_text) return data else: print("未找到id为'__NEXT_DATA__'的script标签。") return None except requests.exceptions.HTTPError as http_err: print(f"HTTP错误: {http_err}") return None except Exception as err: print(f"发生错误: {err}") return None url = "https://www.falabella.com/falabella-cl/category/cat7330051/Mujer?facetSelected=true&f.derived.variant.sellerId=FALABELLA%3A%3ASODIMAC&page=1" data = extract_json_from_falabella(url) if data: with open('falabella_data.json', 'w', encoding='utf-8') as json_file: json.dump(data, json_file, ensure_ascii=False, indent=4) print("数据已保存到'falabella_data.json'") else: print("无法提取JSON数据。")
额外提示
- 如果仍然出现403,可以尝试更换最新版本的
User-Agent,或者使用代理IP分散请求来源。 - 避免短时间内批量请求同一页面,尽量模拟真实用户的浏览节奏。
内容的提问来源于stack exchange,提问作者Marioloteitor
相关产品推荐
相关产品推荐

