You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Falabella部分URL提取JSON时出现403 Error问题排查求助

解决Falabella部分URL爬取返回403错误的问题

部分URL返回403、部分正常的情况,大概率是反爬机制对不同页面的检测强度不同,你的请求模拟程度还不足以绕过检测。以下是针对性的优化方案:

1. 补充完整请求头

仅设置User-Agent不足以绕过反爬,需要添加浏览器请求中常见的其他头信息,让请求更接近真实用户的浏览器行为:

session.headers.update({
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36',
    'Referer': 'https://www.falabella.com/falabella-cl/',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
    'Accept-Language': 'es-CL,es;q=0.8,en-US;q=0.5,en;q=0.3',
    'Upgrade-Insecure-Requests': '1'
})

2. 先获取站点初始Cookies

部分页面会验证会话是否带有从首页跳转的Cookies,建议先让Session访问一次Falabella首页,再请求目标URL:

# 在调用爬取函数前先访问首页获取初始Cookies
session.get('https://www.falabella.com/falabella-cl/')

3. 添加请求延迟

频繁请求会触发反爬阈值,在每次请求前加入随机延迟,模拟真实用户的浏览间隔:

import time
import random

# 在请求前添加1-3秒的随机延迟
time.sleep(random.uniform(1, 3))
response = session.get(url)

修改后的完整代码

import requests
import json
from bs4 import BeautifulSoup
import time
import random

session = requests.Session()
session.headers.update({
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36',
    'Referer': 'https://www.falabella.com/falabella-cl/',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
    'Accept-Language': 'es-CL,es;q=0.8,en-US;q=0.5,en;q=0.3',
    'Upgrade-Insecure-Requests': '1'
})

# 预加载首页获取Cookies
session.get('https://www.falabella.com/falabella-cl/')

def extract_json_from_falabella(url):
    try:
        # 添加随机延迟
        time.sleep(random.uniform(1, 3))
        response = session.get(url)
        response.raise_for_status()

        soup = BeautifulSoup(response.content, 'html.parser')
        script_tag = soup.find('script', id='__NEXT_DATA__')

        if script_tag:
            json_text = script_tag.string.strip()
            data = json.loads(json_text)
            return data
        else:
            print("未找到id为'__NEXT_DATA__'的script标签。")
            return None

    except requests.exceptions.HTTPError as http_err:
        print(f"HTTP错误: {http_err}")
        return None
    except Exception as err:
        print(f"发生错误: {err}")
        return None

url = "https://www.falabella.com/falabella-cl/category/cat7330051/Mujer?facetSelected=true&f.derived.variant.sellerId=FALABELLA%3A%3ASODIMAC&page=1"
data = extract_json_from_falabella(url)

if data:
    with open('falabella_data.json', 'w', encoding='utf-8') as json_file:
        json.dump(data, json_file, ensure_ascii=False, indent=4)
    print("数据已保存到'falabella_data.json'")
else:
    print("无法提取JSON数据。")

额外提示

  • 如果仍然出现403,可以尝试更换最新版本的User-Agent,或者使用代理IP分散请求来源。
  • 避免短时间内批量请求同一页面,尽量模拟真实用户的浏览节奏。

内容的提问来源于stack exchange,提问作者Marioloteitor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 03:03:26