You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用httpx和requests爬取Blibli商品API超时问题求助

爬取Blibli商品API超时/无响应问题解决方案

我来帮你分析下这个问题——Blibli的后端API大概率做了反爬验证,仅靠随机User-Agent肯定不够,下面是一步步的修复方案:

1. 补全关键请求头

电商API通常会校验多个请求头来判断是否是合法的浏览器请求,你需要在现有基础上添加这些必要的头:

  • Accept: 告诉服务器你接受JSON格式的响应
  • Accept-Encoding: 声明支持gzip压缩(响应头显示服务器用了gzip,请求里必须匹配)
  • Referer: 模拟从商品详情页发起的请求,值直接填你要爬的商品页面URL
  • Origin: 声明请求来源是Blibli官网

修改后的headers示例:

headers = {
    'User-Agent': USER_AGENT,
    'Accept': 'application/json, text/plain, */*',
    'Accept-Encoding': 'gzip, deflate, br',
    'Referer': 'https://www.blibli.com/p/facial-tissue-tisu-wajah-250-s-paseo/is--LO1-70001-00049-00003?seller_id=LO1-70001&sku_id=LO1-70001-00049-00001&sclid=7zuGEaS4hh5SowAA6tnfd5i2wKjR6e3p&sid=c5746ccfbb298d3b&pid=LO1-70001-00049-00001&pickupPointCode=PP-3227395',
    'Origin': 'https://www.blibli.com'
}

2. 使用会话保持Cookie

从响应头的set-cookie可以看到,Blibli会给合法请求分配会话Cookie,直接单次请求会丢失这些验证信息。建议用httpx.Client或requests.Session来自动维护会话:

httpx版本修复代码

from fake_useragent import UserAgent
import httpx

ua = UserAgent()
USER_AGENT = ua.random

headers = {
    'User-Agent': USER_AGENT,
    'Accept': 'application/json, text/plain, */*',
    'Accept-Encoding': 'gzip, deflate, br',
    'Referer': 'https://www.blibli.com/p/facial-tissue-tisu-wajah-250-s-paseo/is--LO1-70001-00049-00003?seller_id=LO1-70001&sku_id=LO1-70001-00049-00001&sclid=7zuGEaS4hh5SowAA6tnfd5i2wKjR6e3p&sid=c5746ccfbb298d3b&pid=LO1-70001-00049-00001&pickupPointCode=PP-3227395',
    'Origin': 'https://www.blibli.com'
}

api_url = "https://www.blibli.com/backend/product-detail/products/is--LO1-70001-00049-00003/_summary?pickupPointCode=PP-3227395"
product_url = "https://www.blibli.com/p/facial-tissue-tisu-wajah-250-s-paseo/is--LO1-70001-00049-00003"

# 用Client创建持久会话,设置30秒超时避免无限等待
with httpx.Client(headers=headers, timeout=30) as client:
    # 先访问商品页面获取初始会话Cookie
    client.get(product_url)
    # 再请求API
    response = client.get(api_url)
    # 检查响应状态码
    if response.status_code == 200:
        print(response.json())
    else:
        print(f"请求失败,状态码: {response.status_code}")

requests版本修复代码

from fake_useragent import UserAgent
import requests

ua = UserAgent()
USER_AGENT = ua.random

headers = {
    'User-Agent': USER_AGENT,
    'Accept': 'application/json, text/plain, */*',
    'Accept-Encoding': 'gzip, deflate, br',
    'Referer': 'https://www.blibli.com/p/facial-tissue-tisu-wajah-250-s-paseo/is--LO1-70001-00049-00003?seller_id=LO1-70001&sku_id=LO1-70001-00049-00001&sclid=7zuGEaS4hh5SowAA6tnfd5i2wKjR6e3p&sid=c5746ccfbb298d3b&pid=LO1-70001-00049-00001&pickupPointCode=PP-3227395',
    'Origin': 'https://www.blibli.com'
}

api_url = "https://www.blibli.com/backend/product-detail/products/is--LO1-70001-00049-00003/_summary?pickupPointCode=PP-3227395"
product_url = "https://www.blibli.com/p/facial-tissue-tisu-wajah-250-s-paseo/is--LO1-70001-00049-00003"

# 用Session维护会话
session = requests.Session()
session.headers.update(headers)

try:
    # 先访问商品页面获取Cookie
    session.get(product_url, timeout=10)
    # 请求API
    response = session.get(api_url, timeout=30)
    response.raise_for_status()  # 抛出HTTP错误
    print(response.json())
except requests.exceptions.RequestException as e:
    print(f"请求出错: {e}")

3. 额外排查点

如果以上方法还是不行,你可以:

  • 打开浏览器开发者工具(F12),找到真实的API请求,复制所有请求头到代码里,完全模拟浏览器请求
  • 检查是否有x-requested-with: XMLHttpRequest头(很多AJAX API会校验这个)
  • 注意Akamai的反爬机制(响应头里的akamai-grn),如果遇到更严格的验证,可能需要用playwright或selenium模拟真实浏览器行为

内容的提问来源于stack exchange,提问作者Hal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 18:45:46