You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中Requests请求不同URL返回重复数据及访问拒绝问题求助

爬虫问题排查与解决方案

问题概述

现有待爬取数据集,爬虫可运行但存在两个核心问题:

  1. 不同URL返回相同数据:网页端查看不同URL响应内容不同且均返回200状态码,每个URL对应唯一dictionary_id,示例如下:
爬取的商品名称urldictionary_id
Hansaplast aqua protect 6shttps://www.blibli.com/p/hansaplast-aqua-protect-6s/ps--RAM-70107-07546?ds=RAM-70107-07546-00001&source=MERCHANT_PAGE&sid=7cb21b003f4516bb&cnc=false&pickupPointCode=PP-3239816&pid=RAM-70107-07546418
Hansaplast aqua protect 6shttps://www.blibli.com/p/hansaplast-aqua-protect-6s/ps--RAM-70107-07546?ds=RAM-70107-07546-00001&source=MERCHANT_PAGE&sid=7cb21b003f4516bb&cnc=false&pickupPointCode=PP-3239816&pid=RAM-70107-07546181
  1. 间歇性Access Denied错误:爬取若干URL后返回Access Denied错误且无数据,但重新运行又返回200状态码,需明确该错误对返回数据的影响。

现有代码

from random import randint
import requests

rancherrors=pd.DataFrame()
ranchdf=pd.DataFrame()
for id, input,url in zip(ranch['product_id'],ranch['urls'],ranch['urls2']):
    
    headers={
        "user-agent" : f"{UserAgent().random}",
        'referer':'https://www.blibli.com/'
        }
    response =requests.get(url,headers=headers,verify=True)
    #catches error
    ####################
    if response.status_code != 200:
        datum={
            'id':id,
            'url':url,
        'date_key':today
        }
        rancherrors=rancherrors.append(pd.DataFrame([datum]))
        print(f'{url} error')

        sleep(randint(5,15))

    else:
    #runs scraper
    ################################
        try:
            price=str(response.json()['data']['price']['listed']).replace(".0","")
            discount=str(response.json()['data']['price']['totalDiscount'])
        except:
            price="0"
            discount="0"
        try:
            unit=str(response.json()['data']['uniqueSellingPoint']).replace("• ","")
        except:
            unit=""

        datranch={
            'product_name':response.json()['data']['name'],
            'normal_price':price,
            'discount':discount,
            'competitor_id':response.json()['data']['ean'],
            'url':input,
            'unit':unit,
            'product_id':id,
            'date_key':today,
            'web':'ranch market'
            }
        ranchdf=ranchdf.append(pd.DataFrame([datranch]))

问题分析与解决方案

问题1:不同URL返回相同数据

原因推测

  • 目标网站可能基于会话标识(如sid参数)或请求指纹做缓存,相同sid导致返回相同缓存内容;
  • 代码未关联dictionary_id逻辑,实际请求的是同一资源入口;
  • 部分URL参数(如pickupPointCode)未生效,路由到同一商品数据。

解决步骤

  1. 清理冗余URL参数:移除固定sid参数,让请求自动生成新会话;
  2. 关联dictionary_id:抓包确认网页端请求是否携带该标识,通过headers或params传入请求;
  3. 禁用请求缓存:在requests.get中添加headers={'Cache-Control': 'no-cache'}强制刷新;
  4. 验证响应唯一性:打印response.json()['data']中的唯一标识(如dictionary_id),确认返回数据与目标匹配。

问题2:间歇性Access Denied错误

影响说明

  • 出现该错误时,请求未获取有效数据,对应URL会被记录到rancherrors,不会写入ranchdf;
  • 重新运行恢复200状态码属于网站临时反爬限制(如请求频率过高、会话过期),不影响已爬取数据,但会导致当前批次数据缺失。

解决优化

  1. 添加重试机制:使用HTTPAdapter设置自动重试,避免单次错误丢失数据:
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

session = requests.Session()
retry = Retry(total=3, backoff_factor=1, status_forcelist=[403, 500, 502, 503, 504])
adapter = HTTPAdapter(max_retries=retry)
session.mount('https://', adapter)
session.mount('http://', adapter)

# 后续请求改用session.get(url, headers=headers, verify=True)
  1. 优化请求间隔:改用指数退避间隔,降低触发频率限制的概率:
import time
from random import uniform

# 替代原sleep逻辑,retry_count为当前重试次数
sleep_time = uniform(2, 5) * (2 ** retry_count)
time.sleep(sleep_time)
  1. 完善错误处理:捕获403状态码时更换UA,必要时切换代理:
if response.status_code == 403:
    print(f"Access Denied for {url}, retrying with new UA...")
    headers["user-agent"] = UserAgent().random
    # 可选:切换代理
    # proxies = {"https": "http://your-proxy:port"}
    # response = session.get(url, headers=headers, verify=True, proxies=proxies)
  1. 校验数据完整性:每次运行后对比ranchdf的product_id与原数据集,从rancherrors提取遗漏URL重新爬取。

内容的提问来源于stack exchange,提问作者Hal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 19:20:23