You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

将requests.post转为Scrapy.Request后回调函数未触发的问题求助

问题分析与解决方案

你的核心问题有两个:

  1. 未正确将Scrapy Request提交至引擎:_get_menu_item调用_get_super_item后,仅把生成器对象存入列表,未将生成的Request交给Scrapy调度处理。
  2. POST请求Body格式错误:_get_super_item直接将字典传给body参数,但Scrapy的body需要传入JSON字符串(和你使用requests时的json.dumps写法对应)。

另外还存在一个语法错误:generate_request函数定义未接收参数,但start_requests调用时传入了restaurant_url,会触发参数不匹配报错。


修复后的完整代码示例

import json
from scrapy import Request
from scrapy.loader import ItemLoader

def start_requests(self):
    for listing_item in self.get_listing_items():
        yield self.generate_request(listing_item.get('restaurant_url'))

# 修正参数接收逻辑
def generate_request(self, restaurant_url):
    self.url = restaurant_url  # 根据实际业务调整url赋值逻辑
    headers = {
        'content-type': 'application/json',
        'x-csrf-token': self.x_csrf_token,
    }
    payload = {'var_c': var_c}
    return Request(
        url=self.url,
        headers=headers,
        method='POST',
        body=json.dumps(payload),
        callback=self.parse_restaurant,
        priority=2
    )

def parse_restaurant(self, response):
    try:
        data = json.loads(response.body)
        restaurant = data['data']
    except Exception as e:
        self.logger.debug(f'Invalid response: {response.body}, error: {str(e)}')
        return
    
    # 直接yield Request,交给Scrapy引擎调度
    for request in self._get_menu_item(restaurant):
        yield request

# 改为生成器,直接yield Request
def _get_menu_item(self, restaurant):    
    for x_item in self.x_items:
        yield from self._get_super_item(x_item, restaurant, cookie)

def _get_super_item(self, selector, restaurant, cookie):
    # 修正body为JSON字符串
    yield scrapy.Request(
        url=self.url,
        method='POST',
        body=json.dumps({
            "var_a": "var_a",
            "var_b": "var_b",
        }),
        headers={
            'content-type': 'application/json',
            'cookie': cookie,
            'x-csrf-token': 'x'
        },
        callback=self._get_super_item_v2,
        # 添加errback排查请求失败原因
        errback=self._handle_request_error,
    )

def _get_super_item_v2(self, response):
    print('resp:', response.json())

# 新增错误处理函数,排查请求失败问题
def _handle_request_error(self, failure):
    self.logger.error(f'Request failed: {failure.request.url}, error: {str(failure.value)}')

关键修复点说明

  • Request调度逻辑:Scrapy仅会处理回调函数(如parse_restaurant)或start_requests中yield的Request,之前的代码仅将生成器存入列表,未触发调度。
  • Body格式修正:当请求头content-type设为application/json时,必须将Body转为JSON字符串,否则服务器无法解析请求,可能直接返回错误导致回调不触发。
  • errback辅助排查:添加errback可以捕获请求失败的具体原因(如网络错误、403/404状态码),快速定位回调未触发的问题。

内容的提问来源于stack exchange,提问作者Shivanshu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 10:35:33