You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy.Request未执行回调函数且start_requests报错求助

Scrapy爬虫问题:process_listing函数未执行且报错TypeError

问题现象

  • 期望Visual Studio控制台输出“HIT”,但process_listing函数从未执行
  • 运行命令 scrapy crawl foo -O foo.json 时出现错误:

start_requests = iter(self.spider.start_requests()) TypeError: 'NoneType' object is not iterable

爬虫代码

import json
import re
import os
import requests
import scrapy
import time
from scrapy.selector import Selector
from scrapy.http import HtmlResponse
import html2text

class FooSpider(scrapy.Spider):
    name = 'foo'
    start_urls = ['https://www.example.com/item.json?lang=en']

    def start_requests(self):
        r = requests.get(self.start_urls[0])
        cont = r.json()
        self.parse(cont)

    def parse(self, response):
        for o in response['objects']:
            if o.get('option') == "buy" and o.get('is_available'):   
                listing_url = "https://www.example.com/" + \
                     o.get('brand').lower().replace(' ','-') + "-" + \
                     o.get('model').lower() + "-"
                if o.get('make') is not None:
                    listing_url += o.get('make') + "-"
                listing_url += o.get('year').lower() 
                print(listing_url) #a valid url is printed here

                yield scrapy.Request(
                    url=response.urljoin(listing_url), 
                    callback=self.process_listing
                )

    
    def process_listing(self, response):
        #this function is never executed
        print('HIT')
        yield item

已尝试的方案

  • 使用 url=response.urljoin(listing_url)
  • 使用 url=listing_url

问题原因及修复

1. start_requests方法未返回可迭代对象

你的start_requests方法没有返回任何内容(默认返回None),Scrapy需要该方法返回可迭代的Request对象,因此会抛出NoneType不可迭代的错误。同时你直接调用self.parse(cont),并没有把parse方法中生成的Request传递给Scrapy调度器,导致后续的请求根本不会被处理,process_listing自然不会执行。

2. parse方法参数类型错误

你在start_requests中把json字典cont传给了parse,所以parse中的response参数是字典而非Scrapy的Response对象,此时调用response.urljoin()会抛出AttributeError(字典没有这个方法),只是你可能没注意到这个错误。

修复后的代码示例

方案一:改用Scrapy原生Request获取json(推荐,符合异步设计)

import scrapy
import html2text

class FooSpider(scrapy.Spider):
    name = 'foo'
    start_urls = ['https://www.example.com/item.json?lang=en']

    def start_requests(self):
        # 用Scrapy的Request异步获取json数据
        yield scrapy.Request(
            url=self.start_urls[0],
            callback=self.parse_json
        )

    def parse_json(self, response):
        # 解析json数据
        cont = response.json()
        # 把生成的请求传递给调度器
        yield from self.parse(cont)

    def parse(self, json_data):
        for o in json_data['objects']:
            if o.get('option') == "buy" and o.get('is_available'):   
                listing_url = "https://www.example.com/" + \
                     o.get('brand').lower().replace(' ','-') + "-" + \
                     o.get('model').lower() + "-"
                if o.get('make') is not None:
                    listing_url += o.get('make') + "-"
                listing_url += str(o.get('year')).lower() 
                print(listing_url)

                # 直接用完整的listing_url,不需要urljoin
                yield scrapy.Request(
                    url=listing_url, 
                    callback=self.process_listing
                )

    def process_listing(self, response):
        print('HIT')
        # 这里需要定义item并赋值,示例中补充基础结构
        item = {}
        item['url'] = response.url
        yield item

方案二:保留requests.get但修正start_requests返回值

如果你坚持要用requests同步获取数据,需要修改start_requests返回parse生成的迭代器:

def start_requests(self):
    r = requests.get(self.start_urls[0])
    cont = r.json()
    # 用yield from返回parse生成的所有Request
    yield from self.parse(cont)

同时要删除parse中的response.urljoin(),改用直接的listing_url,因为此时response是json字典,没有urljoin方法。

这样修改后,Scrapy就能正确调度parse生成的Request,process_listing函数会被执行,控制台也会输出“HIT”。

内容的提问来源于stack exchange,提问作者Adam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 05:16:37