You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy设置NO_CALLBACK仍触发回调报错问题求助

问题描述

编写Scrapy爬虫时,在parse回调中生成设置了NO_CALLBACK的请求,按照文档说明NO_CALLBACK表示请求无需触发回调,但实际却触发了回调并抛出RuntimeError错误。尝试移除errback、cb_kwargs、meta及设置dont_filter=True均无效。

代码示例

from scrapy import Spider
from scrapy.http import TextResponse
from scrapy.http.request import NO_CALLBACK


class AppsSpider(Spider):
    name = "Apps"
    allowed_domains = ['steampowered.com', 'steamstatic.com']
    start_urls = ['https://store.steampowered.com/app/20']

    def parse(self, response: TextResponse):
        # preview media
        preview_section = response.css('#game_highlights')
        main_image_selector = '.game_header_image_full::attr("src")'
        preview_img_selector = '.highlight_screenshot a::attr("href")'
        preview_videos_selector = '.highlight_movie::attr("data-mp4-hd-source")'
        links = preview_section.css(
            ', '.join([main_image_selector, preview_img_selector, preview_videos_selector])).getall()

        # description section media
        description_section = response.css('#aboutThisGame')
        description_img_gif_selector = 'img::attr("src")'
        links += description_section.css(description_img_gif_selector).getall()

        yield from response.follow_all(links, callback=NO_CALLBACK)

报错回溯

2023-04-02 21:12:11 [scrapy.core.scraper] ERROR: Spider error processing <GET https://cdn.cloudflare.steamstatic.com/steam/apps/20/0000000164.1920x1080.jpg?t=1579634708> (referer: https://store.steampowered.com/app/20)
Traceback (most recent call last):
  File "C:\Users\leiver\miniconda3\envs\steam-scraping\Lib\site-packages\twisted\internet\defer.py", line 892, in _runCallbacks
    current.result = callback(  # type: ignore[misc]
                     ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "C:\Users\leiver\miniconda3\envs\steam-scraping\Lib\site-packages\scrapy\http\request\__init__.py", line 40, in NO_CALLBACK
    raise RuntimeError(
RuntimeError: The NO_CALLBACK callback has been called. This is a special callback value intended for requests whose callback is never meant to be called.
解决方案

将请求的callback=NO_CALLBACK替换为callback=None即可解决问题:

yield from response.follow_all(links, callback=None)

原因说明

NO_CALLBACK是Scrapy内部用于标记“回调逻辑不应被执行”的特殊标识,但它本身是一个会主动抛出RuntimeError的函数。当请求完成后,Scrapy仍会尝试调用该回调,从而触发报错。而设置callback=None才是Scrapy官方认可的、告知框架无需执行任何回调的正确方式。

如果你的需求是下载这些媒体资源,更推荐使用Scrapy内置的FilesPipeline或ImagesPipeline,它们能更高效地处理资源下载、去重、存储等逻辑。

内容的提问来源于stack exchange,提问作者limg21

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 21:17:07