You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy多回调函数失效求助:无法触发get_title回调

Fixing the Unfired get_title Callback in Your Scrapy Spider

Hey there, let's break down why your get_title callback isn't triggering, and fix it step by step.

The Core Issue

The main problem is that you're not yielding the Request object created by response.follow in your books_detail method. Scrapy's scheduler only processes requests that you explicitly yield—if you just assign the request to a variable (like title = response.follow(...)) and do nothing with it, the request is never sent, so get_title never runs.

Additional Problem: Data Association

On top of that, your current approach of collecting price/availability in books_detail and trying to get titles from get_title won't work as expected. Scrapy runs requests asynchronously, so the yas dictionary you yield will be sent before any get_title requests complete, and you won't be able to map titles to their corresponding prices/availability.

The Fix

We'll adjust two key parts:

  1. Yield the response.follow request to trigger get_title
  2. Pass the price/availability/category data from books_detail to get_title using Scrapy's meta parameter, so we can yield a complete, matched item in the end.

Here's the revised code:

class quotes(scrapy.Spider):
    name = 'quotes'
    start_urls = ['http://books.toscrape.com/?']
    
    def parse(self, response):
        all_links = response.css('.nav-list ul li')
        for links in all_links:
            link = links.css('a::attr(href)').get()
            yield response.follow(link, callback=self.books_detail)
    
    def books_detail(self, response):
        # Get the category once per page
        category = response.css('h1::text').get()
        all_divs = response.css('.col-lg-3')
        
        for div in all_divs:
            # Extract price and availability for this book
            price = div.css('.price_color::text').get()
            # Clean up availability text (strip whitespace)
            availability = div.css('.availability::text')[1].get().strip()
            # Get the book detail link
            book_link = div.css('.product_pod a::attr(href)').get()
            
            # Yield the request, passing data via meta
            yield response.follow(
                book_link,
                callback=self.get_title,
                meta={
                    'category': category,
                    'price': price,
                    'availability': availability
                }
            )
    
    def get_title(self, response):
        # Retrieve data from meta
        category = response.meta['category']
        price = response.meta['price']
        availability = response.meta['availability']
        # Extract the title
        title = response.css('h1::text').get()
        
        # Yield a complete item with all matched data
        yield {
            'title': title,
            'price': price,
            'availability': availability,
            'category': category
        }

Key Changes Explained

  • Yielding the Request: In books_detail, we now yield response.follow(...) instead of assigning it to a variable. This tells Scrapy to send the request and run the get_title callback when the response comes back.
  • Using meta: We pass the category, price, and availability data through the meta parameter of the request. This ensures that when get_title runs, it has access to the corresponding data for that specific book.
  • Cleaner Data Extraction: We use .get() instead of .extract() (since .get() returns a single string instead of a list, which is cleaner for single-value fields) and strip whitespace from the availability text.
  • Single Item Yield: Now each book's complete data is yielded as one item, with all fields correctly matched.

Quick Notes

  • Always remember: Scrapy only processes requests that you yield. If you create a Request but don't yield it, it's ignored.
  • When working with nested callbacks, use meta to pass data between them—this is the standard way to keep related data paired in asynchronous scraping.

内容的提问来源于stack exchange,提问作者ahsan mukhtar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 19:17:48