Scrapy多回调函数失效求助:无法触发get_title回调
get_title Callback in Your Scrapy Spider Hey there, let's break down why your get_title callback isn't triggering, and fix it step by step.
The Core Issue
The main problem is that you're not yielding the Request object created by response.follow in your books_detail method. Scrapy's scheduler only processes requests that you explicitly yield—if you just assign the request to a variable (like title = response.follow(...)) and do nothing with it, the request is never sent, so get_title never runs.
Additional Problem: Data Association
On top of that, your current approach of collecting price/availability in books_detail and trying to get titles from get_title won't work as expected. Scrapy runs requests asynchronously, so the yas dictionary you yield will be sent before any get_title requests complete, and you won't be able to map titles to their corresponding prices/availability.
The Fix
We'll adjust two key parts:
- Yield the
response.followrequest to triggerget_title - Pass the price/availability/category data from
books_detailtoget_titleusing Scrapy'smetaparameter, so we can yield a complete, matched item in the end.
Here's the revised code:
class quotes(scrapy.Spider): name = 'quotes' start_urls = ['http://books.toscrape.com/?'] def parse(self, response): all_links = response.css('.nav-list ul li') for links in all_links: link = links.css('a::attr(href)').get() yield response.follow(link, callback=self.books_detail) def books_detail(self, response): # Get the category once per page category = response.css('h1::text').get() all_divs = response.css('.col-lg-3') for div in all_divs: # Extract price and availability for this book price = div.css('.price_color::text').get() # Clean up availability text (strip whitespace) availability = div.css('.availability::text')[1].get().strip() # Get the book detail link book_link = div.css('.product_pod a::attr(href)').get() # Yield the request, passing data via meta yield response.follow( book_link, callback=self.get_title, meta={ 'category': category, 'price': price, 'availability': availability } ) def get_title(self, response): # Retrieve data from meta category = response.meta['category'] price = response.meta['price'] availability = response.meta['availability'] # Extract the title title = response.css('h1::text').get() # Yield a complete item with all matched data yield { 'title': title, 'price': price, 'availability': availability, 'category': category }
Key Changes Explained
- Yielding the Request: In
books_detail, we nowyield response.follow(...)instead of assigning it to a variable. This tells Scrapy to send the request and run theget_titlecallback when the response comes back. - Using
meta: We pass the category, price, and availability data through themetaparameter of the request. This ensures that whenget_titleruns, it has access to the corresponding data for that specific book. - Cleaner Data Extraction: We use
.get()instead of.extract()(since.get()returns a single string instead of a list, which is cleaner for single-value fields) and strip whitespace from the availability text. - Single Item Yield: Now each book's complete data is yielded as one item, with all fields correctly matched.
Quick Notes
- Always remember: Scrapy only processes requests that you yield. If you create a Request but don't yield it, it's ignored.
- When working with nested callbacks, use
metato pass data between them—this is the standard way to keep related data paired in asynchronous scraping.
内容的提问来源于stack exchange,提问作者ahsan mukhtar

