Scrapy爬虫可爬取页面但无法提取Item问题求助
Hey there, let's figure out why your Scrapy spider is crawling pages but failing to extract items. Looking at your code snippet, there are a few critical issues that are almost certainly causing this problem:
1. rules is a dictionary instead of a list
Scrapy requires the rules attribute to be a list of Rule objects, but you’ve used curly braces {} which creates a dictionary. This breaks the crawler’s link-following logic—your parse_item callback might not even be running at all! Fix this by switching to square brackets:
rules = [ Rule(LinkExtractor(allow=r'.*/consultation/\d+'), callback="parse_item", follow=True), ]
2. Typo in the response variable
In your parse_item method, you wrote Selector(respons...—missing the final 'e' in response. That’s a syntax error that will crash the method before it can extract any data. Correct it to Selector(response...) (though, as we’ll see next, you might not even need this line).
3. No proper Item instantiation or return logic
Your code initializes items = [] but there’s no code to create a QuestionItem instance, populate its fields, add it to the list, or return the items. Even if your selector logic was perfect, you’re not actually outputting any items. Here’s how to fix this section:
def parse_item(self, response): # Create an instance of your QuestionItem item = QuestionItem() # Replace these with your actual field selectors (match the target page's HTML) item['title'] = response.css('h1::text').get() item['content'] = response.css('.consultation-content::text').getall() # Yield the item—Scrapy handles collecting and processing items automatically yield item
4. Redundant Selector creation
You don’t need to explicitly create a Selector from response—Scrapy’s response object already has built-in css() and xpath() methods that work directly with the page content. Using these is cleaner and more efficient than making a new Selector instance.
Corrected Full Spider Code
Here’s what your spider should look like after fixing all these issues:
from scrapy.spiders import CrawlSpider, Rule from scrapy.linkextractors import LinkExtractor from Diplom.items import QuestionItem class ConsultSpider(CrawlSpider): name = "consultation" allowed_domains = ['health.mail.ru'] start_urls = ['https://health.mail.ru/consultation/1579497'] rules = [ Rule(LinkExtractor(allow=r'.*/consultation/\d+'), callback="parse_item", follow=True), ] def parse_item(self, response): item = QuestionItem() # Update these selectors to match the actual elements on health.mail.ru item['question_title'] = response.xpath('//h1/text()').get() item['question_body'] = response.xpath('//div[contains(@class, "question-text")]/text()').getall() yield item
Don’t Forget Your Item Definition!
Double-check that your QuestionItem in Diplom/items.py is properly defined with the fields you want to extract. For example:
import scrapy class QuestionItem(scrapy.Item): question_title = scrapy.Field() question_body = scrapy.Field() # Add any other fields you need here
After making these fixes, run your spider again. It should now crawl pages and extract items as expected.
内容的提问来源于stack exchange,提问作者Konstantin Lysyy

