You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫可爬取页面但无法提取Item问题求助

Troubleshooting Your Scrapy Spider's Item Extraction Issue

Hey there, let's figure out why your Scrapy spider is crawling pages but failing to extract items. Looking at your code snippet, there are a few critical issues that are almost certainly causing this problem:

1. rules is a dictionary instead of a list

Scrapy requires the rules attribute to be a list of Rule objects, but you’ve used curly braces {} which creates a dictionary. This breaks the crawler’s link-following logic—your parse_item callback might not even be running at all! Fix this by switching to square brackets:

rules = [
    Rule(LinkExtractor(allow=r'.*/consultation/\d+'), callback="parse_item", follow=True),
]

2. Typo in the response variable

In your parse_item method, you wrote Selector(respons...—missing the final 'e' in response. That’s a syntax error that will crash the method before it can extract any data. Correct it to Selector(response...) (though, as we’ll see next, you might not even need this line).

3. No proper Item instantiation or return logic

Your code initializes items = [] but there’s no code to create a QuestionItem instance, populate its fields, add it to the list, or return the items. Even if your selector logic was perfect, you’re not actually outputting any items. Here’s how to fix this section:

def parse_item(self, response):
    # Create an instance of your QuestionItem
    item = QuestionItem()
    
    # Replace these with your actual field selectors (match the target page's HTML)
    item['title'] = response.css('h1::text').get()
    item['content'] = response.css('.consultation-content::text').getall()
    
    # Yield the item—Scrapy handles collecting and processing items automatically
    yield item

4. Redundant Selector creation

You don’t need to explicitly create a Selector from response—Scrapy’s response object already has built-in css() and xpath() methods that work directly with the page content. Using these is cleaner and more efficient than making a new Selector instance.

Corrected Full Spider Code

Here’s what your spider should look like after fixing all these issues:

from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor
from Diplom.items import QuestionItem

class ConsultSpider(CrawlSpider):
    name = "consultation"
    allowed_domains = ['health.mail.ru']
    start_urls = ['https://health.mail.ru/consultation/1579497']
    
    rules = [
        Rule(LinkExtractor(allow=r'.*/consultation/\d+'), callback="parse_item", follow=True),
    ]

    def parse_item(self, response):
        item = QuestionItem()
        
        # Update these selectors to match the actual elements on health.mail.ru
        item['question_title'] = response.xpath('//h1/text()').get()
        item['question_body'] = response.xpath('//div[contains(@class, "question-text")]/text()').getall()
        
        yield item

Don’t Forget Your Item Definition!

Double-check that your QuestionItem in Diplom/items.py is properly defined with the fields you want to extract. For example:

import scrapy

class QuestionItem(scrapy.Item):
    question_title = scrapy.Field()
    question_body = scrapy.Field()
    # Add any other fields you need here

After making these fixes, run your spider again. It should now crawl pages and extract items as expected.

内容的提问来源于stack exchange,提问作者Konstantin Lysyy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 10:09:03