You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫运行后无数据抓取,含3条DEBUG与1条ERROR求助

排查Scrapy爬虫无数据抓取问题的解决方案

Hey there, let's walk through fixing your Scrapy spider step by step—there are a few key issues in your code and settings that are causing the spider to fail silently or throw errors.


1. First: Fix the Robots Protocol Block

Looking at your settings log, ROBOTSTXT_OBEY = True is enabled. Best Buy's robots.txt file almost certainly blocks scrapers from accessing product pages, which means Scrapy is getting stopped before it even can request data.

Update your settings.py:

ROBOTSTXT_OBEY = False

Also, add a download delay to avoid getting blocked by the site's anti-scraping measures:

DOWNLOAD_DELAY = 2  # Wait 2 seconds between requests

2. Fix Broken Selectors (The Core Issue)

Your spider's parse methods use incorrect selectors that can't find elements on Best Buy's pages. Let's rewrite these to target the right elements:

Rewrite the parse method (homepage handling)

Your original code tried to target a non-existent "item" link—instead, we'll directly grab product links from the list and handle pagination properly:

def parse(self, response):
    # Grab all product links from the list page
    product_links = response.css(".sku-item a.sku-header::attr(href)").getall()
    for link in product_links:
        full_link = response.urljoin(link)  # Convert relative URL to absolute
        yield scrapy.Request(full_link, callback=self.parse_covers)
    
    # Handle pagination (next page)
    next_page = response.css("a.pagination-btn.next::attr(href)").get()
    if next_page:
        full_next_link = response.urljoin(next_page)
        yield scrapy.Request(full_next_link, callback=self.parse)

Fix the parse_covers method (product detail page handling)

Your original selectors had syntax errors (missing . for classes) and targeted the wrong elements. Here's the corrected version with safety checks to avoid errors:

def parse_covers(self, response):
    # Grab product image URL
    image_url = response.css(".primary-image img::attr(src)").get(default="")
    
    # Grab product details with safe defaults and text extraction
    name = response.css(".sku-title::text").get(default="").strip()
    price = response.css(".priceView-purchase-price::text").get(default="").strip()
    model = response.css(".sku-value::text").get(default="").strip()
    sku_raw = response.css(".sku-id::text").get(default="").strip()
    sku = sku_raw[:-2] if sku_raw else ""  # Only slice if we have a value
    
    # Only yield the item if we have at least some valid data
    if name or price:
        yield Refrigerator(
            name=name,
            price=price,
            model=model,
            sku=sku,
            file_urls=[image_url] if image_url else []
        )

3. Fix File Download Pipeline

Your custom pipeline might be causing issues if it's not properly set up. For file downloads, use Scrapy's built-in FilesPipeline instead (unless you have custom logic you need to keep):

Update settings.py:

ITEM_PIPELINES = {
    'scrapy.pipelines.files.FilesPipeline': 1,  # Built-in file download pipeline
    # Keep your custom pipeline only if it adds unique functionality
    # 'refrigeratorspider.pipelines.RefrigeratorspiderPipeline': 300,
}
FILES_STORE = "/Users/Berkeley/refrigeratorspider/refrigeratorspider/output"

Make sure the FILES_STORE directory exists before running the spider—create it manually if needed.


4. Test the Spider

Run your spider again with the command:

scrapy crawl pyimagesearch-refrigerator-spider -o output.json

You should now see the spider requesting product pages, extracting data, and saving items to output.json, plus downloading images to your specified directory.

内容的提问来源于stack exchange,提问作者Berkeley Fondren

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:12:35