Scrapy爬虫运行后无数据抓取,含3条DEBUG与1条ERROR求助
Hey there, let's walk through fixing your Scrapy spider step by step—there are a few key issues in your code and settings that are causing the spider to fail silently or throw errors.
1. First: Fix the Robots Protocol Block
Looking at your settings log, ROBOTSTXT_OBEY = True is enabled. Best Buy's robots.txt file almost certainly blocks scrapers from accessing product pages, which means Scrapy is getting stopped before it even can request data.
Update your settings.py:
ROBOTSTXT_OBEY = False
Also, add a download delay to avoid getting blocked by the site's anti-scraping measures:
DOWNLOAD_DELAY = 2 # Wait 2 seconds between requests
2. Fix Broken Selectors (The Core Issue)
Your spider's parse methods use incorrect selectors that can't find elements on Best Buy's pages. Let's rewrite these to target the right elements:
Rewrite the parse method (homepage handling)
Your original code tried to target a non-existent "item" link—instead, we'll directly grab product links from the list and handle pagination properly:
def parse(self, response): # Grab all product links from the list page product_links = response.css(".sku-item a.sku-header::attr(href)").getall() for link in product_links: full_link = response.urljoin(link) # Convert relative URL to absolute yield scrapy.Request(full_link, callback=self.parse_covers) # Handle pagination (next page) next_page = response.css("a.pagination-btn.next::attr(href)").get() if next_page: full_next_link = response.urljoin(next_page) yield scrapy.Request(full_next_link, callback=self.parse)
Fix the parse_covers method (product detail page handling)
Your original selectors had syntax errors (missing . for classes) and targeted the wrong elements. Here's the corrected version with safety checks to avoid errors:
def parse_covers(self, response): # Grab product image URL image_url = response.css(".primary-image img::attr(src)").get(default="") # Grab product details with safe defaults and text extraction name = response.css(".sku-title::text").get(default="").strip() price = response.css(".priceView-purchase-price::text").get(default="").strip() model = response.css(".sku-value::text").get(default="").strip() sku_raw = response.css(".sku-id::text").get(default="").strip() sku = sku_raw[:-2] if sku_raw else "" # Only slice if we have a value # Only yield the item if we have at least some valid data if name or price: yield Refrigerator( name=name, price=price, model=model, sku=sku, file_urls=[image_url] if image_url else [] )
3. Fix File Download Pipeline
Your custom pipeline might be causing issues if it's not properly set up. For file downloads, use Scrapy's built-in FilesPipeline instead (unless you have custom logic you need to keep):
Update settings.py:
ITEM_PIPELINES = { 'scrapy.pipelines.files.FilesPipeline': 1, # Built-in file download pipeline # Keep your custom pipeline only if it adds unique functionality # 'refrigeratorspider.pipelines.RefrigeratorspiderPipeline': 300, } FILES_STORE = "/Users/Berkeley/refrigeratorspider/refrigeratorspider/output"
Make sure the FILES_STORE directory exists before running the spider—create it manually if needed.
4. Test the Spider
Run your spider again with the command:
scrapy crawl pyimagesearch-refrigerator-spider -o output.json
You should now see the spider requesting product pages, extracting data, and saving items to output.json, plus downloading images to your specified directory.
内容的提问来源于stack exchange,提问作者Berkeley Fondren

