Scrapy嵌套URL爬取问题:CSV输出为空排查求助
Alright, let's work through your issues step by step, focusing first on the core logic problems as you requested:
1. Core Logic Errors in Your Crawl
There are several critical issues in your code that are causing empty CSV exports, even though the console says all pages are traversed:
- Undefined variable in
parse_item: You’re usinga.xpath(...)in this method, butais never defined here. This throws an uncaught exception, and Scrapy silently drops the item instead of yielding it. That’s the biggest reason your CSV is empty. - Redundant Item instantiation: In
parse_target_page, you createitems = KartonageItem()then pass it via meta, but inparse_itemyou re-instantiate it again before overwriting withresponse.meta['items']. This is harmless but unnecessary, and can lead to confusion. - Broken Link field assignment: In
parse_target_page, you already grab the product detail link withlink = a.xpath('a/@href'), but then assignitems['Link']using a different (and likely incorrect) XPath that returns a list (viaextract()). You should just uselink.get()here to store the single URL string. - Category traversal gap: If a category page has no subcategories (and directly shows products), your
parse_category_cartonsmethod won’t yield any requests—since it only looks forcat-thumbnailslinks. This means those product pages are completely skipped.
2. Do You Need to Switch to CrawlSpider?
No, you don’t. Your crawl has a clear, predictable hierarchy (start page → subcategory → sub-subcategory → product list → detail page), so using a regular scrapy.Spider with explicit callbacks is perfectly suitable. CrawlSpider is better for cases where you have lots of irregular link patterns to auto-match, but your current structure is straightforward enough to handle with manual callbacks.
3. Quick Fixes to Get Your Crawl Working
Here are the most impactful changes to test first:
- Fix the undefined variable in
parse_item:def parse_item(self,response): # Remove the redundant KartonageItem instantiation items = response.meta['items'] # Use response.xpath (or the table variable you defined) instead of undefined 'a' items['Weight'] = response.xpath('.//span[@class="staffelpreise-small"]/text()').get() items['Volume'] = response.xpath('.//td[@class="icon_contenct"][7]/text()').get() yield items - Close the category traversal gap:
def parse_category_cartons(self, response): # Grab all subcategory links directly subcat_links = response.xpath('//div[@class="cat-thumbnails"]/a/@href') if subcat_links: for link in subcat_links: yield response.follow(url=link.get(), callback=self.parse_target_page) else: # If no subcategories, this page is a product list—process it directly yield from self.parse_target_page(response) - Fix the Link field in
parse_target_page:items['Link'] = link.get() # Store single URL string instead of a list - Enable debug logging: Run your spider with
scrapy crawl kartons12 -s LOG_LEVEL=DEBUGto see exactly where XPaths fail or exceptions are thrown—this will help you debug remaining issues quickly.
内容的提问来源于stack exchange,提问作者kekw

