You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy嵌套URL爬取问题:CSV输出为空排查求助

Alright, let's work through your issues step by step, focusing first on the core logic problems as you requested:

1. Core Logic Errors in Your Crawl

There are several critical issues in your code that are causing empty CSV exports, even though the console says all pages are traversed:

  • Undefined variable in parse_item: You’re using a.xpath(...) in this method, but a is never defined here. This throws an uncaught exception, and Scrapy silently drops the item instead of yielding it. That’s the biggest reason your CSV is empty.
  • Redundant Item instantiation: In parse_target_page, you create items = KartonageItem() then pass it via meta, but in parse_item you re-instantiate it again before overwriting with response.meta['items']. This is harmless but unnecessary, and can lead to confusion.
  • Broken Link field assignment: In parse_target_page, you already grab the product detail link with link = a.xpath('a/@href'), but then assign items['Link'] using a different (and likely incorrect) XPath that returns a list (via extract()). You should just use link.get() here to store the single URL string.
  • Category traversal gap: If a category page has no subcategories (and directly shows products), your parse_category_cartons method won’t yield any requests—since it only looks for cat-thumbnails links. This means those product pages are completely skipped.

2. Do You Need to Switch to CrawlSpider?

No, you don’t. Your crawl has a clear, predictable hierarchy (start page → subcategory → sub-subcategory → product list → detail page), so using a regular scrapy.Spider with explicit callbacks is perfectly suitable. CrawlSpider is better for cases where you have lots of irregular link patterns to auto-match, but your current structure is straightforward enough to handle with manual callbacks.

3. Quick Fixes to Get Your Crawl Working

Here are the most impactful changes to test first:

  • Fix the undefined variable in parse_item:
    def parse_item(self,response):
        # Remove the redundant KartonageItem instantiation
        items = response.meta['items']
        # Use response.xpath (or the table variable you defined) instead of undefined 'a'
        items['Weight'] = response.xpath('.//span[@class="staffelpreise-small"]/text()').get()
        items['Volume'] = response.xpath('.//td[@class="icon_contenct"][7]/text()').get()
        yield items
    
  • Close the category traversal gap:
    def parse_category_cartons(self, response):
        # Grab all subcategory links directly
        subcat_links = response.xpath('//div[@class="cat-thumbnails"]/a/@href')
        if subcat_links:
            for link in subcat_links:
                yield response.follow(url=link.get(), callback=self.parse_target_page)
        else:
            # If no subcategories, this page is a product list—process it directly
            yield from self.parse_target_page(response)
    
  • Fix the Link field in parse_target_page:
    items['Link'] = link.get()  # Store single URL string instead of a list
    
  • Enable debug logging: Run your spider with scrapy crawl kartons12 -s LOG_LEVEL=DEBUG to see exactly where XPaths fail or exceptions are thrown—this will help you debug remaining issues quickly.

内容的提问来源于stack exchange,提问作者kekw

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 17:33:01