You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Scrapy爬虫中添加多XPath规则并导出CSV文件

解决Scrapy爬虫中添加多个CSS/XPath规则的问题

嘿,我懂你的困扰啦!你现在遇到的核心问题是Python类里的同名方法会被覆盖——你写了两个parse方法,第二个直接把第一个给替换掉了,所以爬虫只会执行第二个parse里的逻辑,第一个规则自然就没生效啦。

下面给你几种可行的解决方案,你可以根据自己的需求选:

方案1:在同一个parse方法里处理多组规则

直接把两个提取逻辑放在同一个parse方法里,这样爬虫会依次执行这两段代码,把两组数据都yield出来。导出CSV的时候,Scrapy会自动把所有出现过的字段作为列,缺失的字段会填充为空值,完全不影响结果。

修改后的代码示例:

import scrapy
class TestSetSpider(scrapy.Spider):
    name = "test_spider"
    start_urls = ['https://example.html']
    
    def parse(self, response):
        # 提取product-name下的name字段
        for brickset in response.xpath('//div[@class="product-name"]'):
            yield {
                'name': brickset.xpath('h1/text()').extract_first(),
                'color': None  # 补充color字段,避免CSV列缺失
            }
        
        # 提取another-class下的color字段
        for attempt in response.xpath('//div[@class="another-class"]'):
            yield {
                'name': None,  # 补充name字段,避免CSV列缺失
                'color': attempt.xpath('h1/a/text()').extract_first()
            }

如果你的页面结构里,product-name和another-class是属于同一个商品的不同模块(比如每个商品同时有名称和颜色),那可以调整XPath,一次性提取同一商品的所有信息,这样数据会更规整:

def parse(self, response):
    # 假设每个商品的容器是product-container,根据实际页面结构修改
    for product in response.xpath('//div[@class="product-container"]'):
        yield {
            'name': product.xpath('.//div[@class="product-name"]/h1/text()').extract_first(),
            'color': product.xpath('.//div[@class="another-class"]/h1/a/text()').extract_first()
        }

方案2:拆分提取逻辑到辅助方法

如果规则很多,把所有代码堆在parse里会显得杂乱,你可以把不同的提取逻辑拆成独立的辅助方法,然后在parse里调用它们,这样代码可读性更强:

import scrapy
class TestSetSpider(scrapy.Spider):
    name = "test_spider"
    start_urls = ['https://example.html']
    
    def parse(self, response):
        # 调用提取名称的方法
        yield from self.extract_product_names(response)
        # 调用提取颜色的方法
        yield from self.extract_product_colors(response)
    
    def extract_product_names(self, response):
        for brickset in response.xpath('//div[@class="product-name"]'):
            yield {
                'name': brickset.xpath('h1/text()').extract_first(),
                'color': None
            }
    
    def extract_product_colors(self, response):
        for attempt in response.xpath('//div[@class="another-class"]'):
            yield {
                'name': None,
                'color': attempt.xpath('h1/a/text()').extract_first()
            }

这里用yield from来遍历辅助方法返回的生成器,和直接循环yield效果是一样的,只是写法更简洁。

补充说明

运行爬虫的命令还是你原来的:

scrapy crawl test_spider -o test.csv

导出的CSV文件会包含name和color两列,对应的数据会正确填充,缺失的位置显示为空。

内容的提问来源于stack exchange,提问作者pali112

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:10:34