如何在Scrapy爬虫中添加多XPath规则并导出CSV文件
解决Scrapy爬虫中添加多个CSS/XPath规则的问题
嘿,我懂你的困扰啦!你现在遇到的核心问题是Python类里的同名方法会被覆盖——你写了两个parse方法,第二个直接把第一个给替换掉了,所以爬虫只会执行第二个parse里的逻辑,第一个规则自然就没生效啦。
下面给你几种可行的解决方案,你可以根据自己的需求选:
方案1:在同一个parse方法里处理多组规则
直接把两个提取逻辑放在同一个parse方法里,这样爬虫会依次执行这两段代码,把两组数据都yield出来。导出CSV的时候,Scrapy会自动把所有出现过的字段作为列,缺失的字段会填充为空值,完全不影响结果。
修改后的代码示例:
import scrapy class TestSetSpider(scrapy.Spider): name = "test_spider" start_urls = ['https://example.html'] def parse(self, response): # 提取product-name下的name字段 for brickset in response.xpath('//div[@class="product-name"]'): yield { 'name': brickset.xpath('h1/text()').extract_first(), 'color': None # 补充color字段,避免CSV列缺失 } # 提取another-class下的color字段 for attempt in response.xpath('//div[@class="another-class"]'): yield { 'name': None, # 补充name字段,避免CSV列缺失 'color': attempt.xpath('h1/a/text()').extract_first() }
如果你的页面结构里,product-name和another-class是属于同一个商品的不同模块(比如每个商品同时有名称和颜色),那可以调整XPath,一次性提取同一商品的所有信息,这样数据会更规整:
def parse(self, response): # 假设每个商品的容器是product-container,根据实际页面结构修改 for product in response.xpath('//div[@class="product-container"]'): yield { 'name': product.xpath('.//div[@class="product-name"]/h1/text()').extract_first(), 'color': product.xpath('.//div[@class="another-class"]/h1/a/text()').extract_first() }
方案2:拆分提取逻辑到辅助方法
如果规则很多,把所有代码堆在parse里会显得杂乱,你可以把不同的提取逻辑拆成独立的辅助方法,然后在parse里调用它们,这样代码可读性更强:
import scrapy class TestSetSpider(scrapy.Spider): name = "test_spider" start_urls = ['https://example.html'] def parse(self, response): # 调用提取名称的方法 yield from self.extract_product_names(response) # 调用提取颜色的方法 yield from self.extract_product_colors(response) def extract_product_names(self, response): for brickset in response.xpath('//div[@class="product-name"]'): yield { 'name': brickset.xpath('h1/text()').extract_first(), 'color': None } def extract_product_colors(self, response): for attempt in response.xpath('//div[@class="another-class"]'): yield { 'name': None, 'color': attempt.xpath('h1/a/text()').extract_first() }
这里用yield from来遍历辅助方法返回的生成器,和直接循环yield效果是一样的,只是写法更简洁。
补充说明
运行爬虫的命令还是你原来的:
scrapy crawl test_spider -o test.csv
导出的CSV文件会包含name和color两列,对应的数据会正确填充,缺失的位置显示为空。
内容的提问来源于stack exchange,提问作者pali112
相关产品推荐
相关产品推荐

