如何用Scrapy去除重复数据?页面实际695条却爬取954条
解决Scrapy爬虫爬取数据重复问题的方法
问题根源
当前爬虫把翻页请求逻辑写在了遍历单条数据的循环内部,导致每处理一条记录就会发起一次翻页请求,同一个页面被重复爬取多次,最终产生大量重复数据。
解决步骤
1. 调整翻页逻辑位置
将翻页请求代码移到遍历数据的循环外部,确保每个页面仅被请求一次,从根源减少重复爬取。
2. 基于唯一标识去重
以Регистрационный номер адвоката в реестре(律师注册编号)作为唯一标识,因为注册编号具有唯一性,通过维护一个集合存储已爬取的编号,过滤重复记录。
修正后的完整代码
import scrapy from scrapy.http import Request class PushpaSpider(scrapy.Spider): name = 'test' start_urls = ['http://www.palatakd.ru/list/'] page_number = 1 # 存储已爬取的注册编号,用于去重 crawled_registrations = set() def parse(self, response): details = response.xpath("//p[@class='detail_block']") for detail in details: registration = detail.xpath(".//span[contains(.,'Регистрационный номер адвоката в реестре')]//following-sibling::span//text()").get() # 跳过已爬取过的记录 if registration in self.crawled_registrations: continue self.crawled_registrations.add(registration) address = detail.xpath(".//span[contains(.,'Адрес')]//following-sibling::span//text()").get() phone = detail.xpath(".//span[contains(.,'Телефон')]//following-sibling::span//text()").get() fax = detail.xpath(".//span[contains(.,'Факс')]//following-sibling::span//text()").get() yield { 'Телефон': phone, 'Факс': fax, 'Регистрационный номер адвоката в реестре': registration, 'Адрес': address } # 翻页逻辑移到循环外,确保每个页面只请求一次 next_page = 'http://www.palatakd.ru/list/?PAGEN_1=' + str(PushpaSpider.page_number) if PushpaSpider.page_number <= 3: PushpaSpider.page_number += 1 yield response.follow(next_page, callback=self.parse)
补充说明
如果遇到注册编号为空的情况,可以组合地址+电话等多个字段作为唯一标识,确保去重效果。
内容的提问来源于stack exchange,提问作者Amen Aziz
相关产品推荐
相关产品推荐

