如何消除爬虫CSV中因网页关键词重复出现导致的重复链接
解决爬虫重复记录关键词与URL的问题
问题根源
你的代码会遍历网页内关键词的每一次出现,并输出一条对应记录。如果同一关键词在页面中多次出现,就会生成多条重复的URL记录。
解决方案
我们只需检测关键词是否在网页中存在,而非遍历每一次出现的位置。同时可以用集合记录已处理过的(关键词,URL)对,避免单页内或跨页面的重复记录。
修改后的代码
allowed_domains = ["www.geo.tv"] start_urls = ["https://www.geo.tv/"] rules = [Rule(LinkExtractor(), follow=True, callback="check_buzzwords")] crawl_count = 0 words_found = 0 # 存储已记录的(关键词,URL)组合,避免重复 recorded_pairs = set() def check_buzzwords(self, response): self.__class__.crawl_count += 1 crawl_count = self.__class__.crawl_count wordlist = [ "Imran", "Hello", "Nauman", ] url = response.url contenttype = response.headers.get("content-type", "").decode('utf-8').lower() data = response.body.decode('utf-8') for word in wordlist: # 检查关键词是否存在于页面内容中 if word in data: pair = (word, url) # 检查该组合是否已被记录 if pair not in self.__class__.recorded_pairs: self.__class__.recorded_pairs.add(pair) self.__class__.words_found += 1 print(f"{word};{url};") return Item()
关键修改点
- 移除对
find_all_substrings的遍历逻辑,改为直接判断关键词是否存在于页面内容 - 添加全局集合
recorded_pairs,确保同一(关键词,URL)组合仅被记录一次 - 简化输出逻辑,只在关键词存在且未被记录时生成记录
内容的提问来源于stack exchange,提问作者Nauman Asif
相关产品推荐
相关产品推荐

