You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何消除爬虫CSV中因网页关键词重复出现导致的重复链接

解决爬虫重复记录关键词与URL的问题

问题根源

你的代码会遍历网页内关键词的每一次出现,并输出一条对应记录。如果同一关键词在页面中多次出现,就会生成多条重复的URL记录。

解决方案

我们只需检测关键词是否在网页中存在,而非遍历每一次出现的位置。同时可以用集合记录已处理过的(关键词,URL)对,避免单页内或跨页面的重复记录。

修改后的代码

allowed_domains = ["www.geo.tv"]
start_urls = ["https://www.geo.tv/"]
rules = [Rule(LinkExtractor(), follow=True, callback="check_buzzwords")]

crawl_count = 0
words_found = 0
# 存储已记录的(关键词,URL)组合,避免重复
recorded_pairs = set()

def check_buzzwords(self, response):
    self.__class__.crawl_count += 1
    crawl_count = self.__class__.crawl_count

    wordlist = [
        "Imran",
        "Hello",
        "Nauman",
    ]

    url = response.url
    contenttype = response.headers.get("content-type", "").decode('utf-8').lower()
    data = response.body.decode('utf-8')

    for word in wordlist:
        # 检查关键词是否存在于页面内容中
        if word in data:
            pair = (word, url)
            # 检查该组合是否已被记录
            if pair not in self.__class__.recorded_pairs:
                self.__class__.recorded_pairs.add(pair)
                self.__class__.words_found += 1
                print(f"{word};{url};")
    return Item()

关键修改点

  • 移除对find_all_substrings的遍历逻辑,改为直接判断关键词是否存在于页面内容
  • 添加全局集合recorded_pairs,确保同一(关键词,URL)组合仅被记录一次
  • 简化输出逻辑,只在关键词存在且未被记录时生成记录

内容的提问来源于stack exchange,提问作者Nauman Asif

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 06:08:12