You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬取/解析过程中如何自动去除重复内容?

爬取/解析过程中自动去重的实现方案

你现在的核心问题是没在爬取阶段做去重逻辑,导致最终输出有重复内容。最直接高效的解决方式是用**集合(set)**存储爬取到的内容——集合会自动忽略重复元素,从根源上避免重复项产生,比事后再处理要高效得多。

关键修改点

  • 把存储内容的容器从列表改成集合,实现爬取时实时去重
  • 用WebDriverWait等待元素加载,替代固定的time.sleep,提升稳定性和效率
  • 修正原代码里的类型注解错误、变量名冲突等细节问题
  • 逐行处理爬取到的文本,确保每一行都被去重

修改后的完整代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.chrome.options import Options
import time

# URL LIST:
urls = [
    'URL0',
    'URL1',
]
# OPTIONS/REQUESTS:
options = Options()
options.add_argument("--disable-extensions")
options.add_argument('--disable-blink-features=AutomationControlled')
options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')
# options.add_argument("--headless")  # 修正原拼写错误:是headless不是headerless
options.page_load_strategy = 'normal'

driver = webdriver.Chrome(options=options)
driver.implicitly_wait(10)
wait = WebDriverWait(driver, 10)

# 用集合存储内容,自动去重
spike: set[str] = set()

# 遍历URL列表(修正原变量名冲突:把循环变量url改为避免和列表名重复)
for target_url in urls:
    driver.get(target_url)
    try:
        # 等待目标pre元素加载完成,替代固定sleep
        wait.until(EC.presence_of_all_elements_located((By.XPATH, '(//pre)[position() < 10]')))
        # 遍历每个pre元素
        for ele in driver.find_elements(By.XPATH, '(//pre)[position() < 10]'):
            # 拆分每一行,逐个添加到集合(自动去重)
            lines = ele.text.splitlines(keepends=False)
            for line in lines:
                if line.strip():  # 可选:过滤空行,不需要可以删除此行
                    spike.add(line)
            time.sleep(1)  # 缩短不必要的等待时间
    except Exception as e:
        print(f"处理URL {target_url}时出错: {e}")
        continue

# 关闭浏览器
driver.quit()

# 保存去重后的内容到文件
# 集合是无序的,转成有序列表保证输出顺序稳定
sorted_lines = sorted(spike)
with open("text.txt", "wt", encoding="utf-8") as file:
    file.write("\n".join(sorted_lines))

代码说明

  1. 集合去重:spike.add(line)会自动跳过已经存在的行,不需要额外写判断逻辑,直接在爬取阶段完成去重。
  2. 等待优化:用WebDriverWait等待元素加载,比固定time.sleep(10)更灵活——元素加载完成就继续执行,不会浪费不必要的时间。
  3. 细节修正:修正了原代码里的变量名冲突、拼写错误等问题,让代码更健壮。
  4. 有序输出:集合本身是无序的,用sorted()转成有序列表后再保存,保证输出的文本有稳定的顺序。

内容的提问来源于stack exchange,提问作者spike666spike666

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 04:16:38