You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python Selenium移除HTML特定标签并实现精准网站更新监测?

解决方案:避免Selenium提取HTML时的脚本误报问题

针对你的需求,有两个核心解决方向:一是彻底清理页面中的script元素,二是聚焦核心监控目标(PDF文件)直接提取关键信息,后者更精准且能从根源避免误报。

方案一:正确移除页面所有script标签

你之前尝试的driver.find_elements(By.TAG_NAME, 'script').remove()无效,原因是:

  • find_elements返回的是Python元素列表,调用remove()只是操作列表,不会修改页面DOM
  • 必须通过JavaScript直接操作浏览器DOM,才能真正删除页面中的script元素

代码示例:

from selenium import webdriver

driver = webdriver.Chrome()
driver.get("目标网站URL")

# 隐式等待页面基本加载完成
driver.implicitly_wait(10)

# 执行JS删除所有script标签(包括内联、外部加载的脚本)
driver.execute_script("""
    const allScripts = document.querySelectorAll('script');
    allScripts.forEach(script => script.parentNode.removeChild(script));
""")

# 获取清理后的HTML
cleaned_html = driver.page_source

# 后续对比逻辑...
driver.quit()

方案二:聚焦PDF文件提取(推荐)

你的核心需求是监控新上传的PDF,完全没必要对比整个HTML。直接提取PDF的关键信息(链接、标题、更新时间),和历史记录对比,彻底避开广告、Analytics等无关脚本的干扰。

代码示例:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import json
import os

# 加载历史记录(首次运行会自动创建)
HISTORY_FILE = "pdf_history.json"
if os.path.exists(HISTORY_FILE):
    with open(HISTORY_FILE, 'r') as f:
        history = json.load(f)
else:
    history = {}

driver = webdriver.Chrome()
target_url = "目标网站URL"
driver.get(target_url)

# 显式等待PDF链接加载完成(适配动态渲染的页面)
wait = WebDriverWait(driver, 15)
pdf_elements = wait.until(EC.presence_of_all_elements_located((By.XPATH, "//a[contains(@href, '.pdf')]")))

# 提取PDF关键信息
current_pdfs = []
for elem in pdf_elements:
    pdf_url = elem.get_attribute('href')
    pdf_title = elem.text.strip()
    # 尝试提取更新时间(根据目标网站的DOM结构调整xpath)
    try:
        update_time = elem.find_element(By.XPATH, "./parent::div//span[@class='update-date']").text.strip()
    except:
        update_time = "未知"
    current_pdfs.append({"url": pdf_url, "title": pdf_title, "update_time": update_time})

# 对比历史记录,找出新增/更新的PDF
new_pdfs = [pdf for pdf in current_pdfs if pdf["url"] not in history]
updated_pdfs = []
for pdf in current_pdfs:
    if pdf["url"] in history and history[pdf["url"]]["update_time"] != pdf["update_time"]:
        updated_pdfs.append(pdf)

# 发送通知(示例:打印到控制台,可替换为邮件/微信机器人等)
if new_pdfs:
    print(f"发现新PDF:{[p['title'] for p in new_pdfs]}")
if updated_pdfs:
    print(f"发现更新的PDF:{[p['title'] for p in updated_pdfs]}")

# 更新历史记录
for pdf in current_pdfs:
    history[pdf["url"]] = pdf
with open(HISTORY_FILE, 'w') as f:
    json.dump(history, f, indent=2)

driver.quit()

额外优化建议

  • 针对50个网站,可将每个网站的PDF提取规则(如xpath、时间元素选择器)存为配置文件(如JSON),批量遍历执行
  • 对于无限滚动的页面,可添加滚动逻辑确保所有PDF加载完成
  • 使用无头浏览器模式(options.add_argument('--headless=new'))减少资源占用
  • 对比时可对PDF链接做去重处理,避免重复链接触发误报

内容的提问来源于stack exchange,提问作者P Raj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 19:52:40