You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python+Selenium/BeautifulSoup实现网页关键词监控与通知?

如何用Python实现特定关键词的网页监控与通知?

当然可以实现!结合Python的Selenium、BeautifulSoup这类工具完全能搞定你要的「特定关键词监控+触发通知+提取详情」需求。先看看你现有代码的局限——它只是对比整个HTML的变化,没法精准定位关键词,也没法提取对应内容和触发针对性通知。下面给你一套完整的解决方案:

核心思路

  • 用BeautifulSoup(搭配lxml解析器)解析HTML,精准抓取所有包含「GST」的元素文本,不管它在<p>还是<h1>标签里
  • 用Selenium处理动态渲染的网页(如果目标新闻网站是JS加载内容,urllib2这类工具拿不到完整页面)
  • 新增去重机制,避免重复通知相同内容
  • 实现通知(先做控制台通知,可扩展为邮件/桌面弹窗)和详情保存功能
  • 持续定期监控网页变化

准备工作

先安装需要的依赖库:

pip install beautifulsoup4 selenium lxml

完整实现代码

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import time
import os

# 配置Selenium无头模式(后台运行,不弹出浏览器窗口)
chrome_options = Options()
chrome_options.add_argument("--headless=new")
chrome_options.add_argument("--disable-gpu")
driver = webdriver.Chrome(options=chrome_options)

# 配置监控参数
TARGET_URL = "https://www.example.com"  # 替换成你的目标新闻网站URL
KEYWORD = "GST"
MONITOR_INTERVAL = 60  # 监控间隔(秒),建议不要设置太短
DETAIL_FILE = "gst_news_details.txt"  # 保存新闻详情的文件路径
recorded_content = set()  # 存储已记录的内容,避免重复通知

def fetch_html():
    """获取网页内容,支持动态渲染页面"""
    driver.get(TARGET_URL)
    time.sleep(3)  # 等待页面加载完成,可根据网站情况调整
    return driver.page_source

def extract_keyword_content(html):
    """解析HTML,提取所有包含关键词的元素文本"""
    soup = BeautifulSoup(html, "lxml")
    keyword_contents = []
    # 查找所有包含关键词的文本节点,自动匹配所有标签类型
    for element in soup.find_all(text=lambda text: text and KEYWORD in text.strip()):
        clean_text = element.strip()
        # 去重并过滤空内容
        if clean_text and clean_text not in keyword_contents:
            keyword_contents.append(clean_text)
    return keyword_contents

def save_to_file(content):
    """将新的新闻详情追加到文件"""
    with open(DETAIL_FILE, "a", encoding="utf-8") as f:
        timestamp = time.strftime("%Y-%m-%d %H:%M:%S")
        f.write(f"[{timestamp}] 新内容:\n{content}\n\n")

def send_notification(content):
    """触发通知,这里用控制台打印,可扩展为邮件/桌面弹窗"""
    print(f"\n⚠️ 检测到含「{KEYWORD}」的新内容!")
    print(content)
    print("----------------------------------------")

def monitor():
    """主监控逻辑"""
    print(f"开始监控目标网站:{TARGET_URL}")
    print(f"监控关键词:{KEYWORD},间隔:{MONITOR_INTERVAL}秒\n")
    
    # 初始化:抓取初始页面的关键词内容并记录
    initial_html = fetch_html()
    initial_contents = extract_keyword_content(initial_html)
    for content in initial_contents:
        recorded_content.add(content)
        save_to_file(content)
    print(f"初始化完成,已记录{len(initial_contents)}条含关键词内容")
    
    # 持续监控循环
    while True:
        time.sleep(MONITOR_INTERVAL)
        current_html = fetch_html()
        current_contents = extract_keyword_content(current_html)
        
        # 筛选出未记录的新内容
        new_contents = [c for c in current_contents if c not in recorded_content]
        if new_contents:
            for content in new_contents:
                send_notification(content)
                save_to_file(content)
                recorded_content.add(content)
            print(f"已处理{len(new_contents)}条新内容")
        else:
            timestamp = time.strftime("%H:%M:%S")
            print(f"[{timestamp}] 未检测到新的含「{KEYWORD}」内容")

if __name__ == "__main__":
    try:
        monitor()
    except KeyboardInterrupt:
        print("\n监控已手动停止")
        driver.quit()

关键细节说明

  1. 动态页面适配:如果你的目标新闻网站是静态的(不需要JS加载内容),可以把Selenium替换成requests库,代码会更轻量——只需把fetch_html函数改成用requests.get()获取页面内容即可。
  2. 关键词匹配:代码里用lambda text: text and KEYWORD in text.strip()确保只匹配包含关键词的非空文本,避免无效内容。
  3. 去重机制:用set存储已记录的内容,因为集合的去重特性可以轻松避免重复通知相同的新闻内容。
  4. 通知扩展:如果需要更实用的通知方式,可以把send_notification函数改成:
    • 发送邮件:用Python的smtplib库实现
    • 桌面弹窗:用plyer库的notification模块
  5. 监控间隔:建议把间隔设置为60秒以上,避免频繁请求给目标网站造成压力,也降低被封禁IP的风险。

内容的提问来源于stack exchange,提问作者Naiswita

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:28:11