如何用Python+Selenium/BeautifulSoup实现网页关键词监控与通知?
如何用Python实现特定关键词的网页监控与通知?
当然可以实现!结合Python的Selenium、BeautifulSoup这类工具完全能搞定你要的「特定关键词监控+触发通知+提取详情」需求。先看看你现有代码的局限——它只是对比整个HTML的变化,没法精准定位关键词,也没法提取对应内容和触发针对性通知。下面给你一套完整的解决方案:
核心思路
- 用BeautifulSoup(搭配lxml解析器)解析HTML,精准抓取所有包含「GST」的元素文本,不管它在
<p>还是<h1>标签里 - 用Selenium处理动态渲染的网页(如果目标新闻网站是JS加载内容,urllib2这类工具拿不到完整页面)
- 新增去重机制,避免重复通知相同内容
- 实现通知(先做控制台通知,可扩展为邮件/桌面弹窗)和详情保存功能
- 持续定期监控网页变化
准备工作
先安装需要的依赖库:
pip install beautifulsoup4 selenium lxml
完整实现代码
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options import time import os # 配置Selenium无头模式(后台运行,不弹出浏览器窗口) chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("--disable-gpu") driver = webdriver.Chrome(options=chrome_options) # 配置监控参数 TARGET_URL = "https://www.example.com" # 替换成你的目标新闻网站URL KEYWORD = "GST" MONITOR_INTERVAL = 60 # 监控间隔(秒),建议不要设置太短 DETAIL_FILE = "gst_news_details.txt" # 保存新闻详情的文件路径 recorded_content = set() # 存储已记录的内容,避免重复通知 def fetch_html(): """获取网页内容,支持动态渲染页面""" driver.get(TARGET_URL) time.sleep(3) # 等待页面加载完成,可根据网站情况调整 return driver.page_source def extract_keyword_content(html): """解析HTML,提取所有包含关键词的元素文本""" soup = BeautifulSoup(html, "lxml") keyword_contents = [] # 查找所有包含关键词的文本节点,自动匹配所有标签类型 for element in soup.find_all(text=lambda text: text and KEYWORD in text.strip()): clean_text = element.strip() # 去重并过滤空内容 if clean_text and clean_text not in keyword_contents: keyword_contents.append(clean_text) return keyword_contents def save_to_file(content): """将新的新闻详情追加到文件""" with open(DETAIL_FILE, "a", encoding="utf-8") as f: timestamp = time.strftime("%Y-%m-%d %H:%M:%S") f.write(f"[{timestamp}] 新内容:\n{content}\n\n") def send_notification(content): """触发通知,这里用控制台打印,可扩展为邮件/桌面弹窗""" print(f"\n⚠️ 检测到含「{KEYWORD}」的新内容!") print(content) print("----------------------------------------") def monitor(): """主监控逻辑""" print(f"开始监控目标网站:{TARGET_URL}") print(f"监控关键词:{KEYWORD},间隔:{MONITOR_INTERVAL}秒\n") # 初始化:抓取初始页面的关键词内容并记录 initial_html = fetch_html() initial_contents = extract_keyword_content(initial_html) for content in initial_contents: recorded_content.add(content) save_to_file(content) print(f"初始化完成,已记录{len(initial_contents)}条含关键词内容") # 持续监控循环 while True: time.sleep(MONITOR_INTERVAL) current_html = fetch_html() current_contents = extract_keyword_content(current_html) # 筛选出未记录的新内容 new_contents = [c for c in current_contents if c not in recorded_content] if new_contents: for content in new_contents: send_notification(content) save_to_file(content) recorded_content.add(content) print(f"已处理{len(new_contents)}条新内容") else: timestamp = time.strftime("%H:%M:%S") print(f"[{timestamp}] 未检测到新的含「{KEYWORD}」内容") if __name__ == "__main__": try: monitor() except KeyboardInterrupt: print("\n监控已手动停止") driver.quit()
关键细节说明
- 动态页面适配:如果你的目标新闻网站是静态的(不需要JS加载内容),可以把Selenium替换成
requests库,代码会更轻量——只需把fetch_html函数改成用requests.get()获取页面内容即可。 - 关键词匹配:代码里用
lambda text: text and KEYWORD in text.strip()确保只匹配包含关键词的非空文本,避免无效内容。 - 去重机制:用
set存储已记录的内容,因为集合的去重特性可以轻松避免重复通知相同的新闻内容。 - 通知扩展:如果需要更实用的通知方式,可以把
send_notification函数改成:- 发送邮件:用Python的
smtplib库实现 - 桌面弹窗:用
plyer库的notification模块
- 发送邮件:用Python的
- 监控间隔:建议把间隔设置为60秒以上,避免频繁请求给目标网站造成压力,也降低被封禁IP的风险。
内容的提问来源于stack exchange,提问作者Naiswita
相关产品推荐
相关产品推荐

