如何使用线程实现网页抓取后的字典对比与变化监控
网页监控功能:检测网页变化并输出提示
需求背景
开发网页监控功能,通过GET请求抓取网页数据并存储为字典,定期重复抓取后对比前后数据,检测网页是否变化,在title或repo_count变化时输出提示,脚本需24小时持续运行。现有代码已完成请求发送和数据解析部分,缺少变化检测的核心逻辑。
解决方案
核心是维护一个存储历史数据的全局字典,每次抓取后对比当前数据与历史数据,发现差异则输出提示,同时更新历史数据。
修改后的完整代码
import random import threading import time from concurrent.futures import as_completed from concurrent.futures.thread import ThreadPoolExecutor import requests from bs4 import BeautifulSoup URLS = [ 'https://github.com/search?q=hello+world', 'https://github.com/search?q=python+3', 'https://github.com/search?q=world', 'https://github.com/search?q=i+love+python', 'https://github.com/search?q=sport+today', 'https://github.com/search?q=how+to+code', 'https://github.com/search?q=banana', 'https://github.com/search?q=android+vs+iphone', 'https://github.com/search?q=please+help+me', 'https://github.com/search?q=batman', ] # 全局字典存储每个URL的上次抓取数据 previous_data = {} def doRequest(url): response = requests.get(url) time.sleep(random.randint(10, 30)) return response, url def doScrape(response): soup = BeautifulSoup(response.text, 'html.parser') return { 'title': soup.find("input", {"name": "q"})['value'], 'repo_count': soup.find("span", {"data-search-type": "Repositories"}).text.strip() } def checkDifference(parsed, url): global previous_data # 首次抓取该URL,仅存储数据 if url not in previous_data: previous_data[url] = parsed print(f"首次抓取URL: {url},已存储初始数据") return # 对比数据 old_data = previous_data[url] changes = [] if parsed['title'] != old_data['title']: changes.append(f"title: 从「{old_data['title']}」变为「{parsed['title']}」") if parsed['repo_count'] != old_data['repo_count']: changes.append(f"repo_count: 从「{old_data['repo_count']}」变为「{parsed['repo_count']}」") # 输出变化提示 if changes: current_time = time.strftime("%Y-%m-%d %H:%M:%S", time.localtime()) print(f"[{current_time}] URL {url} 发生变化:") for change in changes: print(f" - {change}") # 更新历史数据 previous_data[url] = parsed else: print(f"[{time.strftime('%Y-%m-%d %H:%M:%S', time.localtime())}] URL {url} 无变化") def threadPoolLoop(): with ThreadPoolExecutor(max_workers=1) as executor: future_tasks = [ executor.submit(doRequest, url) for url in URLS ] for future in as_completed(future_tasks): response, url = future.result() if response.status_code == 200: checkDifference(doScrape(response), url) if __name__ == "__main__": # 设置每次监控循环的间隔时间(秒),比如3600秒=1小时 CHECK_INTERVAL = 3600 while True: t = threading.Thread(target=threadPoolLoop) t.start() print('等待当前监控线程完成...') t.join() print(f"本轮监控完成,等待{CHECK_INTERVAL}秒后开始下一轮...") time.sleep(CHECK_INTERVAL)
关键修改点
- 新增
previous_data全局字典,用于存储每个URL的历史抓取数据 - 实现
checkDifference函数:- 处理首次抓取的情况,直接存储数据
- 对比当前数据与历史数据的
title和repo_count字段 - 发现变化时,带时间戳输出具体变化内容
- 无变化时输出状态提示(可根据需求删除)
- 新增
CHECK_INTERVAL变量,控制每轮监控的间隔时间,避免过于频繁请求 - 优化主循环逻辑,添加循环间隔等待
内容的提问来源于stack exchange,提问作者PythonNewbie
相关产品推荐
相关产品推荐

