如何基于Thread与ThreadPoolExecutor对比请求实现URL内容变更监控
基于ThreadPoolExecutor实现URL内容变更监控
核心实现逻辑非常直接:你只需要维护一份各URL历史抓取结果的缓存,每次拿到新的抓取结果时和对应URL的历史值做比对,发现差异就打印告警,再更新缓存即可。以下是可以直接运行的改造后代码,完全兼容你现有的线程池逻辑:
import random import threading import time from concurrent.futures import as_completed from concurrent.futures.thread import ThreadPoolExecutor import requests from bs4 import BeautifulSoup URLS = [ 'https://github.com/search?q=hello+world', 'https://github.com/search?q=python+3', 'https://github.com/search?q=world', 'https://github.com/search?q=i+love+python', 'https://github.com/search?q=sport+today', 'https://github.com/search?q=how+to+code', 'https://github.com/search?q=banana', 'https://github.com/search?q=android+vs+iphone', 'https://github.com/search?q=please+help+me', 'https://github.com/search?q=batman', ] # 存储每个URL上一次的抓取结果,key是URL,value是repository_results字段值 last_results = {} # 多线程读写共享字典加锁,避免竞态问题 cache_lock = threading.Lock() # 每轮全量抓取完成后的等待间隔,单位秒,可自行调整避免触发反爬 LOOP_INTERVAL = 60 def doScrape(response): try: soup = BeautifulSoup(response.text, 'html.parser') result_ele = soup.find("div", {"class": "codesearch-results"}).find("h3") return { 'url': response.url, 'repository_results': result_ele.text.strip() } except Exception as e: print(f"解析页面失败,URL: {response.url}, 错误: {str(e)}") return None def doRequest(url): try: response = requests.get(url, timeout=10) time.sleep(random.randint(1, 3)) return response except Exception as e: print(f"请求失败,URL: {url}, 错误: {str(e)}") return None def check_diff(new_result): """比对新结果和历史结果,存在变更则打印提示""" url = new_result['url'] new_val = new_result['repository_results'] with cache_lock: old_val = last_results.get(url) # 第一次抓取该URL,直接存缓存不告警 if old_val is None: last_results[url] = new_val return # 值不一致触发告警 if old_val != new_val: print(f"*[内容变更提示]* URL: {url}") print(f" 旧结果: {old_val}") print(f" 新结果: {new_val}") # 更新缓存为最新值 last_results[url] = new_val def ourLoop(): with ThreadPoolExecutor(max_workers=2) as executor: while True: # 提交本轮所有抓取任务 future_tasks = [executor.submit(doRequest, url) for url in URLS] for future in as_completed(future_tasks): response = future.result() # 请求失败直接跳过 if not response or response.status_code != 200: continue scrape_result = doScrape(response) # 解析失败直接跳过 if not scrape_result: continue # 比对结果判断是否变更 check_diff(scrape_result) # 本轮所有URL处理完成,等待指定时间后进入下一轮 print(f"本轮监控完成,等待{LOOP_INTERVAL}秒后进入下一轮监控") time.sleep(LOOP_INTERVAL) if __name__ == "__main__": ourLoop()
关键改造点说明
- 移除了原代码外层手动反复创建Thread的冗余逻辑:ThreadPoolExecutor本身已经实现了线程生命周期管理,直接在线程池内部跑无限监控循环即可,减少不必要的线程创建销毁开销。
- 新增线程安全的结果缓存:用字典存储各URL的历史抓取结果,搭配线程锁保证多线程读写共享数据时不会出现错乱。
- 变更判断逻辑:首次抓取到某URL的结果时直接存入缓存不触发告警,后续每次抓取到新结果就和缓存值比对,发现差异立即打印新旧值提示,之后更新缓存为最新结果。
- 增加异常兜底:覆盖请求超时、请求失败、页面结构变动导致解析失败等场景,单个URL出错不会导致整个监控程序崩溃。
- 新增轮询间隔配置:每轮全量抓取完成后等待固定时间再发起下一轮请求,你可以根据实际需求调整间隔时长,避免请求频率过高被目标站点拦截。
内容的提问来源于stack exchange,提问作者PythonNewbie
相关产品推荐
相关产品推荐

