You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Thread与ThreadPoolExecutor对比请求实现URL内容变更监控

基于ThreadPoolExecutor实现URL内容变更监控

核心实现逻辑非常直接:你只需要维护一份各URL历史抓取结果的缓存,每次拿到新的抓取结果时和对应URL的历史值做比对,发现差异就打印告警,再更新缓存即可。以下是可以直接运行的改造后代码,完全兼容你现有的线程池逻辑:

import random
import threading
import time
from concurrent.futures import as_completed
from concurrent.futures.thread import ThreadPoolExecutor

import requests
from bs4 import BeautifulSoup

URLS = [
    'https://github.com/search?q=hello+world',
    'https://github.com/search?q=python+3',
    'https://github.com/search?q=world',
    'https://github.com/search?q=i+love+python',
    'https://github.com/search?q=sport+today',
    'https://github.com/search?q=how+to+code',
    'https://github.com/search?q=banana',
    'https://github.com/search?q=android+vs+iphone',
    'https://github.com/search?q=please+help+me',
    'https://github.com/search?q=batman',
]

# 存储每个URL上一次的抓取结果,key是URL,value是repository_results字段值
last_results = {}
# 多线程读写共享字典加锁,避免竞态问题
cache_lock = threading.Lock()
# 每轮全量抓取完成后的等待间隔,单位秒,可自行调整避免触发反爬
LOOP_INTERVAL = 60


def doScrape(response):
    try:
        soup = BeautifulSoup(response.text, 'html.parser')
        result_ele = soup.find("div", {"class": "codesearch-results"}).find("h3")
        return {
            'url': response.url,
            'repository_results': result_ele.text.strip()
        }
    except Exception as e:
        print(f"解析页面失败,URL: {response.url}, 错误: {str(e)}")
        return None


def doRequest(url):
    try:
        response = requests.get(url, timeout=10)
        time.sleep(random.randint(1, 3))
        return response
    except Exception as e:
        print(f"请求失败,URL: {url}, 错误: {str(e)}")
        return None


def check_diff(new_result):
    """比对新结果和历史结果,存在变更则打印提示"""
    url = new_result['url']
    new_val = new_result['repository_results']
    with cache_lock:
        old_val = last_results.get(url)
        # 第一次抓取该URL,直接存缓存不告警
        if old_val is None:
            last_results[url] = new_val
            return
        # 值不一致触发告警
        if old_val != new_val:
            print(f"*[内容变更提示]* URL: {url}")
            print(f"  旧结果: {old_val}")
            print(f"  新结果: {new_val}")
            # 更新缓存为最新值
            last_results[url] = new_val


def ourLoop():
    with ThreadPoolExecutor(max_workers=2) as executor:
        while True:
            # 提交本轮所有抓取任务
            future_tasks = [executor.submit(doRequest, url) for url in URLS]
            for future in as_completed(future_tasks):
                response = future.result()
                # 请求失败直接跳过
                if not response or response.status_code != 200:
                    continue
                scrape_result = doScrape(response)
                # 解析失败直接跳过
                if not scrape_result:
                    continue
                # 比对结果判断是否变更
                check_diff(scrape_result)
            # 本轮所有URL处理完成,等待指定时间后进入下一轮
            print(f"本轮监控完成,等待{LOOP_INTERVAL}秒后进入下一轮监控")
            time.sleep(LOOP_INTERVAL)


if __name__ == "__main__":
    ourLoop()

关键改造点说明

  • 移除了原代码外层手动反复创建Thread的冗余逻辑:ThreadPoolExecutor本身已经实现了线程生命周期管理,直接在线程池内部跑无限监控循环即可,减少不必要的线程创建销毁开销。
  • 新增线程安全的结果缓存:用字典存储各URL的历史抓取结果,搭配线程锁保证多线程读写共享数据时不会出现错乱。
  • 变更判断逻辑:首次抓取到某URL的结果时直接存入缓存不触发告警,后续每次抓取到新结果就和缓存值比对,发现差异立即打印新旧值提示,之后更新缓存为最新结果。
  • 增加异常兜底:覆盖请求超时、请求失败、页面结构变动导致解析失败等场景,单个URL出错不会导致整个监控程序崩溃。
  • 新增轮询间隔配置:每轮全量抓取完成后等待固定时间再发起下一轮请求,你可以根据实际需求调整间隔时长,避免请求频率过高被目标站点拦截。

内容的提问来源于stack exchange,提问作者PythonNewbie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 17:16:17