You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用线程实现网页抓取后的字典对比与变化监控

网页监控功能:检测网页变化并输出提示

需求背景

开发网页监控功能,通过GET请求抓取网页数据并存储为字典,定期重复抓取后对比前后数据,检测网页是否变化,在title或repo_count变化时输出提示,脚本需24小时持续运行。现有代码已完成请求发送和数据解析部分,缺少变化检测的核心逻辑。

解决方案

核心是维护一个存储历史数据的全局字典,每次抓取后对比当前数据与历史数据,发现差异则输出提示,同时更新历史数据。

修改后的完整代码

import random
import threading
import time
from concurrent.futures import as_completed
from concurrent.futures.thread import ThreadPoolExecutor

import requests
from bs4 import BeautifulSoup

URLS = [
    'https://github.com/search?q=hello+world',
    'https://github.com/search?q=python+3',
    'https://github.com/search?q=world',
    'https://github.com/search?q=i+love+python',
    'https://github.com/search?q=sport+today',
    'https://github.com/search?q=how+to+code',
    'https://github.com/search?q=banana',
    'https://github.com/search?q=android+vs+iphone',
    'https://github.com/search?q=please+help+me',
    'https://github.com/search?q=batman',
]

# 全局字典存储每个URL的上次抓取数据
previous_data = {}

def doRequest(url):
    response = requests.get(url)
    time.sleep(random.randint(10, 30))
    return response, url

def doScrape(response):
    soup = BeautifulSoup(response.text, 'html.parser')
    return {
        'title': soup.find("input", {"name": "q"})['value'],
        'repo_count': soup.find("span", {"data-search-type": "Repositories"}).text.strip()
    }

def checkDifference(parsed, url):
    global previous_data
    # 首次抓取该URL,仅存储数据
    if url not in previous_data:
        previous_data[url] = parsed
        print(f"首次抓取URL: {url},已存储初始数据")
        return
    
    # 对比数据
    old_data = previous_data[url]
    changes = []
    if parsed['title'] != old_data['title']:
        changes.append(f"title: 从「{old_data['title']}」变为「{parsed['title']}」")
    if parsed['repo_count'] != old_data['repo_count']:
        changes.append(f"repo_count: 从「{old_data['repo_count']}」变为「{parsed['repo_count']}」")
    
    # 输出变化提示
    if changes:
        current_time = time.strftime("%Y-%m-%d %H:%M:%S", time.localtime())
        print(f"[{current_time}] URL {url} 发生变化:")
        for change in changes:
            print(f"  - {change}")
        # 更新历史数据
        previous_data[url] = parsed
    else:
        print(f"[{time.strftime('%Y-%m-%d %H:%M:%S', time.localtime())}] URL {url} 无变化")

def threadPoolLoop():
    with ThreadPoolExecutor(max_workers=1) as executor:
        future_tasks = [
            executor.submit(doRequest, url) for url in URLS
        ]

        for future in as_completed(future_tasks):
            response, url = future.result()
            if response.status_code == 200:
                checkDifference(doScrape(response), url)

if __name__ == "__main__":
    # 设置每次监控循环的间隔时间(秒),比如3600秒=1小时
    CHECK_INTERVAL = 3600
    while True:
        t = threading.Thread(target=threadPoolLoop)
        t.start()
        print('等待当前监控线程完成...')
        t.join()
        print(f"本轮监控完成,等待{CHECK_INTERVAL}秒后开始下一轮...")
        time.sleep(CHECK_INTERVAL)

关键修改点

  • 新增previous_data全局字典,用于存储每个URL的历史抓取数据
  • 实现checkDifference函数:
    • 处理首次抓取的情况,直接存储数据
    • 对比当前数据与历史数据的title和repo_count字段
    • 发现变化时,带时间戳输出具体变化内容
    • 无变化时输出状态提示(可根据需求删除)
  • 新增CHECK_INTERVAL变量,控制每轮监控的间隔时间,避免过于频繁请求
  • 优化主循环逻辑,添加循环间隔等待

内容的提问来源于stack exchange,提问作者PythonNewbie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 12:54:27