You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python脚本监测WHM配额页特定国家状态文本变化

问题

我想编写Python脚本,监测澳大利亚移民局WHM项目配额状态页面中特定国家(如厄瓜多尔)的状态显示文本变化。请问该如何追踪该特定文本的变更情况?目前我尝试通过时间间隔对比网页哈希值的方法,但即便页面视觉无变化,哈希值仍每次都改变。当前代码如下:

input_website = 'https://immi.homeaffairs.gov.au/what-we-do/whm-program/status-of-country-caps'
time_delay = 60

#Monitor the website
def monitor_website():
    # Run the loop the keep monitoring
    while True:
        # Visit the website to know if it is up
        status = urllib.request.urlopen(input_website).getcode()
        # If it returns 200, the website is up
        if status != 200:
            # Call email function
            send_email("The website is DOWN")
        else:
            send_email("The website is UP")
            # Open url and create the hash code
            response = urllib.request.urlopen(input_website).read()
            current_hash = hashlib.sha224(response).hexdigest()
            # Revisit the website after time delay
            time.sleep(time_delay)
            # Visit the website after delay, and generate the new website
            response = urllib.request.urlopen(input_website).read()
            new_hash = hashlib.sha224(response).hexdigest()
            # Check the hash codes
            if new_hash != current_hash:
                send_email("The website CHANGED")
解决方案

哈希值每次变化的核心原因是:页面包含动态生成的内容(比如加载时间戳、随机验证token、广告脚本参数),这些内容每次请求都会更新,但和你关注的国家配额状态无关。正确的做法是只提取并对比目标国家的状态文本,而非整个页面的哈希值。

具体实现步骤

  1. 使用HTML解析库精准定位目标内容,忽略页面其他动态元素
  2. 存储目标状态的历史值,仅在状态发生变化时触发通知
  3. 优化请求逻辑,减少重复请求并添加异常处理

修改后的代码示例

import urllib.request
import time
from bs4 import BeautifulSoup

input_website = 'https://immi.homeaffairs.gov.au/what-we-do/whm-program/status-of-country-caps'
time_delay = 60
target_country = "Ecuador"  # 可替换为你关注的国家
previous_status = None

def send_email(message):
    # 保留你原有的邮件发送逻辑,这里用打印做测试示例
    print(message)

def get_country_status():
    try:
        response = urllib.request.urlopen(input_website)
        if response.getcode() != 200:
            return None, "网站无法访问"
        
        # 解析HTML页面
        soup = BeautifulSoup(response.read(), 'html.parser')
        # 遍历表格行,定位目标国家的状态单元格
        for row in soup.find_all('tr'):
            cells = row.find_all('td')
            if len(cells) >= 2 and target_country.strip() in cells[0].text.strip():
                return cells[1].text.strip(), None
        return None, f"未找到目标国家: {target_country}"
    except Exception as e:
        return None, f"请求出错: {str(e)}"

def monitor_website():
    global previous_status
    while True:
        current_status, error_msg = get_country_status()
        
        if error_msg:
            send_email(error_msg)
        else:
            if previous_status is None:
                # 首次运行,记录初始状态
                previous_status = current_status
                send_email(f"监测启动,{target_country}当前状态: {current_status}")
            elif current_status != previous_status:
                # 状态变更时发送通知
                send_email(f"{target_country}状态已更新!旧状态: {previous_status} | 新状态: {current_status}")
                previous_status = current_status
        
        time.sleep(time_delay)

if __name__ == "__main__":
    monitor_website()

代码说明

  • 用BeautifulSoup解析HTML,精准提取目标国家的状态文本,彻底避免动态内容干扰
  • 仅对比目标状态的文本内容,解决了哈希值无故变化的问题
  • 添加异常处理,捕获请求、解析过程中的错误,提升脚本稳定性
  • 仅在首次启动或状态变更时发送通知,减少冗余邮件

内容的提问来源于stack exchange,提问作者Yos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 04:18:16