You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何判断站点地图变更项所属列表,实现增删改状态识别?

站点地图变更监控:识别增/删状态

需求说明

已实现站点地图变更输出功能,需进一步明确变更的新增、删除状态:

  • 旧列表无、新列表有 → 新增
  • 新列表无、旧列表有 → 删除
    尝试过difflib但输出混乱,希望自行实现逻辑,最终目标是监控sitemap.xml,输出所有变更及对应状态。

现有代码问题

原代码直接将BeautifulSoup的loc对象存入集合,导致集合对比异常;CSV写入逻辑错误,无法记录变更状态;输出内容杂乱,无法区分增删类型。

修正后的代码

import requests
from bs4 import BeautifulSoup
import time
from datetime import datetime
import csv

# CSV字段名
fields = ['变更类型', 'URL', '检测时间'] 

# 目标站点地图URL
url = "https://www.huntermichaelseo.com/testing.xml"

# 请求头模拟浏览器
headers = {'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/39.0.2171.95 Safari/537.36'}

prev_urls = set()
first_run = True

while True:
    # 请求并解析站点地图
    response = requests.get(url, headers=headers)
    soup = BeautifulSoup(response.text, "xml")
    current_loc_tags = soup.find_all('loc')
    current_urls = set(tag.get_text().strip() for tag in current_loc_tags)

    if prev_urls != current_urls:
        if first_run:
            prev_urls = current_urls
            first_run = False
            print(f"开始监控 {url} - {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}")
        else:
            detect_time = datetime.now().strftime('%Y-%m-%d %H:%M:%S')
            print(f"\n检测到变更 - {detect_time}")

            # 计算新增和删除的URL
            added_urls = current_urls - prev_urls
            removed_urls = prev_urls - current_urls

            # 输出变更详情
            if added_urls:
                print("✅ 新增URL:")
                for url_item in added_urls:
                    print(f"  - {url_item}")
            if removed_urls:
                print("❌ 删除URL:")
                for url_item in removed_urls:
                    print(f"  - {url_item}")

            # 写入CSV文件
            with open('sitemap_changes.csv', 'a', newline='', encoding='utf-8') as f:
                writer = csv.writer(f)
                # 首次写入时添加表头
                if f.tell() == 0:
                    writer.writerow(fields)
                # 写入新增记录
                for url_item in added_urls:
                    writer.writerow(['新增', url_item, detect_time])
                # 写入删除记录
                for url_item in removed_urls:
                    writer.writerow(['删除', url_item, detect_time])

            # 更新历史URL集合
            prev_urls = current_urls
    else:
        print(f"\n无变更 - {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}")
    
    time.sleep(5)

关键修改说明

  • 提取URL文本:不再直接存储loc对象,而是提取标签内的URL文本并去重,确保集合对比准确
  • 明确增删逻辑:通过集合差集直接计算新增(current_urls - prev_urls)和删除(prev_urls - current_urls)的URL
  • 优化输出格式:用清晰标识区分新增/删除,提升可读性
  • CSV写入优化:采用追加模式避免覆盖历史记录,自动判断表头写入,完整记录变更类型、URL和检测时间
  • 代码规范:变量名语义化,使用f-string提升可读性

内容的提问来源于stack exchange,提问作者Hunter Michael

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 06:39:18