如何用Python实现类似原生Watchdog的功能?解决脚本需重启问题
问题修复与实时监控实现方案
一、解决现有脚本需重启才能整理的问题
你的脚本必须重启才能处理新文件,核心原因是只在启动时扫描一次HTML文件,循环过程中不会重新检测新生成的文件。下面是修复后的完整脚本,同时做了多项优化:
import os import re import shutil import time report_path = "/home/user/reports/" source_dir = "/home/user/" # 替代自定义EndsWith类,用原生方法筛选HTML文件 def get_html_files(dir_path): return [f for f in os.listdir(dir_path) if f.endswith('.html')] while True: html_files = get_html_files(source_dir) if not html_files: # 暂停1秒降低CPU占用 time.sleep(1) continue # 提取第一个文件的前缀作为文件夹名,兼容无下划线的文件 first_file = html_files[0] match_result = re.match(r"^[^_]+(?=_)", first_file) folder_name = match_result.group() if match_result else "unclassified" target_folder = os.path.join(report_path, folder_name) # 自动创建文件夹,已存在则跳过 os.makedirs(target_folder, exist_ok=True) # 逐个移动文件,避免批量移动误操作 for file in html_files: source_path = os.path.join(source_dir, file) target_path = os.path.join(target_folder, file) shutil.move(source_path, target_path) time.sleep(1)
优化细节说明
- 去掉冗余的自定义
EndsWith类,原生endswith足够满足筛选需求,代码更简洁。 - 把文件扫描逻辑放到循环内部,每次循环都重新获取最新的HTML文件列表。
- 用
shutil.move替代os.system调用系统命令,更安全且符合Python规范。 - 增加文件名无下划线的兼容处理,避免正则匹配报错。
- 加入
time.sleep减少循环对CPU的消耗。
二、实现类似Watchdog的实时监控功能
如果想要像Watchdog那样实时响应文件创建事件,而不是定时轮询,可以分两种方式实现:
1. 优化版轮询监控(简单易实现)
通过记录上次扫描的时间,只处理新增或修改的文件,比原始轮询更高效:
import os import re import shutil import time report_path = "/home/user/reports/" source_dir = "/home/user/" last_scan_timestamp = time.time() def get_new_html_files(dir_path, last_time): new_files = [] for f in os.listdir(dir_path): if f.endswith('.html'): file_path = os.path.join(dir_path, f) # 只处理上次扫描后新增/修改的文件 if os.path.getmtime(file_path) > last_time: new_files.append(f) return new_files while True: current_time = time.time() new_html_files = get_new_html_files(source_dir, last_scan_timestamp) if new_html_files: first_file = new_html_files[0] match_result = re.match(r"^[^_]+(?=_)", first_file) folder_name = match_result.group() if match_result else "unclassified" target_folder = os.path.join(report_path, folder_name) os.makedirs(target_folder, exist_ok=True) for file in new_html_files: source_path = os.path.join(source_dir, file) target_path = os.path.join(target_folder, file) # 跳过正在写入的文件,避免移动失败 try: with open(source_path, 'r'): shutil.move(source_path, target_path) except PermissionError: continue last_scan_timestamp = current_time time.sleep(1)
2. 系统原生事件监控(高效实时)
利用操作系统的文件系统事件API(如Linux的inotify、macOS的FSEvents),实现真正的实时监控。以Linux为例,用inotify_simple库实现:
import os import re import shutil import time from inotify_simple import INotify, flags report_path = "/home/user/reports/" source_dir = "/home/user/" # 初始化inotify监听 inotify = INotify() # 监听文件创建和移动到文件夹的事件 watch_flags = flags.CREATE | flags.MOVED_TO wd = inotify.add_watch(source_dir, watch_flags) def process_new_file(file_name): if not file_name.endswith('.html'): return match_result = re.match(r"^[^_]+(?=_)", file_name) folder_name = match_result.group() if match_result else "unclassified" target_folder = os.path.join(report_path, folder_name) os.makedirs(target_folder, exist_ok=True) source_path = os.path.join(source_dir, file_name) # 等待文件写入完成 time.sleep(0.5) try: shutil.move(source_path, os.path.join(target_folder, file_name)) except Exception as e: print(f"移动文件失败: {e}") try: while True: # 读取系统事件 for event in inotify.read(): if event.mask & (flags.CREATE | flags.MOVED_TO): process_new_file(event.name) except KeyboardInterrupt: # 清理监听 inotify.rm_watch(wd)
注:需要先安装依赖库:pip install inotify-simple,如果不想用第三方库,可以通过ctypes直接调用Linux的inotify底层API,但代码会更繁琐。
内容的提问来源于stack exchange,提问作者Jugert Mucoimaj
相关产品推荐
相关产品推荐

