如何每分钟检测RESULTS.ODB文件并验证写入完成后触发子进程?
Nice problem to solve! Let's walk through a practical implementation that covers all your requirements—timed checks, file size validation, ensuring the file isn't being written to, and spawning a subprocess to extract data. Here's how to do it:
1. Set Up Timed File Checks
First, you need a way to run your detection logic every minute. You have two main options:
Option A: Python Scheduler (Cross-Platform)
Use the schedule library to handle timing directly in your script. Install it first with pip install schedule, then use this skeleton:
import schedule import time import os def check_odb_files(): # All detection/validation logic goes here pass # Schedule the check to run every minute schedule.every(1).minutes.do(check_odb_files) # Keep the script running while True: schedule.run_pending() time.sleep(1)
Option B: System-Level Scheduling
For a more lightweight approach, use native system tools:
- Linux/macOS: Add a cron job with
crontab -eand insert:* * * * * /usr/bin/python3 /full/path/to/your/script.py - Windows: Use Task Scheduler to create a task that runs your script every minute.
2. Locate the Target File & Validate Size
Next, add logic to find RESULTS.ODB and check if it's larger than 1.5GB (1610612736 bytes):
TARGET_FILE = "RESULTS.ODB" MIN_SIZE = 1.5 * 1024**3 # 1.5GB in bytes def check_odb_files(): target_dir = "/path/to/your/target/directory" # Replace with your actual path for root, _, files in os.walk(target_dir): if TARGET_FILE in files: file_path = os.path.join(root, TARGET_FILE) file_size = os.path.getsize(file_path) if file_size >= MIN_SIZE: # Now verify the file is ready to read if is_file_ready(file_path): start_data_extraction(file_path)
3. Ensure the File Isn't Being Written To
This is critical—you don't want to read a partial file. The method varies by OS:
Linux/macOS
Use file locking or lsof to check for active writes:
Method 1: Exclusive File Lock
import fcntl def is_file_ready(file_path): try: with open(file_path, 'rb') as f: # Try to get an exclusive non-blocking lock fcntl.flock(f, fcntl.LOCK_EX | fcntl.LOCK_NB) return True except BlockingIOError: # File is locked by another process (being written) return False except Exception as e: print(f"Error checking file status: {str(e)}") return False
Method 2: Check with lsof
import subprocess def is_file_ready(file_path): try: # Check if any process has the file open for writing result = subprocess.run( ["lsof", "-w", "-F", "n", file_path], capture_output=True, text=True ) # If no output, the file is not in use return not result.stdout except Exception as e: print(f"Error running lsof: {str(e)}") return False
Windows
Use the pywin32 library to check for file locks, or verify stable modification times:
import win32file import win32con import time def is_file_ready(file_path): # Method 1: Try exclusive access try: handle = win32file.CreateFile( file_path, win32con.GENERIC_READ, 0, # No shared access None, win32con.OPEN_EXISTING, win32con.FILE_ATTRIBUTE_NORMAL, None ) win32file.CloseHandle(handle) return True except win32file.error as e: # Error 32 = file is in use by another process if e.winerror == 32: return False print(f"Error checking file: {str(e)}") return False # Method 2: Check if modification time is stable # first_mtime = os.path.getmtime(file_path) # time.sleep(2) # second_mtime = os.path.getmtime(file_path) # return first_mtime == second_mtime
4. Spawn a Subprocess to Extract Data
Use Python's subprocess module to launch your data extraction tool without blocking the main script:
import subprocess def start_data_extraction(file_path): # Replace with your actual extraction command/script extraction_cmd = [ "/path/to/your/extraction_tool", file_path, "--output", "/path/to/save/results" ] try: # Start the process in the background subprocess.Popen(extraction_cmd) print(f"Successfully started data extraction for {file_path}") except Exception as e: print(f"Failed to start extraction process: {str(e)}")
Extra Tips for Robustness
- Avoid Duplicate Runs: Track processed files (e.g., with a log file or simple text database) to prevent spawning multiple subprocesses for the same completed file.
- Add Logging: Use Python's
loggingmodule to record check times, file statuses, and subprocess outcomes—this makes debugging way easier. - Handle Permissions: Make sure the user running the script has read access to the target directory and execute permissions for the extraction tool.
- Limit Concurrent Processes: If extraction is resource-heavy, add a check to ensure only one extraction runs at a time.
内容的提问来源于stack exchange,提问作者Nils

