如何用Python获取最新两个文件及提取最新两个AMI ID?
Hey there! Let's work through your technical needs and that great question you had about using epoch timestamps in filenames.
You’ve got two solid options here—using Python’s built-in pathlib (modern, clean) or the older os module. Both will fetch files sorted by their modification time, then grab the top two:
Using pathlib (Recommended for Python 3.4+)
from pathlib import Path def get_latest_two_files(target_dir): # Grab only files (skip subdirectories) all_files = [file for file in Path(target_dir).iterdir() if file.is_file()] # Sort files by last modified time, newest first sorted_files = sorted(all_files, key=lambda f: f.stat().st_mtime, reverse=True) # Return the top two return sorted_files[:2] # Example usage latest_files = get_latest_two_files("/path/to/your/directory") for file in latest_files: print(f"Latest file: {file.resolve()}")
Using os Module
import os def get_latest_two_files_os(target_dir): # Build full paths for all files in the directory file_paths = [os.path.join(target_dir, filename) for filename in os.listdir(target_dir) if os.path.isfile(os.path.join(target_dir, filename))] # Sort by modification time sorted_files = sorted(file_paths, key=lambda f: os.path.getmtime(f), reverse=True) return sorted_files[:2]
This depends on where your output content is coming from—here are the two most common scenarios:
Scenario A: Parsing Raw Text Output
If you have text output (e.g., logs, CLI stdout) with lines like ami-x1 2024-05-22 16:45:00, use regex to extract AMI IDs and their timestamps, then sort:
import re from datetime import datetime # Sample output content (replace with your actual data) output_content = """ ami-x1 2024-05-22 16:45:00 ami-x3 2024-05-20 10:20:00 ami-x9 2024-05-22 17:10:00 ami-x5 2024-05-19 09:30:00 """ # Regex to match AMI ID and timestamp ami_pattern = r'(ami-[a-z0-9]+)\s+(\d{4}-\d{2}-\d{2}\s+\d{2}:\d{2}:\d{2})' matches = re.findall(ami_pattern, output_content) # Convert to (datetime object, AMI ID) pairs and sort newest first ami_with_timestamps = [(datetime.strptime(ts, "%Y-%m-%d %H:%M:%S"), ami) for ami, ts in matches] sorted_amis = sorted(ami_with_timestamps, key=lambda x: x[0], reverse=True) # Grab the top two AMI IDs latest_two_amis = [ami for _, ami in sorted_amis[:2]] print(f"Latest two AMIs: {latest_two_amis}") # Output: ['ami-x9', 'ami-x1']
Scenario B: Fetching Directly from AWS (Most Accurate)
If you’re pulling AMI data from AWS, use boto3 to query the API with built-in sorting—no parsing needed:
import boto3 # Initialize EC2 client ec2 = boto3.client('ec2') # Fetch images sorted by creation date (newest first), limit to 2 response = ec2.describe_images( Owners=['self'], # Replace with your AWS account ID or 'amazon' for official images Filters=[{'Name': 'state', 'Values': ['available']}], SortBy='creationDate', SortOrder='descending', MaxResults=2 ) # Extract AMI IDs latest_two_amis = [image['ImageId'] for image in response['Images']] print(f"Latest two AMIs from AWS: {latest_two_amis}")
Absolutely—this is a great approach, and here’s why:
- Faster & Less IO: You don’t need to read file metadata (like modification time) via
stat()—just parse the timestamp directly from the filename. - No DateTime Parsing Headaches: Epoch timestamps are integers, so sorting is as simple as comparing numbers (no dealing with string formats, time zones, or parsing errors).
- More Reliable: File modification times can get altered (e.g., when copying/moving files), but the epoch in the filename is tied to when the AMI was generated—so it’s a true reflection of creation order.
Example Implementation for Epoch-Named Files
Assuming your filenames look like ami-x1_1716385500.id (AMI ID + underscore + epoch timestamp + .id):
from pathlib import Path def get_latest_two_amis_from_epoch_filenames(target_dir): # Grab all .id files id_files = [file for file in Path(target_dir).iterdir() if file.suffix == '.id'] ami_epoch_pairs = [] for file in id_files: # Split the filename to extract AMI ID and epoch stem = file.stem # Removes the .id suffix, e.g., "ami-x1_1716385500" ami_id, epoch_str = stem.split('_') ami_epoch_pairs.append( (-int(epoch_str), ami_id) ) # Negative for reverse sort # Sort (negative epoch puts newest first) ami_epoch_pairs.sort() # Grab the top two AMIs latest_two = [ami for _, ami in ami_epoch_pairs[:2]] return latest_two # Example usage latest_amis = get_latest_two_amis_from_epoch_filenames("/path/to/your/id/files") print(f"Latest two AMIs from epoch filenames: {latest_amis}")
- If you can generate filenames with epoch timestamps, use that method for extracting the latest AMIs—it’s the most efficient and reliable.
- When fetching AMIs directly from AWS, always use
boto3with theSortBy='creationDate'parameter—this avoids any parsing errors and gives you the most accurate results. - For general file sorting,
pathlibis cleaner and more readable than theosmodule.
内容的提问来源于stack exchange,提问作者user6481062

