如何用Python读取目录中对应视频的JPG文件并按规则排序
Got it, let's break down how to solve this problem cleanly. You need to sort your JPG frame files by their video identifier (everything except the final number in the filename) and then group them into a dictionary for easy processing. Here's a step-by-step approach:
Step 1: Filter and Sort Files
First, we'll filter out only JPG files, then define a custom sorting key to order them by video ID and frame number. We can use either string splitting or regular expressions—both work, but regex is more robust if your filename structure might have minor variations.
Option 1: Using String Splitting (Simple, for strict filename formats)
This works perfectly if all your filenames follow the exact structure you provided:
import os # Replace with your target directory path target_path = "/path/to/your/jpg/files" # First, get all JPG files (case-insensitive) jpg_files = [file for file in os.listdir(target_path) if file.lower().endswith(".jpg")] def sort_key(filename): # Remove file extension name_no_ext = filename.split(".")[0] # Split the name by underscores parts = name_no_ext.split("_") # Extract video ID (all parts except the frame number and trailing empty string) video_id = "_".join(parts[:-2]) + "_" # Convert frame number to integer for numerical sorting frame_number = int(parts[-2]) # Sort first by video ID, then by frame number return (video_id, frame_number) # Sort the files using our custom key sorted_files = sorted(jpg_files, key=sort_key)
Option 2: Using Regular Expressions (Robust for varied formats)
If you want to handle small inconsistencies in filenames, regex is a better choice. It explicitly captures the video identifier and frame number:
import os import re target_path = "/path/to/your/jpg/files" jpg_files = [file for file in os.listdir(target_path) if file.lower().endswith(".jpg")] # Regex pattern to match your filename structure: captures video ID and frame number filename_pattern = re.compile(r"(img_\d+_.+?)_(\d+)_\.jpg") def sort_key(filename): match = filename_pattern.match(filename) if match: # Reconstruct the full video ID (add back the trailing underscore) video_id = match.group(1) + "_" frame_number = int(match.group(2)) return (video_id, frame_number) # Push invalid files to the end of the sorted list return ("", float("inf")) sorted_files = sorted(jpg_files, key=sort_key)
Step 2: Group into a Dictionary
Now that the files are sorted, we can organize them into a dictionary where each key is a video identifier, and the value is a list of sorted frame file paths:
video_frame_dict = {} for file in sorted_files: # Extract video ID using the same method as your sort key name_no_ext = file.split(".")[0] parts = name_no_ext.split("_") video_id = "_".join(parts[:-2]) + "_" # Get full path to the file full_file_path = os.path.join(target_path, file) # Add to the dictionary (create list if video ID doesn't exist yet) if video_id not in video_frame_dict: video_frame_dict[video_id] = [] video_frame_dict[video_id].append(full_file_path) # Alternative one-liner using dict.setdefault() # video_frame_dict = {} # for file in sorted_files: # name_no_ext = file.split(".")[0] # parts = name_no_ext.split("_") # video_id = "_".join(parts[:-2]) + "_" # video_frame_dict.setdefault(video_id, []).append(os.path.join(target_path, file))
Key Notes
- Case Insensitivity: We use
file.lower().endswith(".jpg")to handle both.jpgand.JPGfiles. - Numerical Sorting: Converting the frame number to an integer ensures frames are sorted numerically (1, 2, 3...) instead of lexicographically (1, 10, 2...).
- Invalid Files: The regex approach pushes any files that don't match your format to the end of the sorted list—you can also filter them out entirely by adding
and filename_pattern.match(file)to the initialjpg_fileslist comprehension.
内容的提问来源于stack exchange,提问作者Jose Ramon

