修改Python索引函数生成子列表并关联至DataFrame的方法
First, let's adjust your function to generate a list of sublists where each sublist contains all indices for a specific original image (including its duplicates marked with rep_). The key change is to iterate over each original image in rep_list first, then collect all matching indices for that image instead of building a flat list.
Modified Function (Based on Your Original Logic)
def get_grouped_indices(rep_list, img_labels): grouped_indices = [] for original_img in rep_list: # Initialize a sublist for this original image's indices img_indices = [] for idx, label in enumerate(img_labels): # Check if label is the original or a rep_ variant # Note: Updated the condition to ensure exact match after 'rep_' if label == original_img or (label.startswith('rep_') and label[4:] == original_img): img_indices.append(idx) grouped_indices.append(img_indices) return grouped_indices
Key Changes from Your Original Code:
- Instead of a single flat list, we create a sublist for each image in
rep_list. - Fixed the duplicate check:
label[4:] == original_imgensures that the part afterrep_exactly matches the original image name (your originalendswithcheck could accidentally match partial suffixes, e.g.,rep_img12would be considered a duplicate ofimg1otherwise). - The function now returns a list of sublists, where each sublist corresponds to an image in
rep_list.
Adding to Your DataFrame
Assuming your DataFrame (let's call it duplicate_imgs_df) has one row per original duplicate image (113 rows total), you can add the grouped indices as a new column like this:
# Generate the grouped indices list indices_list = get_grouped_indices(duplicate_imgs_df['original_image_column'].tolist(), img_labels) # Add as a new column to the DataFrame duplicate_imgs_df['image_indices'] = indices_list
Replace 'original_image_column' with the actual column name in your DataFrame that contains the 113 original image strings.
Optimized Version for Larger Datasets
If you're working with even larger lists (beyond 5k items), this optimized approach using a dictionary will be much faster (reduces nested loops to linear passes):
from collections import defaultdict def get_grouped_indices_fast(rep_list, img_labels): # Create a set of original images for quick lookups original_images = set(rep_list) # Map each label to its original image (if applicable) label_to_original = {} for label in img_labels: if label in original_images: label_to_original[label] = label elif label.startswith('rep_'): candidate_original = label[4:] if candidate_original in original_images: label_to_original[label] = candidate_original # Group indices by original image index_groups = defaultdict(list) for idx, label in enumerate(img_labels): original = label_to_original.get(label) if original is not None: index_groups[original].append(idx) # Return groups in the same order as rep_list return [index_groups[img] for img in rep_list]
This version only loops through img_labels twice, making it significantly more efficient for large datasets.
Content of the question originated from Stack Exchange, asked by Cameron Lester

