You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

修改Python索引函数生成子列表并关联至DataFrame的方法

Solution: Grouping Image Indices with Their Duplicates

First, let's adjust your function to generate a list of sublists where each sublist contains all indices for a specific original image (including its duplicates marked with rep_). The key change is to iterate over each original image in rep_list first, then collect all matching indices for that image instead of building a flat list.

Modified Function (Based on Your Original Logic)

def get_grouped_indices(rep_list, img_labels):
    grouped_indices = []
    for original_img in rep_list:
        # Initialize a sublist for this original image's indices
        img_indices = []
        for idx, label in enumerate(img_labels):
            # Check if label is the original or a rep_ variant
            # Note: Updated the condition to ensure exact match after 'rep_'
            if label == original_img or (label.startswith('rep_') and label[4:] == original_img):
                img_indices.append(idx)
        grouped_indices.append(img_indices)
    return grouped_indices

Key Changes from Your Original Code:

  • Instead of a single flat list, we create a sublist for each image in rep_list.
  • Fixed the duplicate check: label[4:] == original_img ensures that the part after rep_ exactly matches the original image name (your original endswith check could accidentally match partial suffixes, e.g., rep_img12 would be considered a duplicate of img1 otherwise).
  • The function now returns a list of sublists, where each sublist corresponds to an image in rep_list.

Adding to Your DataFrame

Assuming your DataFrame (let's call it duplicate_imgs_df) has one row per original duplicate image (113 rows total), you can add the grouped indices as a new column like this:

# Generate the grouped indices list
indices_list = get_grouped_indices(duplicate_imgs_df['original_image_column'].tolist(), img_labels)

# Add as a new column to the DataFrame
duplicate_imgs_df['image_indices'] = indices_list

Replace 'original_image_column' with the actual column name in your DataFrame that contains the 113 original image strings.

Optimized Version for Larger Datasets

If you're working with even larger lists (beyond 5k items), this optimized approach using a dictionary will be much faster (reduces nested loops to linear passes):

from collections import defaultdict

def get_grouped_indices_fast(rep_list, img_labels):
    # Create a set of original images for quick lookups
    original_images = set(rep_list)
    # Map each label to its original image (if applicable)
    label_to_original = {}
    
    for label in img_labels:
        if label in original_images:
            label_to_original[label] = label
        elif label.startswith('rep_'):
            candidate_original = label[4:]
            if candidate_original in original_images:
                label_to_original[label] = candidate_original
    
    # Group indices by original image
    index_groups = defaultdict(list)
    for idx, label in enumerate(img_labels):
        original = label_to_original.get(label)
        if original is not None:
            index_groups[original].append(idx)
    
    # Return groups in the same order as rep_list
    return [index_groups[img] for img in rep_list]

This version only loops through img_labels twice, making it significantly more efficient for large datasets.

Content of the question originated from Stack Exchange, asked by Cameron Lester

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 15:48:12