You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

读取图片路径与标签文件,去重标签后合并生成新文件的方法

Solution to Combine Image Paths with Deduplicated Labels

Got it, let's walk through how to solve this problem step by step using Python—it's simple, readable, and handles exactly what you need.

Step-by-Step Breakdown

  • Read both input files: We'll read lines from the image path file and label file in parallel to keep them paired correctly.
  • Clean up the labels: For each line of labels, split the string into individual labels, remove duplicates (while preserving the original order of first occurrence), then rejoin them into a clean string.
  • Combine and write output: Pair each cleaned label string with its corresponding image path, then write the result to a new file.

Full Python Code

# Define your file paths (update these to match your actual file locations)
image_paths_file = "image_paths.txt"
labels_file = "labels.txt"
output_file = "combined_output.txt"

# Open all files in one context manager to handle resources properly
with open(image_paths_file, 'r') as img_file, \
     open(labels_file, 'r') as lbl_file, \
     open(output_file, 'w') as out_file:
    
    # Iterate over corresponding lines from both input files
    for img_line, lbl_line in zip(img_file, lbl_file):
        # Remove any leading/trailing whitespace (including newlines)
        img_path = img_line.strip()
        raw_label_str = lbl_line.strip()
        
        # Split the label string into a list of individual labels
        label_list = raw_label_str.split(', ')
        
        # Remove duplicates while preserving the original order (Python 3.7+)
        # Using dict.fromkeys() since it retains insertion order
        unique_labels = list(dict.fromkeys(label_list))
        
        # Rejoin the unique labels into a single string
        cleaned_labels = ', '.join(unique_labels)
        
        # Combine image path and cleaned labels, then write to output
        output_line = f"{img_path} {cleaned_labels}\n"
        out_file.write(output_line)

Key Details Explained

  • Parallel line reading: Using zip() ensures we process the first line of the image file with the first line of the label file, and so on.
  • Order-preserving deduplication: dict.fromkeys(label_list) creates a dictionary where keys are the labels (automatically removing duplicates), and since Python 3.7+ dictionaries preserve insertion order, we get the first occurrence of each label in the original order. If you don't care about order, you could use list(set(label_list)) instead—but that will shuffle the labels.
  • Whitespace handling: strip() removes any extra newlines or leading/trailing spaces from each line, which prevents messy output.

Example Output

When run with your sample input files, the combined_output.txt will contain exactly what you requested:

/home/picture/I10045.jpg Cat, Dog
/home/picture/I10056.jpg Mouse

Note

This code assumes both input files have the exact same number of lines. If there's a chance they don't, you might want to add a check to handle mismatched line counts (e.g., log a warning or pad missing entries).

内容的提问来源于stack exchange,提问作者Ken

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:19:27