You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中按图片文件名相同合并文本文件行的实现求助

问题描述

我有一个格式如下的文本文件:

0.jpg 12,13,14,15,16
0.jpg 13,14,15,16,17
1.jpg 1,2,3,4,5
1.jpg 2,3,4,5,6

需要将文件名相同的行合并,输出格式如下:

0.jpg 12,13,14,15,16 13,14,15,16,17
1.jpg 1,2,3,4,5 2,3,4,5,6

我尝试了两段代码,但不知道如何正确对比文件名并实现合并逻辑:

第一段尝试代码:

with open("file.txt", "r") as input:       # Read all data lines.
    data = input.readlines()
with open("out_file.txt", "w") as output:  # Create output file.
    for line in data:                      # Iterate over data lines.
        line_elements = line.split()       # Split line by spaces.
        line_updated = [line_elements[0]]  # Initialize fixed line (without undesired patterns) with image's name.
        if line_elements[0] = (next line's line_elements[0])???:
            for i in line_elements[1:]:    # Iterate over groups of numbers in current line.
               tmp = i.split(',')          # Split current group by commas.
               if len(tmp) == 5:
                  line_updated.append(','.join(tmp))

            if len(line_updated) > 1:      # If the fixed line is valid, write it to output file.
               output.write(f"{' '.join(line_updated)}\n")

第二段尝试代码:

for i in range (len(data)):
if line_elements[0] in line[i] == line_elements[0] in line[i+1]:

   line_updated = [line_elements[0]]
   for i in line_elements[1:]:    # Iterate over groups of numbers in current line.
      tmp = i.split(',')          # Split current group by commas.
      if len(tmp) == 5:
         line_updated.append(','.join(tmp))

   if len(line_updated) > 1:      # If the fixed line is valid, write it to output file.
      output.write(f"{' '.join(line_updated)}\n")
解决方案

核心思路是用字典存储每个文件名对应的所有数字组:键为文件名,值为该文件对应的数字组列表。遍历每一行时将数字组添加到对应文件名的列表中,最后把字典内容按要求格式写入文件。

完整实现代码:

# 初始化字典存储文件名和对应的数字组
file_data = {}

# 读取输入文件
with open("file.txt", "r") as input_file:
    for line in input_file:
        line = line.strip()  # 去掉换行符和首尾空格
        if not line:
            continue  # 跳过空行
        elements = line.split()
        if len(elements) < 2:
            continue  # 跳过格式不正确的行
        filename = elements[0]
        num_group = elements[1]
        # 验证数字组格式是否为5个数字(可选,按需保留)
        tmp = num_group.split(',')
        if len(tmp) != 5:
            continue
        # 将数字组添加到对应文件名的列表中
        if filename not in file_data:
            file_data[filename] = []
        file_data[filename].append(num_group)

# 写入输出文件
with open("out_file.txt", "w") as output_file:
    for filename, num_groups in file_data.items():
        # 拼接成要求的格式:文件名 + 空格分隔的数字组
        output_line = f"{filename} {' '.join(num_groups)}\n"
        output_file.write(output_line)

代码说明

  1. 字典分组:用file_data字典天然实现文件名的分组,避免逐行对比的复杂逻辑。
  2. 行过滤:清理空行和格式错误的行,仅保留有效数据。
  3. 格式验证:保留了你原代码中对数字组长度的校验,确保只处理符合要求的5个数字组合。
  4. 结果写入:遍历字典键值对,按指定格式拼接内容并写入输出文件。

内容的提问来源于stack exchange,提问作者Vanq

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 11:05:17