Python中按图片文件名相同合并文本文件行的实现求助
问题描述
我有一个格式如下的文本文件:
0.jpg 12,13,14,15,16 0.jpg 13,14,15,16,17 1.jpg 1,2,3,4,5 1.jpg 2,3,4,5,6
需要将文件名相同的行合并,输出格式如下:
0.jpg 12,13,14,15,16 13,14,15,16,17 1.jpg 1,2,3,4,5 2,3,4,5,6
我尝试了两段代码,但不知道如何正确对比文件名并实现合并逻辑:
第一段尝试代码:
with open("file.txt", "r") as input: # Read all data lines. data = input.readlines() with open("out_file.txt", "w") as output: # Create output file. for line in data: # Iterate over data lines. line_elements = line.split() # Split line by spaces. line_updated = [line_elements[0]] # Initialize fixed line (without undesired patterns) with image's name. if line_elements[0] = (next line's line_elements[0])???: for i in line_elements[1:]: # Iterate over groups of numbers in current line. tmp = i.split(',') # Split current group by commas. if len(tmp) == 5: line_updated.append(','.join(tmp)) if len(line_updated) > 1: # If the fixed line is valid, write it to output file. output.write(f"{' '.join(line_updated)}\n")
第二段尝试代码:
for i in range (len(data)): if line_elements[0] in line[i] == line_elements[0] in line[i+1]: line_updated = [line_elements[0]] for i in line_elements[1:]: # Iterate over groups of numbers in current line. tmp = i.split(',') # Split current group by commas. if len(tmp) == 5: line_updated.append(','.join(tmp)) if len(line_updated) > 1: # If the fixed line is valid, write it to output file. output.write(f"{' '.join(line_updated)}\n")
解决方案
核心思路是用字典存储每个文件名对应的所有数字组:键为文件名,值为该文件对应的数字组列表。遍历每一行时将数字组添加到对应文件名的列表中,最后把字典内容按要求格式写入文件。
完整实现代码:
# 初始化字典存储文件名和对应的数字组 file_data = {} # 读取输入文件 with open("file.txt", "r") as input_file: for line in input_file: line = line.strip() # 去掉换行符和首尾空格 if not line: continue # 跳过空行 elements = line.split() if len(elements) < 2: continue # 跳过格式不正确的行 filename = elements[0] num_group = elements[1] # 验证数字组格式是否为5个数字(可选,按需保留) tmp = num_group.split(',') if len(tmp) != 5: continue # 将数字组添加到对应文件名的列表中 if filename not in file_data: file_data[filename] = [] file_data[filename].append(num_group) # 写入输出文件 with open("out_file.txt", "w") as output_file: for filename, num_groups in file_data.items(): # 拼接成要求的格式:文件名 + 空格分隔的数字组 output_line = f"{filename} {' '.join(num_groups)}\n" output_file.write(output_line)
代码说明
- 字典分组:用
file_data字典天然实现文件名的分组,避免逐行对比的复杂逻辑。 - 行过滤:清理空行和格式错误的行,仅保留有效数据。
- 格式验证:保留了你原代码中对数字组长度的校验,确保只处理符合要求的5个数字组合。
- 结果写入:遍历字典键值对,按指定格式拼接内容并写入输出文件。
内容的提问来源于stack exchange,提问作者Vanq
相关产品推荐
相关产品推荐

