文本文件二级字段排序:按样本数排序并保留raw1目录顺序
解决目录内CSV文件按样本数升序排序的问题
核心思路
先按raw1/100、raw1/101这类目录分组,严格保留原目录的出现顺序,再对每个目录下的条目,提取文件名中csv后的数字(如10、100)并转为整数,以此为依据升序排序,最后将各组按原目录顺序拼接输出。
Python 自动化实现脚本
假设你的结果文件每行格式为文件路径,对应数值(如raw1/100/csv10_samples_f1.json,0.78),可以用以下脚本处理:
import os import re from collections import OrderedDict # 配置输入输出文件路径 input_file = "你的输入结果文件.txt" output_file = "排序后的结果文件.txt" # 按目录分组,保留目录顺序 dir_groups = OrderedDict() with open(input_file, 'r', encoding='utf-8') as f: for line in f: line = line.strip() if not line: continue # 分割路径和数值(如果你的分隔符不是逗号,修改这里的分隔符) path, value = line.split(',', 1) dir_name = os.path.dirname(path) # 将条目加入对应目录组 if dir_name not in dir_groups: dir_groups[dir_name] = [] dir_groups[dir_name].append((path, value)) # 对每个目录下的条目按样本数排序 sorted_lines = [] for dir_name, entries in dir_groups.items(): # 提取csv后的数字作为排序key def sort_key(entry): path = entry[0] # 匹配csv后的数字,适配你的文件名格式 match = re.search(r'csv(\d+)_samples_f1.json', path) if match: return int(match.group(1)) return 0 # 匹配失败的条目放最后 # 升序排序 sorted_entries = sorted(entries, key=sort_key) # 转成原格式的行 sorted_lines.extend([f"{path},{value}" for path, value in sorted_entries]) # 写入结果文件 with open(output_file, 'w', encoding='utf-8') as f: f.write('\n'.join(sorted_lines))
使用说明
- 修改
input_file和output_file为你的实际文件路径 - 如果你的行格式不是
路径,数值(比如用空格分隔),修改split(',', 1)中的分隔符 - 如果文件名格式有变化,调整正则表达式
r'csv(\d+)_samples_f1.json',确保能正确提取样本数
备选Shell实现(适用于Linux/macOS)
如果习惯用命令行,可以用awk结合sort处理,同样保留目录顺序:
# 假设输入文件为input.txt,输出为output.txt awk -F',' '{ dir = substr($1, 1, index($1, "/csv")-1); match($1, /csv([0-9]+)_/, arr); num = arr[1]; print dir "\t" num "\t" $0; }' input.txt | sort -k1,1 -k2n | cut -f3- > output.txt
- 原理:先给每行加上目录和样本数字段,按目录字符串排序(保留原顺序)+ 样本数数值排序,最后去掉临时字段
内容的提问来源于stack exchange,提问作者JohnJ
相关产品推荐
相关产品推荐

