如何用Python/Unix命令按首词分组打乱大型文本数据集?
没问题!我帮你整理了两种靠谱的实现方法——Python脚本和Unix命令行工具组合,都能满足「按首词分组打乱组顺序、组内行保持原序」的需求:
Python 实现方案
核心思路是先按首词把所有行分组,再随机打乱分组的顺序,最后按打乱后的顺序输出每个组的所有行。
假设你的输入是每行一条记录(比如 input.txt 里每行是 "ABC", 21, 15, 45 这样的格式),可以用下面的脚本:
import random from collections import defaultdict # 读取输入文件并分组 groups = defaultdict(list) with open('input.txt', 'r') as input_file: for line in input_file: cleaned_line = line.strip() if not cleaned_line: continue # 提取首词:去掉引号和空格,取第一个逗号前的内容 first_word = cleaned_line.split(',')[0].strip().strip('"') groups[first_word].append(cleaned_line) # 随机打乱分组的顺序 shuffled_groups = list(groups.keys()) random.shuffle(shuffled_groups) # 写入输出文件 with open('output.txt', 'w') as output_file: for group_key in shuffled_groups: for line in groups[group_key]: output_file.write(f"{line}\n")
如果需要从标准输入读取、标准输出写入(比如配合管道使用),可以把文件读写换成 sys.stdin 和 sys.stdout:
import sys import random from collections import defaultdict groups = defaultdict(list) for line in sys.stdin: cleaned_line = line.strip() if not cleaned_line: continue first_word = cleaned_line.split(',')[0].strip().strip('"') groups[first_word].append(cleaned_line) shuffled_groups = list(groups.keys()) random.shuffle(shuffled_groups) for group_key in shuffled_groups: for line in groups[group_key]: print(line)
Unix 命令行实现方案
如果习惯用命令行工具处理文本,推荐用 awk + shuf 的组合,简洁高效:
# 假设输入文件是 input.txt,输出直接打印到终端 awk -F ',' '{ key = $1; gsub(/[" ]/, "", key) if (prev_key != key && NR > 1) print "\x00" prev_key = key print $0 }' input.txt | shuf --separator=$'\x00' | tr -d '\x00'
步骤解释:
awk预处理:提取每行的首词作为分组key,当分组切换时插入一个特殊分隔符\x00(避免和文本内容冲突),把同一组的行连在一起。shuf打乱分组:用--separator参数指定按\x00分隔的块(也就是每个分组)随机打乱顺序。tr清理分隔符:去掉插入的\x00,得到最终的分组打乱结果。
如果需要把结果写入文件,直接加重定向即可:
# 输出到 output.txt awk -F ',' '{ key = $1; gsub(/[" ]/, "", key) if (prev_key != key && NR > 1) print "\x00" prev_key = key print $0 }' input.txt | shuf --separator=$'\x00' | tr -d '\x00' > output.txt
内容的提问来源于stack exchange,提问作者Touhidul Alam
相关产品推荐
相关产品推荐

