You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python/Unix命令按首词分组打乱大型文本数据集?

没问题!我帮你整理了两种靠谱的实现方法——Python脚本和Unix命令行工具组合,都能满足「按首词分组打乱组顺序、组内行保持原序」的需求:

Python 实现方案

核心思路是先按首词把所有行分组,再随机打乱分组的顺序,最后按打乱后的顺序输出每个组的所有行。

假设你的输入是每行一条记录(比如 input.txt 里每行是 "ABC", 21, 15, 45 这样的格式),可以用下面的脚本:

import random
from collections import defaultdict

# 读取输入文件并分组
groups = defaultdict(list)
with open('input.txt', 'r') as input_file:
    for line in input_file:
        cleaned_line = line.strip()
        if not cleaned_line:
            continue
        # 提取首词:去掉引号和空格,取第一个逗号前的内容
        first_word = cleaned_line.split(',')[0].strip().strip('"')
        groups[first_word].append(cleaned_line)

# 随机打乱分组的顺序
shuffled_groups = list(groups.keys())
random.shuffle(shuffled_groups)

# 写入输出文件
with open('output.txt', 'w') as output_file:
    for group_key in shuffled_groups:
        for line in groups[group_key]:
            output_file.write(f"{line}\n")

如果需要从标准输入读取、标准输出写入(比如配合管道使用),可以把文件读写换成 sys.stdin 和 sys.stdout:

import sys
import random
from collections import defaultdict

groups = defaultdict(list)
for line in sys.stdin:
    cleaned_line = line.strip()
    if not cleaned_line:
        continue
    first_word = cleaned_line.split(',')[0].strip().strip('"')
    groups[first_word].append(cleaned_line)

shuffled_groups = list(groups.keys())
random.shuffle(shuffled_groups)

for group_key in shuffled_groups:
    for line in groups[group_key]:
        print(line)
Unix 命令行实现方案

如果习惯用命令行工具处理文本,推荐用 awk + shuf 的组合,简洁高效:

# 假设输入文件是 input.txt,输出直接打印到终端
awk -F ',' '{
    key = $1; gsub(/[" ]/, "", key)
    if (prev_key != key && NR > 1) print "\x00"
    prev_key = key
    print $0
}' input.txt | shuf --separator=$'\x00' | tr -d '\x00'

步骤解释:

  1. awk 预处理:提取每行的首词作为分组key,当分组切换时插入一个特殊分隔符 \x00(避免和文本内容冲突),把同一组的行连在一起。
  2. shuf 打乱分组:用 --separator 参数指定按 \x00 分隔的块(也就是每个分组)随机打乱顺序。
  3. tr 清理分隔符:去掉插入的 \x00,得到最终的分组打乱结果。

如果需要把结果写入文件,直接加重定向即可:

# 输出到 output.txt
awk -F ',' '{
    key = $1; gsub(/[" ]/, "", key)
    if (prev_key != key && NR > 1) print "\x00"
    prev_key = key
    print $0
}' input.txt | shuf --separator=$'\x00' | tr -d '\x00' > output.txt

内容的提问来源于stack exchange,提问作者Touhidul Alam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:43:48