如何从文本文件中随机移除100个指定格式的数据块?
嘿,针对你要从大型文本里随机移除100个「数字开头+空行结尾」的数据块的需求,我给你整理了两个实用方案,分别适合不同环境~
方法1:用Unix/Linux命令组合(高效快捷)
这个方案适合在Linux、MacOS或者Windows下用WSL/Git Bash的场景,全程用命令行工具搞定,不用写代码:
给每个数据块编号
首先把每个“数字行+空行”的块标记上序号,方便后续筛选删除:awk -v RS='' -v ORS='\n\n' '{print NR ":" $0}' input.txt > numbered_blocks.txt这里
RS=''让awk把空行作为数据块的分隔符,ORS='\n\n'保证输出后每个块依然保留结尾的空行,NR是当前块的序号。生成要删除的随机块序号
先统计总共有多少个数据块,再随机选出100个要删除的序号:shuf -i 1-$(awk -v RS='' 'END{print NR}' input.txt) -n 100 > to_delete.txtawk -v RS='' 'END{print NR}' input.txt用来统计总块数,shuf则从1到总块数里随机挑100个数字。过滤并生成最终文件
把不在删除列表里的数据块提取出来,去掉之前加的序号:awk -F: 'NR==FNR{del[$1]=1; next} !del[$1]' to_delete.txt numbered_blocks.txt | cut -d: -f2- > output.txt完成后你可以删掉
numbered_blocks.txt和to_delete.txt这两个临时文件。
方法2:Python脚本(跨平台,灵活适配)
如果需要跨平台运行,或者你的数据结构有特殊情况,用Python脚本会更灵活,而且不会一次性把大文件加载到内存,适合超大文件:
import random def process_large_file(input_path, output_path, delete_count=100): # 第一步:先遍历一次文件,统计符合要求的数据块总数 block_count = 0 current_lines = [] with open(input_path, 'r', encoding='utf-8') as f: for line in f: stripped_line = line.rstrip('\n') # 遇到空行,说明当前数据块结束 if stripped_line == '': if current_lines and current_lines[0][0].isdigit(): block_count += 1 current_lines = [] else: # 收集数字开头的行(这里假设每个数据块只有一行,有需要可以调整) if stripped_line[0].isdigit(): current_lines.append(stripped_line) # 处理文件末尾可能没有空行的最后一个数据块 if current_lines and current_lines[0][0].isdigit(): block_count += 1 # 生成要删除的随机块序号(从0开始计数) if block_count <= delete_count: raise ValueError("数据块总数不足100个,无法删除!") to_delete = set(random.sample(range(block_count), delete_count)) # 第二步:再次遍历文件,保留不需要删除的数据块 current_block_idx = 0 current_lines = [] with open(input_path, 'r', encoding='utf-8') as f_in, open(output_path, 'w', encoding='utf-8') as f_out: for line in f_in: stripped_line = line.rstrip('\n') if stripped_line == '': if current_lines and current_lines[0][0].isdigit(): # 如果当前块不在删除列表里,就写入文件 if current_block_idx not in to_delete: f_out.write('\n'.join(current_lines) + '\n\n') current_block_idx += 1 current_lines = [] else: if stripped_line[0].isdigit(): current_lines.append(stripped_line) # 处理最后一个无空行结尾的数据块 if current_lines and current_lines[0][0].isdigit(): if current_block_idx not in to_delete: f_out.write('\n'.join(current_lines) + '\n\n') if __name__ == '__main__': # 替换成你的输入文件和输出文件路径 process_large_file('input.txt', 'output.txt', 100)
额外提醒
- 操作前一定要备份原文件,避免误操作导致数据丢失
- 如果你的数据块里有连续多个空行,方法1的
awk命令依然能正常处理,因为RS=''会把多个空行当成一个分隔符 - 方法2的脚本默认每个数据块只有一行数字开头的内容,如果你的数据块是多行(但开头行是数字,结尾是空行),可以稍微调整
current_lines的收集逻辑
内容的提问来源于stack exchange,提问作者felix
相关产品推荐
相关产品推荐

