You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从文本文件中随机移除100个指定格式的数据块?

嘿,针对你要从大型文本里随机移除100个「数字开头+空行结尾」的数据块的需求,我给你整理了两个实用方案,分别适合不同环境~

方法1:用Unix/Linux命令组合(高效快捷)

这个方案适合在Linux、MacOS或者Windows下用WSL/Git Bash的场景,全程用命令行工具搞定,不用写代码:

  1. 给每个数据块编号
    首先把每个“数字行+空行”的块标记上序号,方便后续筛选删除:

    awk -v RS='' -v ORS='\n\n' '{print NR ":" $0}' input.txt > numbered_blocks.txt
    

    这里RS=''让awk把空行作为数据块的分隔符,ORS='\n\n'保证输出后每个块依然保留结尾的空行,NR是当前块的序号。

  2. 生成要删除的随机块序号
    先统计总共有多少个数据块,再随机选出100个要删除的序号:

    shuf -i 1-$(awk -v RS='' 'END{print NR}' input.txt) -n 100 > to_delete.txt
    

    awk -v RS='' 'END{print NR}' input.txt用来统计总块数,shuf则从1到总块数里随机挑100个数字。

  3. 过滤并生成最终文件
    把不在删除列表里的数据块提取出来,去掉之前加的序号:

    awk -F: 'NR==FNR{del[$1]=1; next} !del[$1]' to_delete.txt numbered_blocks.txt | cut -d: -f2- > output.txt
    

    完成后你可以删掉numbered_blocks.txt和to_delete.txt这两个临时文件。

方法2:Python脚本(跨平台,灵活适配)

如果需要跨平台运行,或者你的数据结构有特殊情况,用Python脚本会更灵活,而且不会一次性把大文件加载到内存,适合超大文件:

import random

def process_large_file(input_path, output_path, delete_count=100):
    # 第一步:先遍历一次文件,统计符合要求的数据块总数
    block_count = 0
    current_lines = []
    with open(input_path, 'r', encoding='utf-8') as f:
        for line in f:
            stripped_line = line.rstrip('\n')
            # 遇到空行,说明当前数据块结束
            if stripped_line == '':
                if current_lines and current_lines[0][0].isdigit():
                    block_count += 1
                    current_lines = []
            else:
                # 收集数字开头的行(这里假设每个数据块只有一行,有需要可以调整)
                if stripped_line[0].isdigit():
                    current_lines.append(stripped_line)
        # 处理文件末尾可能没有空行的最后一个数据块
        if current_lines and current_lines[0][0].isdigit():
            block_count += 1

    # 生成要删除的随机块序号(从0开始计数)
    if block_count <= delete_count:
        raise ValueError("数据块总数不足100个,无法删除!")
    to_delete = set(random.sample(range(block_count), delete_count))

    # 第二步:再次遍历文件,保留不需要删除的数据块
    current_block_idx = 0
    current_lines = []
    with open(input_path, 'r', encoding='utf-8') as f_in, open(output_path, 'w', encoding='utf-8') as f_out:
        for line in f_in:
            stripped_line = line.rstrip('\n')
            if stripped_line == '':
                if current_lines and current_lines[0][0].isdigit():
                    # 如果当前块不在删除列表里,就写入文件
                    if current_block_idx not in to_delete:
                        f_out.write('\n'.join(current_lines) + '\n\n')
                    current_block_idx += 1
                    current_lines = []
            else:
                if stripped_line[0].isdigit():
                    current_lines.append(stripped_line)
        # 处理最后一个无空行结尾的数据块
        if current_lines and current_lines[0][0].isdigit():
            if current_block_idx not in to_delete:
                f_out.write('\n'.join(current_lines) + '\n\n')

if __name__ == '__main__':
    # 替换成你的输入文件和输出文件路径
    process_large_file('input.txt', 'output.txt', 100)

额外提醒

  • 操作前一定要备份原文件,避免误操作导致数据丢失
  • 如果你的数据块里有连续多个空行,方法1的awk命令依然能正常处理,因为RS=''会把多个空行当成一个分隔符
  • 方法2的脚本默认每个数据块只有一行数字开头的内容,如果你的数据块是多行(但开头行是数字,结尾是空行),可以稍微调整current_lines的收集逻辑

内容的提问来源于stack exchange,提问作者felix

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:10:43