You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

50G大文件单次处理提取重复行与剩余内容方案求助

解决方案

方案1:Linux命令行实现(推荐,性能最优)

极简写法(支持进程替换的shell环境可用)

仅执行一次排序,通过tee和进程替换将排序结果同时分发给两个uniq进程分别处理,耗时仅为原方案的50%左右:

sort file.tsv | tee >(uniq -d > duplicateList.tsv) | uniq -u > clean_List.tsv

全兼容awk写法

如果当前shell不支持进程替换,可使用awk直接处理排序后的结果,兼容性更强:

sort file.tsv | awk '
$0 == prev {
    if (cnt == 1) print prev >> "duplicateList.tsv"
    print >> "duplicateList.tsv"
    cnt++
    next
}
{
    if (cnt == 1) print prev >> "clean_List.tsv"
    prev = $0
    cnt = 1
}
END {
    if (cnt == 1) print prev >> "clean_List.tsv"
}'

方案2:Python脚本实现(灵活可扩展)

如果需要后续自定义处理逻辑,可使用Python实现,内存占用极低,仅缓存两行内容,完全适配50G级大文件:

import subprocess

# 调用系统sort处理大文件,避免Python内存溢出
sort_process = subprocess.Popen(
    ["sort", "file.tsv"],
    stdout=subprocess.PIPE,
    text=True,
    bufsize=1
)

prev_line = None
line_count = 0

with open("duplicateList.tsv", "w", encoding="utf-8") as dup_file, \
     open("clean_List.tsv", "w", encoding="utf-8") as clean_file:
    for line in sort_process.stdout:
        if line == prev_line:
            line_count += 1
            # 第一次检测到重复时,先写入之前缓存的那行
            if line_count == 2:
                dup_file.write(prev_line)
            dup_file.write(line)
        else:
            if line_count == 1:
                clean_file.write(prev_line)
            prev_line = line
            line_count = 1
    # 处理最后一行数据
    if line_count == 1:
        clean_file.write(prev_line)

sort_process.wait()

内容的提问来源于stack exchange,提问作者Younes Zaidi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 09:54:07