50G大文件单次处理提取重复行与剩余内容方案求助
解决方案
方案1:Linux命令行实现(推荐,性能最优)
极简写法(支持进程替换的shell环境可用)
仅执行一次排序,通过tee和进程替换将排序结果同时分发给两个uniq进程分别处理,耗时仅为原方案的50%左右:
sort file.tsv | tee >(uniq -d > duplicateList.tsv) | uniq -u > clean_List.tsv
全兼容awk写法
如果当前shell不支持进程替换,可使用awk直接处理排序后的结果,兼容性更强:
sort file.tsv | awk ' $0 == prev { if (cnt == 1) print prev >> "duplicateList.tsv" print >> "duplicateList.tsv" cnt++ next } { if (cnt == 1) print prev >> "clean_List.tsv" prev = $0 cnt = 1 } END { if (cnt == 1) print prev >> "clean_List.tsv" }'
方案2:Python脚本实现(灵活可扩展)
如果需要后续自定义处理逻辑,可使用Python实现,内存占用极低,仅缓存两行内容,完全适配50G级大文件:
import subprocess # 调用系统sort处理大文件,避免Python内存溢出 sort_process = subprocess.Popen( ["sort", "file.tsv"], stdout=subprocess.PIPE, text=True, bufsize=1 ) prev_line = None line_count = 0 with open("duplicateList.tsv", "w", encoding="utf-8") as dup_file, \ open("clean_List.tsv", "w", encoding="utf-8") as clean_file: for line in sort_process.stdout: if line == prev_line: line_count += 1 # 第一次检测到重复时,先写入之前缓存的那行 if line_count == 2: dup_file.write(prev_line) dup_file.write(line) else: if line_count == 1: clean_file.write(prev_line) prev_line = line line_count = 1 # 处理最后一行数据 if line_count == 1: clean_file.write(prev_line) sort_process.wait()
内容的提问来源于stack exchange,提问作者Younes Zaidi
相关产品推荐
相关产品推荐

