如何比较行序无关的超大型CSV文件?仅需返回布尔值
超大型CSV文件行序无关对比方案求助
核心需求
- 行的顺序不影响对比结果
- 仅需返回对比结果(True/False),无需输出具体差异内容
示例说明
文件1内容:
a,b,c,d e,f,g,h i,j,k,l
文件2内容:
a,b,c,d i,j,k,l e,f,g,h
上述两个文件行序不同但内容完全一致,应判定为对比通过;若存在行内容差异、列值不匹配或某行仅存在于单个文件中,则对比失败。
当前困境
处理的CSV文件规模极大:包含1400万至3000万行、10至15列,原始未排序文件大小约1GB,且无可用作排序依据的键列。
已尝试方案
通过自定义归并排序函数对文件排序后,再执行diff file1 file2完成对比。排序函数代码如下:
def batch_sort(self, input, output, key=None, buffer_size=32000, tempdirs=None): if isinstance(tempdirs, str): tempdirs = tempdirs.split(",") if tempdirs is None: tempdirs = [] if not tempdirs: tempdirs.append(gettempdir()) chunks = [] try: with open(input,'rb',64*1024) as input_file: input_iterator = iter(input_file) for tempdir in cycle(tempdirs): current_chunk = list(islice(input_iterator,buffer_size)) if not current_chunk: break current_chunk.sort(key=key) output_chunk = open(os.path.join(tempdir,'%06i'%len(chunks)),'w+b',64*1024) chunks.append(output_chunk) output_chunk.writelines(current_chunk) output_chunk.flush() output_chunk.seek(0) with open(output,'wb',64*1024) as output_file: output_file.writelines(self.merge(key, *chunks)) finally: for chunk in chunks: try: chunk.close() os.remove(chunk.name) except Exception: pass
但该方法在处理1400万行以上的文件时会失效,且排序操作耗时极高。此前尝试过filecmm、difflib等工具,均要求文件预先排序,希望找到无需排序即可实现行序无关对比的方案。
内容的提问来源于stack exchange,提问作者luci5r
相关产品推荐
相关产品推荐

