Python实现超大文本文件去重:低内存磁盘式方案问询
超大文本文件去重(低内存Python实现)
核心思路
由于无法将超大文件完整加载到内存,核心方案是用磁盘存储已出现行的唯一标识(哈希值),避免内存占用爆炸。这里推荐使用SQLite作为磁盘哈希存储(轻量、无需额外服务、Python内置支持),或者分块+归并排序的纯文件操作方案,两种方式都能将内存控制在极低水平。
方案一:SQLite磁盘哈希表(简单高效)
每次仅读取一行到内存,用哈希值记录已出现的行,内存占用几乎恒定。
import sqlite3 import hashlib import os def remove_duplicates_large_file(input_path, output_path): # 创建磁盘SQLite数据库存储哈希值,避免内存溢出 conn = sqlite3.connect('temp_line_hashes.db') cursor = conn.cursor() # 创建唯一约束的哈希表,确保不会存储重复哈希 cursor.execute('''CREATE TABLE IF NOT EXISTS hashes (hash TEXT PRIMARY KEY)''') conn.commit() try: with open(input_path, 'r', encoding='utf-8') as infile, \ open(output_path, 'w', encoding='utf-8') as outfile: line_count = 0 for line in infile: # 计算行的SHA-256哈希,冲突概率可忽略 line_hash = hashlib.sha256(line.encode('utf-8')).hexdigest() # 检查哈希是否已存在 cursor.execute('SELECT 1 FROM hashes WHERE hash = ?', (line_hash,)) if not cursor.fetchone(): outfile.write(line) cursor.execute('INSERT INTO hashes (hash) VALUES (?)', (line_hash,)) # 每1000行提交一次事务,减少磁盘IO频率 line_count += 1 if line_count % 1000 == 0: conn.commit() # 最后提交剩余事务 conn.commit() finally: conn.close() # 清理临时数据库文件 os.remove('temp_line_hashes.db') # 使用示例 remove_duplicates_large_file('large_input.txt', 'deduped_output.txt')
优化细节
- 哈希选择:用SHA-256而非MD5,虽然计算稍慢,但冲突概率极低,避免误判重复行。
- 批量提交:每处理1000行提交一次事务,减少磁盘IO次数,提升整体效率。
- 内存控制:每次仅加载一行数据,数据库仅存储64字符的哈希值,内存占用几乎无增长。
方案二:分块+归并排序(无依赖极端场景)
如果文件大到连哈希表的磁盘存储都有压力,可以采用分块处理+归并排序的纯文件操作方案,完全不依赖数据库:
import os from tempfile import TemporaryDirectory def split_and_dedup(input_path, chunk_size=100*1024*1024): # 按100MB分块 temp_dir = TemporaryDirectory() chunk_paths = [] with open(input_path, 'r', encoding='utf-8') as infile: chunk_num = 0 while True: # 读取一块数据,处理行截断问题 chunk_data = infile.read(chunk_size) if not chunk_data: break lines = chunk_data.splitlines(keepends=True) # 补全可能被截断的最后一行 if not chunk_data.endswith('\n'): next_line = infile.readline() if next_line: lines[-1] += next_line # 块内内存去重并排序 unique_lines = sorted(set(lines)) # 写入临时块文件 chunk_path = os.path.join(temp_dir.name, f'chunk_{chunk_num}.txt') with open(chunk_path, 'w', encoding='utf-8') as chunk_file: chunk_file.writelines(unique_lines) chunk_paths.append(chunk_path) chunk_num += 1 return temp_dir, chunk_paths def merge_dedup(chunk_paths, output_path): # 打开所有临时块文件,创建迭代器 file_handles = [open(path, 'r', encoding='utf-8') for path in chunk_paths] line_iterators = [iter(fh) for fh in file_handles] current_lines = [] # 初始化每个迭代器的第一行 for idx, it in enumerate(line_iterators): try: current_lines.append((next(it), idx)) except StopIteration: pass # 归并排序去重 with open(output_path, 'w', encoding='utf-8') as outfile: last_written_line = None while current_lines: # 按行内容排序,取最小的行 current_lines.sort(key=lambda x: x[0]) current_line, idx = current_lines.pop(0) if current_line != last_written_line: outfile.write(current_line) last_written_line = current_line # 读取当前块的下一行 try: next_line = next(line_iterators[idx]) current_lines.append((next_line, idx)) except StopIteration: file_handles[idx].close() # 关闭剩余文件句柄 for fh in file_handles: if not fh.closed: fh.close() def remove_duplicates_extreme_large(input_path, output_path): temp_dir, chunk_paths = split_and_dedup(input_path) try: merge_dedup(chunk_paths, output_path) finally: # 自动清理临时目录 temp_dir.cleanup() # 使用示例 remove_duplicates_extreme_large('extreme_large_input.txt', 'final_deduped.txt')
适用场景
- 当文件体积远超磁盘剩余空间(哈希表方案需要额外存储哈希)时,分块方案更合适。
- 无SQLite依赖的环境下,纯文件操作更稳妥。
内容的提问来源于stack exchange,提问作者Lexoner
相关产品推荐
相关产品推荐

