You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现超大文本文件去重:低内存磁盘式方案问询

超大文本文件去重(低内存Python实现)

核心思路

由于无法将超大文件完整加载到内存,核心方案是用磁盘存储已出现行的唯一标识(哈希值),避免内存占用爆炸。这里推荐使用SQLite作为磁盘哈希存储(轻量、无需额外服务、Python内置支持),或者分块+归并排序的纯文件操作方案,两种方式都能将内存控制在极低水平。

方案一:SQLite磁盘哈希表(简单高效)

每次仅读取一行到内存,用哈希值记录已出现的行,内存占用几乎恒定。

import sqlite3
import hashlib
import os

def remove_duplicates_large_file(input_path, output_path):
    # 创建磁盘SQLite数据库存储哈希值,避免内存溢出
    conn = sqlite3.connect('temp_line_hashes.db')
    cursor = conn.cursor()
    # 创建唯一约束的哈希表,确保不会存储重复哈希
    cursor.execute('''CREATE TABLE IF NOT EXISTS hashes
                      (hash TEXT PRIMARY KEY)''')
    conn.commit()

    try:
        with open(input_path, 'r', encoding='utf-8') as infile, \
             open(output_path, 'w', encoding='utf-8') as outfile:
            line_count = 0
            for line in infile:
                # 计算行的SHA-256哈希,冲突概率可忽略
                line_hash = hashlib.sha256(line.encode('utf-8')).hexdigest()
                # 检查哈希是否已存在
                cursor.execute('SELECT 1 FROM hashes WHERE hash = ?', (line_hash,))
                if not cursor.fetchone():
                    outfile.write(line)
                    cursor.execute('INSERT INTO hashes (hash) VALUES (?)', (line_hash,))
                # 每1000行提交一次事务,减少磁盘IO频率
                line_count += 1
                if line_count % 1000 == 0:
                    conn.commit()
        # 最后提交剩余事务
        conn.commit()
    finally:
        conn.close()
        # 清理临时数据库文件
        os.remove('temp_line_hashes.db')

# 使用示例
remove_duplicates_large_file('large_input.txt', 'deduped_output.txt')

优化细节

  • 哈希选择:用SHA-256而非MD5,虽然计算稍慢,但冲突概率极低,避免误判重复行。
  • 批量提交:每处理1000行提交一次事务,减少磁盘IO次数,提升整体效率。
  • 内存控制:每次仅加载一行数据,数据库仅存储64字符的哈希值,内存占用几乎无增长。

方案二:分块+归并排序(无依赖极端场景)

如果文件大到连哈希表的磁盘存储都有压力,可以采用分块处理+归并排序的纯文件操作方案,完全不依赖数据库:

import os
from tempfile import TemporaryDirectory

def split_and_dedup(input_path, chunk_size=100*1024*1024):  # 按100MB分块
    temp_dir = TemporaryDirectory()
    chunk_paths = []
    with open(input_path, 'r', encoding='utf-8') as infile:
        chunk_num = 0
        while True:
            # 读取一块数据,处理行截断问题
            chunk_data = infile.read(chunk_size)
            if not chunk_data:
                break
            lines = chunk_data.splitlines(keepends=True)
            # 补全可能被截断的最后一行
            if not chunk_data.endswith('\n'):
                next_line = infile.readline()
                if next_line:
                    lines[-1] += next_line
            # 块内内存去重并排序
            unique_lines = sorted(set(lines))
            # 写入临时块文件
            chunk_path = os.path.join(temp_dir.name, f'chunk_{chunk_num}.txt')
            with open(chunk_path, 'w', encoding='utf-8') as chunk_file:
                chunk_file.writelines(unique_lines)
            chunk_paths.append(chunk_path)
            chunk_num += 1
    return temp_dir, chunk_paths

def merge_dedup(chunk_paths, output_path):
    # 打开所有临时块文件,创建迭代器
    file_handles = [open(path, 'r', encoding='utf-8') for path in chunk_paths]
    line_iterators = [iter(fh) for fh in file_handles]
    current_lines = []

    # 初始化每个迭代器的第一行
    for idx, it in enumerate(line_iterators):
        try:
            current_lines.append((next(it), idx))
        except StopIteration:
            pass

    # 归并排序去重
    with open(output_path, 'w', encoding='utf-8') as outfile:
        last_written_line = None
        while current_lines:
            # 按行内容排序,取最小的行
            current_lines.sort(key=lambda x: x[0])
            current_line, idx = current_lines.pop(0)
            if current_line != last_written_line:
                outfile.write(current_line)
                last_written_line = current_line
            # 读取当前块的下一行
            try:
                next_line = next(line_iterators[idx])
                current_lines.append((next_line, idx))
            except StopIteration:
                file_handles[idx].close()
    # 关闭剩余文件句柄
    for fh in file_handles:
        if not fh.closed:
            fh.close()

def remove_duplicates_extreme_large(input_path, output_path):
    temp_dir, chunk_paths = split_and_dedup(input_path)
    try:
        merge_dedup(chunk_paths, output_path)
    finally:
        # 自动清理临时目录
        temp_dir.cleanup()

# 使用示例
remove_duplicates_extreme_large('extreme_large_input.txt', 'final_deduped.txt')

适用场景

  • 当文件体积远超磁盘剩余空间(哈希表方案需要额外存储哈希)时,分块方案更合适。
  • 无SQLite依赖的环境下,纯文件操作更稳妥。

内容的提问来源于stack exchange,提问作者Lexoner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 02:27:25