You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多进程处理120GB大文件字符替换求助

多进程高效处理大文本文件字符替换

核心优化思路

  • 放弃单进程逐行处理,利用多核CPU并行处理文件块
  • 用str.replace替代re.sub做简单字符替换,性能提升明显
  • 按行边界拆分文件块,避免截断行内容
  • 每个进程独立处理块并写入单独输出文件(不关心行顺序的情况下,无需额外合并步骤;若需单个文件,用系统命令合并更高效)

具体实现代码

文件分块工具函数

用于定位每个进程要处理的起始/结束位置,确保块从行首开始,避免拆分到行中间:

import os
from multiprocessing import Pool

def get_chunk_boundaries(file_path, num_chunks):
    file_size = os.path.getsize(file_path)
    chunk_size = file_size // num_chunks
    boundaries = []
    with open(file_path, 'r', encoding='utf-8') as f:
        start = 0
        for i in range(num_chunks - 1):
            f.seek(start + chunk_size)
            # 跳转到下一行开头
            f.readline()
            end = f.tell()
            boundaries.append((start, end))
            start = end
        boundaries.append((start, file_size))
    return boundaries

单进程处理函数

负责读取指定块、执行字符替换、写入输出文件:

def process_chunk(args):
    input_path, output_path, start, end, replacements = args
    with open(input_path, 'r', encoding='utf-8') as f_in, open(output_path, 'w', encoding='utf-8') as f_out:
        f_in.seek(start)
        while f_in.tell() < end:
            line = f_in.readline()
            # 批量应用替换规则
            for old, new in replacements.items():
                line = line.replace(old, new)
            f_out.write(line)

主函数(启动多进程)

def main():
    input_file = r'C:\Projects\orders.txt'
    output_dir = r'C:\Projects\processed_chunks'
    num_processes = 6  # 匹配你的6核CPU
    # 替换规则,未来扩展只需添加新键值对
    replacements = {
        'ß': 's',
        # 示例扩展:'ä': 'ae', 'ö': 'oe', 'ü': 'ue'
    }

    # 创建输出目录
    os.makedirs(output_dir, exist_ok=True)
    # 获取分块边界
    boundaries = get_chunk_boundaries(input_file, num_processes)
    # 准备进程参数
    process_args = []
    for i, (start, end) in enumerate(boundaries):
        output_file = os.path.join(output_dir, f'order_chunk_{i}.txt')
        process_args.append((input_file, output_file, start, end, replacements))
    
    # 启动多进程池
    with Pool(num_processes) as pool:
        pool.map(process_chunk, process_args)
    
    print("所有块处理完成!")
    # 若需合并为单个文件,可执行系统命令(Windows示例,Linux用cat)
    # import subprocess
    # final_output = r'C:\Projects\orders_new.txt'
    # subprocess.run(f'copy /b {output_dir}\*.txt {final_output}', shell=True)

if __name__ == '__main__':
    main()

关键说明

  1. 分块逻辑:通过seek定位后读取换行符,确保每个块的起始位置都是行首,不会破坏行内容完整性。
  2. 替换效率:str.replace比正则表达式快数倍;如果未来有大量替换规则,可改用str.translate(需提前构建字符映射表)进一步提升性能。
  3. IO优化:每个进程写入独立文件,避免多进程写同一文件的锁竞争;系统级合并命令比Python逐行写入快得多。
  4. 编码注意:必须指定正确的文件编码(如utf-8),避免德语字符乱码。

扩展建议

如果未来替换规则复杂,可将规则存入JSON配置文件动态加载;若文件远超内存,可调整分块大小,避免单个块占用过多内存。

内容的提问来源于stack exchange,提问作者Juki777

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 21:00:54