Python多进程处理120GB大文件字符替换求助
多进程高效处理大文本文件字符替换
核心优化思路
- 放弃单进程逐行处理,利用多核CPU并行处理文件块
- 用
str.replace替代re.sub做简单字符替换,性能提升明显 - 按行边界拆分文件块,避免截断行内容
- 每个进程独立处理块并写入单独输出文件(不关心行顺序的情况下,无需额外合并步骤;若需单个文件,用系统命令合并更高效)
具体实现代码
文件分块工具函数
用于定位每个进程要处理的起始/结束位置,确保块从行首开始,避免拆分到行中间:
import os from multiprocessing import Pool def get_chunk_boundaries(file_path, num_chunks): file_size = os.path.getsize(file_path) chunk_size = file_size // num_chunks boundaries = [] with open(file_path, 'r', encoding='utf-8') as f: start = 0 for i in range(num_chunks - 1): f.seek(start + chunk_size) # 跳转到下一行开头 f.readline() end = f.tell() boundaries.append((start, end)) start = end boundaries.append((start, file_size)) return boundaries
单进程处理函数
负责读取指定块、执行字符替换、写入输出文件:
def process_chunk(args): input_path, output_path, start, end, replacements = args with open(input_path, 'r', encoding='utf-8') as f_in, open(output_path, 'w', encoding='utf-8') as f_out: f_in.seek(start) while f_in.tell() < end: line = f_in.readline() # 批量应用替换规则 for old, new in replacements.items(): line = line.replace(old, new) f_out.write(line)
主函数(启动多进程)
def main(): input_file = r'C:\Projects\orders.txt' output_dir = r'C:\Projects\processed_chunks' num_processes = 6 # 匹配你的6核CPU # 替换规则,未来扩展只需添加新键值对 replacements = { 'ß': 's', # 示例扩展:'ä': 'ae', 'ö': 'oe', 'ü': 'ue' } # 创建输出目录 os.makedirs(output_dir, exist_ok=True) # 获取分块边界 boundaries = get_chunk_boundaries(input_file, num_processes) # 准备进程参数 process_args = [] for i, (start, end) in enumerate(boundaries): output_file = os.path.join(output_dir, f'order_chunk_{i}.txt') process_args.append((input_file, output_file, start, end, replacements)) # 启动多进程池 with Pool(num_processes) as pool: pool.map(process_chunk, process_args) print("所有块处理完成!") # 若需合并为单个文件,可执行系统命令(Windows示例,Linux用cat) # import subprocess # final_output = r'C:\Projects\orders_new.txt' # subprocess.run(f'copy /b {output_dir}\*.txt {final_output}', shell=True) if __name__ == '__main__': main()
关键说明
- 分块逻辑:通过
seek定位后读取换行符,确保每个块的起始位置都是行首,不会破坏行内容完整性。 - 替换效率:
str.replace比正则表达式快数倍;如果未来有大量替换规则,可改用str.translate(需提前构建字符映射表)进一步提升性能。 - IO优化:每个进程写入独立文件,避免多进程写同一文件的锁竞争;系统级合并命令比Python逐行写入快得多。
- 编码注意:必须指定正确的文件编码(如
utf-8),避免德语字符乱码。
扩展建议
如果未来替换规则复杂,可将规则存入JSON配置文件动态加载;若文件远超内存,可调整分块大小,避免单个块占用过多内存。
内容的提问来源于stack exchange,提问作者Juki777
相关产品推荐
相关产品推荐

