如何用Pythonic方式实现该AWK拆分Mbox脚本并支持断点续传?
Pythonic Mbox拆分工具(支持中断恢复)
你的递归思路其实走偏了——递归会反复重新读取整个超大Mbox文件,完全没必要,而且逻辑上很容易陷入无限循环。下面是更符合Python风格的实现,完美复刻AWK命令的拆分逻辑,还额外支持中断后无缝恢复进度:
完整实现代码
import os from pathlib import Path def split_mbox(mbox_path, chunk_size=500, output_dir="./TestChunks/chunks"): # 确保输出目录存在 output_dir = Path(output_dir) output_dir.mkdir(parents=True, exist_ok=True) # 用隐藏文件记录进度,避免中断后从头再来 progress_file = output_dir / ".split_progress" # 初始化进度参数 processed_msgs = 0 file_offset = 0 current_chunk = 0 # 读取已有进度(如果存在) if progress_file.exists(): with open(progress_file, "r") as f: lines = f.read().strip().split("\n") if len(lines) >= 3: processed_msgs = int(lines[0]) file_offset = int(lines[1]) current_chunk = int(lines[2]) print(f"📌 恢复上次进度:已处理 {processed_msgs} 条消息,当前分块 {current_chunk},文件偏移量 {file_offset}") # 打开原始Mbox文件,直接跳转到中断位置 with open(mbox_path, "r") as mbox_file: mbox_file.seek(file_offset) chunk_file = None try: # 根据进度初始化当前分块文件 if processed_msgs > 0: current_chunk_msgs = processed_msgs % chunk_size if current_chunk_msgs != 0: # 当前分块还没填满,继续追加 chunk_file = open(output_dir / f"chunk_{current_chunk}.txt", "a") else: # 分块已满,新建下一个分块 current_chunk += 1 chunk_file = open(output_dir / f"chunk_{current_chunk}.txt", "w") else: # 首次运行,创建第一个分块 chunk_file = open(output_dir / f"chunk_{current_chunk}.txt", "w") # 逐行处理Mbox文件(内存友好,不会加载整个大文件) for line in mbox_file: # 识别Mbox消息的起始标记(标准格式是"From "开头) if line.startswith("From "): processed_msgs += 1 # 检查是否需要切换分块 if processed_msgs % chunk_size == 0: chunk_file.close() current_chunk += 1 chunk_file = open(output_dir / f"chunk_{current_chunk}.txt", "w") print(f"✅ 完成分块 {current_chunk-1},开始处理分块 {current_chunk}") # 写入当前行到分块文件 chunk_file.write(line) # 每处理100条消息更新一次进度(平衡性能和可靠性) if processed_msgs % 100 == 0: with open(progress_file, "w") as f: f.write(f"{processed_msgs}\n{mbox_file.tell()}\n{current_chunk}") finally: # 确保所有文件都被正确关闭,即使中途中断 if chunk_file: chunk_file.close() # 最后一次更新进度,确保数据不丢失 with open(progress_file, "w") as f: f.write(f"{processed_msgs}\n{mbox_file.tell()}\n{current_chunk}") print(f"🎉 处理结束!共处理 {processed_msgs} 条消息,最后一个分块是 {current_chunk}") # 调用示例 if __name__ == "__main__": # 这里可以调整分块大小,比如改成20测试 split_mbox("mbox", chunk_size=500)
关键特性说明
- 中断恢复:通过
.split_progress隐藏文件记录已处理消息数、文件读取位置、当前分块号,下次启动直接从断点继续,不用重新读取整个大文件。 - 内存友好:逐行读取Mbox文件,不会把几GB的超大文件加载到内存,适合处理大型邮箱文件。
- 标准Mbox兼容:严格遵循Mbox格式规范,以
From开头的行作为消息分隔符,拆分结果和你的AWK命令完全一致。 - Pythonic设计:用
pathlib管理路径(比os模块更直观),用上下文管理器和try...finally保证文件资源安全,代码结构清晰易维护。
和你原递归代码的对比
你的递归思路最大问题是每次递归都会重新打开Mbox文件从头读取,不仅效率极低,还会导致无限递归(msg>20就调用自己,新调用又从头开始计数)。上面的迭代实现完全避免了这个问题,而且逻辑更清晰。
内容的提问来源于stack exchange,提问作者ganicus
相关产品推荐
相关产品推荐

