如何在Python中批量捕获stdout输出并按批次保存?
没问题!针对你要处理2300万行输出并分批保存的需求,我给你准备了两种场景的实现方案,你可以根据实际情况选择:
场景1:通过子进程运行外部Python命令(比如调用另一个脚本)
如果你的function_call_here()是需要通过命令行启动的外部脚本,咱们可以用subprocess模块实时捕获输出,逐行处理并分批保存:
import subprocess import pickle def process_large_output(command): # 定义每批次的行数 BATCH_SIZE = 100000 batch_num = 1 current_batch = [] # 启动子进程,实时捕获stdout with subprocess.Popen( command, stdout=subprocess.PIPE, stderr=subprocess.STDOUT, # 可选:如果需要同时捕获stderr,取消注释 text=True, # 以文本模式读取输出,自动解码为字符串 bufsize=1 # 行缓冲模式,确保实时获取输出 ) as proc: # 逐行读取输出 for line in proc.stdout: # 去掉末尾的换行符(可根据需求保留) cleaned_line = line.rstrip('\n') current_batch.append(cleaned_line) # 当当前批次达到指定行数时,保存并重置 if len(current_batch) == BATCH_SIZE: filename = f"output_{batch_num:04d}.pkl" with open(filename, 'wb') as f: pickle.dump(current_batch, f) print(f"read batch {batch_num}") current_batch = [] batch_num += 1 # 处理最后一批不足10万行的剩余内容 if current_batch: filename = f"output_{batch_num:04d}.pkl" with open(filename, 'wb') as f: pickle.dump(current_batch, f) print(f"read batch {batch_num} (final batch)") # 调用示例:替换成你实际的命令,比如运行另一个Python脚本 process_large_output(["python", "your_external_script.py"])
关键细节说明:
- 使用
subprocess.Popen而非subprocess.run,因为前者可以实时逐行读取输出,不会等到命令完全结束才处理,避免内存过载 bufsize=1开启行缓冲,确保输出一行就读取一行text=True让输出以字符串形式返回,省去手动解码的步骤- 最后一定要处理剩余的不足批次大小的内容,避免数据丢失
场景2:直接调用另一个模块的函数(函数内部直接输出到stdout)
如果你的function_call_here()是可以直接导入调用的模块函数,咱们可以通过重定向stdout来捕获输出,再分批保存:
import sys import pickle from io import StringIO def capture_and_batch(func, batch_size=100000): batch_num = 1 # 保存原本的stdout,后面要恢复 old_stdout = sys.stdout # 创建一个StringIO对象来捕获输出 sys.stdout = captured_output = StringIO() try: # 调用目标函数,它的所有print输出都会被捕获 func() # 将捕获的输出按行拆分 output_lines = captured_output.getvalue().splitlines() # 分批处理输出行 for i in range(0, len(output_lines), batch_size): current_batch = output_lines[i:i+batch_size] filename = f"output_{batch_num:04d}.pkl" with open(filename, 'wb') as f: pickle.dump(current_batch, f) print(f"read batch {batch_num}") batch_num += 1 finally: # 恢复原本的stdout,避免影响后续输出 sys.stdout = old_stdout # 调用示例:从你的模块导入目标函数 from your_module import function_call_here capture_and_batch(function_call_here)
额外提示:
- 读取保存的pickle文件很简单:
with open("output_0015.pkl", 'rb') as f: batch_data = pickle.load(f) - 如果输出包含二进制内容,需要调整读取模式(去掉
text=True,改用rb模式处理字节) - 可以根据实际内存情况调整
BATCH_SIZE的值,比如内存紧张可以适当调小
内容的提问来源于stack exchange,提问作者Drew Brady
相关产品推荐
相关产品推荐

