运行Python程序后生成额外输出(数据处理)的方法咨询
耗时Python代码的后续数据加工方案
问题背景
在VSCode中运行耗时较长的Python代码后,需要对已处理完成的数据做后续加工(比如生成运行前未规划的图表),同时想确认pickle方案是否为最优选择。
pickle方案解析
你查到的pickle实现代码如下:
import pickle with open('training_history.pkl', 'wb') as f: pickle.dump(history, f) with open('training_history.pkl', 'rb') as f: loaded_history = pickle.load(f)
优点
- 操作简单,直接支持保存Python复杂对象(如模型训练的history字典、自定义类实例),无需额外格式转换。
- 保存和加载速度快,适合快速临时存储。
缺点
- 安全性不足:加载非可信来源的pickle文件可能触发恶意代码,仅适合自己生成的文件。
- 兼容性有限:不同Python版本、依赖库版本之间可能出现无法加载的情况(比如不同TensorFlow版本保存的history对象)。
- 可读性差:二进制格式无法直接查看内容,不方便调试和共享。
替代方案推荐
1. CSV/JSON格式(推荐用于结构化小数据)
适合保存history中的loss、准确率等结构化数值数据,可读性强,跨语言、跨环境兼容:
# 保存示例(用pandas) import pandas as pd # 假设history是模型训练返回的对象,history.history是字典格式的指标数据 pd.DataFrame(history.history).to_csv('training_history.csv', index=False) # 加载并生成图表 df = pd.read_csv('training_history.csv') df.plot(y=['loss', 'val_loss'], title='训练损失变化')
2. HDF5格式(推荐用于大型数值数据)
适合保存大张量、数据集等大型数值型数据,支持高效读写和压缩,适合机器学习场景:
# 保存示例 import pandas as pd with pd.HDFStore('training_history.h5') as store: store['history'] = pd.DataFrame(history.history) # 加载示例 with pd.HDFStore('training_history.h5') as store: df = store['history']
3. Feather/Parquet格式(高效的DataFrame存储)
专为Python pandas数据帧设计的格式,读写速度快、占用空间小,适合频繁读写数据的场景:
# Feather示例 import pandas as pd df = pd.DataFrame(history.history) df.to_feather('training_history.feather') # 加载 df = pd.read_feather('training_history.feather')
总结
- 如果只是临时自用、快速保存复杂Python对象,pickle完全够用;
- 如果需要长期存储、跨环境使用或与其他工具共享数据,优先选择CSV/JSON(小数据)、HDF5/Parquet(大数据)。
内容的提问来源于stack exchange,提问作者Si14
相关产品推荐
相关产品推荐

