如何高效将数据同时保存至本地与S3对象存储?
优化工件本地与S3存储的pickle复用方案
需求:将一个工件(artifact)同时保存到本地磁盘和S3对象存储,要求:1. 同步写入本地磁盘;2. 异步将文件副本发送至S3。当前代码对同一个
res对象分别调用了pickle.dump()和pickle.dumps(),想知道是否可以复用第一次pickle转换的结果来优化流程?当前代码:
# res can be any python type with open(cachepath, 'wb') as f: logger.info("Writing results to cache (pkl).") pickle.dump(res, f) post_to_object_store(cachepath, pickle.dumps(res), async=True, app=current_app)
当然可以复用第一次pickle序列化的结果,这样能避免对同一个对象做两次序列化操作,节省CPU和内存资源(尤其是当res对象体积较大时,效果更明显)。
推荐两种优化方案:
方案一:先序列化一次,再复用二进制数据
先调用pickle.dumps()把res序列化为二进制数据,再将这份数据同时写入本地文件和传给S3的异步接口:
import pickle import logging logger = logging.getLogger(__name__) # 仅做一次pickle序列化 pickle_data = pickle.dumps(res) # 同步写入本地磁盘 with open(cachepath, 'wb') as f: logger.info("Writing results to cache (pkl).") f.write(pickle_data) # 异步将序列化后的数据发送至S3 post_to_object_store(cachepath, pickle_data, async=True, app=current_app)
方案二:写入本地时同时留存序列化数据
如果需要保留pickle.dump()的写法(比如某些场景下需要流写入),可以在写入文件的同时把数据存下来:
pickle_data = bytearray() with open(cachepath, 'wb') as f: logger.info("Writing results to cache (pkl).") # 创建一个可以同时写入文件和内存的IO对象 class DualWriter: def __init__(self, file_obj, buffer): self.file_obj = file_obj self.buffer = buffer def write(self, data): self.file_obj.write(data) self.buffer.extend(data) writer = DualWriter(f, pickle_data) pickle.dump(res, writer) # 异步发送留存的序列化数据到S3 post_to_object_store(cachepath, bytes(pickle_data), async=True, app=current_app)
两种方案中,方案一更简洁高效,因为它直接一次序列化得到数据,同时满足本地写入和S3传输的需求,完全避免了重复序列化的开销。
内容的提问来源于stack exchange,提问作者Edmund's Echo
相关产品推荐
相关产品推荐

