为TFF创建SQLite格式自定义联邦图像数据集的最佳方法是什么
问题根因
体积膨胀的核心原因是存储的数据形态存在本质差异:
- 你当前是把JPG源文件解码为未压缩的原始RGB像素数组后序列化存储,单张224*224的3通道图片就占150KB左右,FairFace全量10万+样本的原始像素总大小超过14GB,压缩后得到6.4GB属于正常范围
- 源数据集的555MB是JPG压缩后的体积,CIFAR-100因为单张图片尺寸仅为32*32,全量原始像素总大小也不到200MB,就算直接存未压缩像素也不会有体积问题。它支持直接用FixedLenFeature指定shape读取的逻辑,本质还是读取固定长度的原始像素字节流,没有特殊存储逻辑。
推荐实现方案
按需求优先级分两种可选方案:
方案1:完全兼容CIFAR原生接口(无需修改上层业务代码)
如果你需要100%对齐CIFAR-100的存储解析逻辑,不修改任何上层调用代码,可做两处优化降低体积:
- 解析逻辑直接对齐CIFAR写法,省略冗余的解码、reshape步骤,示例如下:
def parse_proto(tensor_proto): parse_spec = { # 直接指定shape和dtype,不需要先读为string再解码 'image': tf.io.FixedLenFeature(shape=(224,224,3), dtype=tf.uint8), 'label': tf.io.FixedLenFeature(shape=(), dtype=tf.int64), } decoded_example = tf.io.parse_example(tensor_proto, parse_spec) return collections.OrderedDict( image=decoded_example['image'], label=decoded_example['label'])
- 调整lzma压缩参数,使用最高压缩等级(预设9级)替换默认压缩等级,可额外减少20%-30%的最终压缩包体积。
方案2:兼顾体积与兼容性(更推荐,上层业务代码零修改)
该方案可以让最终压缩包体积和源数据集基本一致,同时输出的数据集格式和原来完全对齐,上层业务不需要做任何调整:
- 存储环节不要将JPG解码为原始像素,直接读取JPG文件的二进制内容存入
bytes_feature - 解析环节将原来的
decode_raw+reshape替换为tf.io.decode_jpeg即可,示例如下:
# 存储环节修改为直接读JPG二进制 sample_img_binary = open(f"{img_path}/{sample_id}.jpg", 'rb').read() example = tf.train.Example(features=tf.train.Features(feature={ 'image' : bytes_feature(sample_img_binary), 'label' : int64_feature(sample_label) })) # 解析环节修改 def parse_proto(tensor_proto): parse_spec = { 'image': tf.io.FixedLenFeature(shape=(), dtype=tf.string), 'label': tf.io.FixedLenFeature(shape=(), dtype=tf.int64), } decoded_example = tf.io.parse_example(tensor_proto, parse_spec) return collections.OrderedDict( image=tf.io.decode_jpeg(decoded_example['image'], channels=3), label=decoded_example['label'])
内容的提问来源于stack exchange,提问作者Gaurav Goswami
相关产品推荐
相关产品推荐

