不同图像读取方式致字节表示差异——TensorFlow目标检测兼容问询
关于TensorFlow目标检测中图像读取与编码兼容性的修正方案
嘿,我看你在做TensorFlow目标检测的数据增强工作,还在验证不同图像读取方式的兼容性——这里有个关键的小错误需要修正,我给你拆解清楚:
首先,你当前的代码里有个致命问题:把解码后的numpy像素数组直接塞进io.BytesIO是行不通的,因为BytesIO需要的是编码后的图像字节流(比如原始JPG文件的二进制数据),而不是解码后的像素矩阵。
先看你原来的代码片段:
full_path = 'path/to/my/image.jpg' image = PIL.Image.open(full_path) image_np = np.array(image) encoded_jpg_io1 = io.BytesIO(image_np) # 这里错了!image_np是像素数组,不是JPG编码字节
而你用TensorFlow读取的方式是对的,它直接读取了文件的原始JPG编码字节:
with tf.gfile.GFile(full_path, 'rb') as fid: encoded_jpg = fid.read() # 这是原始的JPG二进制数据,可直接用于TFRecord
正确的兼容做法
如果你想用PIL读取图像后,得到和TensorFlow读取一致的编码字节(用于TFRecord),需要把PIL图像重新编码回JPG格式,再存入BytesIO:
import io import PIL.Image import numpy as np import tensorflow as tf full_path = 'path/to/my/image.jpg' # 用PIL读取并重新编码为JPG字节流 image = PIL.Image.open(full_path) buffer = io.BytesIO() image.save(buffer, format='JPEG') # 核心步骤:把PIL图像编码回JPG格式 encoded_jpg_from_pil = buffer.getvalue() # TensorFlow直接读取原始文件字节 with tf.gfile.GFile(full_path, 'rb') as fid: encoded_jpg_from_tf = fid.read() # 验证兼容性:两者的字节流基本一致(除非原图像编码参数和PIL保存时不同) print(np.array_equal(encoded_jpg_from_pil, encoded_jpg_from_tf))
数据增强后的正确流程
如果你的目标是对解码后的像素数组做增强,再编码回JPG存入TFRecord,流程应该是这样:
# 1. 读取并解码为像素数组 image = PIL.Image.open(full_path) image_np = np.array(image) # 2. 执行数据增强操作(比如翻转、裁剪、亮度调整等) augmented_image_np = your_augmentation_function(image_np) # 3. 将增强后的像素数组转回PIL图像,再编码为JPG字节 augmented_image = PIL.Image.fromarray(augmented_image_np) buffer = io.BytesIO() augmented_image.save(buffer, format='JPEG') encoded_augmented_jpg = buffer.getvalue() # 4. 现在这个encoded_augmented_jpg就可以写入TFRecord了
小提醒
如果你的图像是PNG格式,记得把format='JPEG'改成format='PNG',避免编码错误;另外,如果需要精确匹配原文件的编码字节,可以查看原图像的编码参数(比如JPG质量),在image.save()时指定相同参数。
内容的提问来源于stack exchange,提问作者Eypros
相关产品推荐
相关产品推荐

