You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow加载图像报未知图像格式DecodeJpeg错误求助

问题描述

使用TensorFlow批量提取图像特征时触发如下运行报错:

Unknown image file format. One of JPEG, PNG, GIF, BMP required. [[{{node DecodeJpeg}}]] [Op:IteratorGetNext]

特征提取核心实现代码如下:

encode_train = sorted(set(img_name_vector))
image_dataset = tf.data.Dataset.from_tensor_slices(encode_train)
image_dataset = image_dataset.map(load_image, num_parallel_calls=tf.data.experimental.AUTOTUNE).batch(64)

%%time
for img, path in tqdm(image_dataset):
  batch_features = image_features_extract_model(img)
  batch_features = tf.reshape(batch_features,(batch_features.shape[0], -1, batch_features.shape[3]))

  for bf, p in zip(batch_features, path):
    path_of_feature = p.numpy().decode("utf-8")
    np.save(path_of_feature, bf.numpy())

已提前执行格式校验:遍历./new_google_images/目录下所有.png、.jpg后缀文件,基于imghdr校验文件格式是否属于TensorFlow支持的bmp/gif/jpeg/png类型,校验结果显示无异常文件,校验代码如下:

from pathlib import Path
import imghdr

data_dir = "./new_google_images/"
image_extensions = [".png", ".jpg"]  # 待校验的图像后缀列表

img_type_accepted_by_tf = ["bmp", "gif", "jpeg", "png"]

for filepath in Path(data_dir).rglob("*"):
    if filepath.suffix.lower() in image_extensions:
        img_type = imghdr.what(filepath)
        if img_type is None:
            print(f"{filepath} is not an image")
        elif img_type not in img_type_accepted_by_tf:
            print(f"{filepath} is a {img_type}, not accepted by TensorFlow")
报错核心原因
  • 预校验逻辑存在缺陷:imghdr仅读取文件头部少量特征字节判断格式,无法识别文件截断、损坏的情况——例如下载中途中断的jpg文件,头部符合jpeg特征但内容缺失,imghdr会判定为合法文件,但TensorFlow读取全量内容解码时就会抛出格式错误;旧版本imghdr还可能把篡改后缀的webp、heic等TensorFlow不支持的格式误判为jpeg,漏过异常文件。
  • 扫描范围和数据集范围不匹配:构建数据集用的img_name_vector路径列表,和预校验脚本扫描的路径范围不一致,比如大写后缀(.JPG、.Png)的文件、其他子目录下的异常文件被加入数据集,但没有被预校验逻辑覆盖。
  • 解码逻辑硬编码:如果load_image函数内固定使用tf.io.decode_jpeg处理所有读入的文件,哪怕路径里混入png、bmp这类TensorFlow支持的格式,也会因为强制走JPEG解码流程触发报错,和文件本身合法性无关。
  • 异常文件漏判:零字节空文件、软链接指向失效的文件,可能因为扫描时权限问题、路径匹配问题被imghdr校验漏过,进入数据集后解码直接失败。
排查与解决方案
  • 优先定位具体异常文件:给load_image加一层异常捕获,打印触发解码错误的文件路径,直接定位问题根源,参考实现:
def load_image_with_catch(path):
    try:
        return load_image(path)
    except tf.errors.InvalidArgumentError:
        err_path = path.numpy().decode("utf-8")
        print(f"解码失败文件: {err_path}")
        # 返回统一占位张量,避免单个坏文件中断整个批量处理流程
        return tf.zeros([299, 299, 3]), path

将数据集map调用中的load_image替换为上述函数,运行后即可输出所有触发报错的文件路径,针对性删除或修复即可。

  • 补全预校验逻辑:替换原有仅靠imghdr的校验逻辑,增加文件大小校验、TensorFlow原生解码校验,从源头剔除坏文件,参考实现:
valid_img_paths = []
# 覆盖所有常见大小写图像后缀
check_extensions = {".png", ".jpg", ".jpeg", ".bmp", ".gif"}
for filepath in Path(data_dir).rglob("*"):
    if filepath.suffix.lower() not in check_extensions:
        continue
    # 过滤小于1KB的异常小文件/零字节文件
    if filepath.stat().st_size < 1024:
        print(f"剔除异常小文件: {filepath}")
        continue
    # 用TensorFlow自带解码能力做最终校验
    try:
        raw_bytes = tf.io.read_file(str(filepath))
        # 自动识别格式解码,关闭GIF动图扩展避免维度异常
        tf.io.decode_image(raw_bytes, expand_animations=False)
        valid_img_paths.append(str(filepath))
    except Exception as e:
        print(f"剔除解码失败文件 {filepath}: {str(e)}")
# 后续直接用valid_img_paths构建数据集即可
  • 修正解码逻辑:不要在load_image中硬编码使用tf.io.decode_jpeg,替换为tf.io.decode_image,该接口会自动识别文件格式匹配对应解码器,避免合法格式文件被错误走JPEG解码流程触发报错。
  • 校验路径一致性:对比encode_train的最终路径列表和预校验输出的合法路径列表,剔除不存在、不在合法范围内的路径项,避免非图像文件进入数据集。

内容的提问来源于stack exchange,提问作者Karmel Salah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 07:21:28