You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow导入图像数据集时出现UnicodeDecodeError的解决求助

UnicodeDecodeError 解决:TensorFlow image_dataset_from_directory 加载狗品种数据集失败

错误原因

这个错误的核心是数据集目录或文件名包含非UTF-8编码的字符,image_dataset_from_directory在遍历文件时默认用UTF-8解码文件名,遇到不符合的字节就会抛出解码错误——和CSV导入的编码问题本质类似,但这个API本身没有直接提供编码参数,得从文件系统层面或者预处理环节解决。

可行解决方法

  • 方法1:批量重命名文件/目录,统一为UTF-8兼容字符
    手动或者用脚本把数据集里所有带特殊字符(比如非英文符号、系统本地编码字符)的文件、文件夹名改成纯英文/数字组合,避免非UTF-8编码的字节。比如Windows下可通过PowerShell脚本批量清理文件名:

    Get-ChildItem -Recurse | Rename-Item -NewName { $_.Name -replace '[^\w\.]', '' }
    

    注意:执行前务必备份数据集,避免误操作导致文件丢失或重命名错误。

  • 方法2:自定义文件加载逻辑,绕过默认解码
    放弃image_dataset_from_directory,改用TensorFlow低级API手动加载文件,指定正确的编码解析文件名:

    import tensorflow as tf
    import pathlib
    import os
    
    IMAGE_SIZE = 224  # 替换成你的实际尺寸
    BATCH_SIZE = 32
    
    data_dir = pathlib.Path(".\\dataset")
    class_names = [name.name for name in data_dir.iterdir() if name.is_dir()]
    # 遍历所有文件路径,用系统对应编码解码(比如gbk,根据你的系统调整)
    file_paths = []
    for class_dir in data_dir.iterdir():
        if class_dir.is_dir():
            for img_path in class_dir.glob("*"):
                # 替换为你的文件实际编码,比如gbk、latin-1等
                decoded_path = str(img_path).encode('latin-1').decode('gbk')
                file_paths.append(decoded_path)
    
    # 定义加载函数
    def load_image(path):
        img = tf.io.read_file(path)
        img = tf.image.decode_jpeg(img, channels=3)
        img = tf.image.resize(img, (IMAGE_SIZE, IMAGE_SIZE))
        # 获取标签并转为独热编码
        label_str = tf.strings.split(path, os.sep)[-2]
        label_idx = tf.argmax(tf.equal(class_names, label_str))
        return img, tf.one_hot(label_idx, depth=len(class_names))
    
    # 构建数据集
    dataset = tf.data.Dataset.from_tensor_slices(file_paths)
    dataset = dataset.map(load_image, num_parallel_calls=tf.data.AUTOTUNE)
    dataset = dataset.shuffle(1000, seed=123).batch(BATCH_SIZE).prefetch(tf.data.AUTOTUNE)
    

    这里的编码(比如gbk)需要根据你系统的实际文件编码调整:Windows中文系统常用gbk,Linux/macOS一般默认UTF-8。

  • 方法3:重新解压数据集时指定编码
    如果是下载的压缩包解压后出现文件名乱码,重新解压时指定对应字符集。比如用7-Zip解压时,在「选项-字符集」中选择与压缩包匹配的编码(如GB2312),避免文件名解码错误。

内容的提问来源于stack exchange,提问作者Harshaditya sharma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 20:14:59