TensorFlow导入图像数据集时出现UnicodeDecodeError的解决求助
UnicodeDecodeError 解决:TensorFlow image_dataset_from_directory 加载狗品种数据集失败
错误原因
这个错误的核心是数据集目录或文件名包含非UTF-8编码的字符,image_dataset_from_directory在遍历文件时默认用UTF-8解码文件名,遇到不符合的字节就会抛出解码错误——和CSV导入的编码问题本质类似,但这个API本身没有直接提供编码参数,得从文件系统层面或者预处理环节解决。
可行解决方法
方法1:批量重命名文件/目录,统一为UTF-8兼容字符
手动或者用脚本把数据集里所有带特殊字符(比如非英文符号、系统本地编码字符)的文件、文件夹名改成纯英文/数字组合,避免非UTF-8编码的字节。比如Windows下可通过PowerShell脚本批量清理文件名:Get-ChildItem -Recurse | Rename-Item -NewName { $_.Name -replace '[^\w\.]', '' }注意:执行前务必备份数据集,避免误操作导致文件丢失或重命名错误。
方法2:自定义文件加载逻辑,绕过默认解码
放弃image_dataset_from_directory,改用TensorFlow低级API手动加载文件,指定正确的编码解析文件名:import tensorflow as tf import pathlib import os IMAGE_SIZE = 224 # 替换成你的实际尺寸 BATCH_SIZE = 32 data_dir = pathlib.Path(".\\dataset") class_names = [name.name for name in data_dir.iterdir() if name.is_dir()] # 遍历所有文件路径,用系统对应编码解码(比如gbk,根据你的系统调整) file_paths = [] for class_dir in data_dir.iterdir(): if class_dir.is_dir(): for img_path in class_dir.glob("*"): # 替换为你的文件实际编码,比如gbk、latin-1等 decoded_path = str(img_path).encode('latin-1').decode('gbk') file_paths.append(decoded_path) # 定义加载函数 def load_image(path): img = tf.io.read_file(path) img = tf.image.decode_jpeg(img, channels=3) img = tf.image.resize(img, (IMAGE_SIZE, IMAGE_SIZE)) # 获取标签并转为独热编码 label_str = tf.strings.split(path, os.sep)[-2] label_idx = tf.argmax(tf.equal(class_names, label_str)) return img, tf.one_hot(label_idx, depth=len(class_names)) # 构建数据集 dataset = tf.data.Dataset.from_tensor_slices(file_paths) dataset = dataset.map(load_image, num_parallel_calls=tf.data.AUTOTUNE) dataset = dataset.shuffle(1000, seed=123).batch(BATCH_SIZE).prefetch(tf.data.AUTOTUNE)这里的编码(比如
gbk)需要根据你系统的实际文件编码调整:Windows中文系统常用gbk,Linux/macOS一般默认UTF-8。方法3:重新解压数据集时指定编码
如果是下载的压缩包解压后出现文件名乱码,重新解压时指定对应字符集。比如用7-Zip解压时,在「选项-字符集」中选择与压缩包匹配的编码(如GB2312),避免文件名解码错误。
内容的提问来源于stack exchange,提问作者Harshaditya sharma
相关产品推荐
相关产品推荐

