如何解决tf.keras.utils.image_dataset_from_directory的无效续字节错误?
解决Windows下TensorFlow处理非拉丁文件名的UnicodeDecodeError问题
问题根源在于Ubuntu默认使用UTF-8编码,而Windows系统默认编码(如GBK/CP1252)与UTF-8不兼容,导致TensorFlow的image_dataset_from_directory无法正确解析阿拉伯语/波斯语文件名。以下是几种无需重命名文件的解决方案:
1. 开启Windows系统UTF-8全局支持
这是最彻底的解决方式,让Windows和Linux保持一致的编码环境:
- 打开「控制面板」→「区域」→「管理」标签页
- 点击「更改系统区域设置」,勾选「Beta版:使用Unicode UTF-8提供全球语言支持」
- 重启系统后,重新运行代码
2. 使用pathlib处理路径
用Python的pathlib模块替代字符串路径,它能更好地处理跨平台编码问题:
from pathlib import Path import tensorflow as tf image_path = Path('D:/Desktop/tfmm/Personas/') train_ds = tf.keras.utils.image_dataset_from_directory( image_path, labels="inferred", label_mode="int", color_mode="rgb", batch_size=32, image_size=(256, 256), shuffle=True, seed=123, validation_split=0.2, subset='training', interpolation="bilinear", follow_links=False, crop_to_aspect_ratio=False )
3. 自定义数据加载管道
如果上述方法无效,可以绕过image_dataset_from_directory,手动实现数据加载流程,完全控制编码处理:
import tensorflow as tf from pathlib import Path import os def load_image(file_path): # 从父文件夹获取标签并转换为整数 class_names = tf.constant([dir.name for dir in Path('D:/Desktop/tfmm/Personas/').iterdir() if dir.is_dir()]) label = tf.strings.split(file_path, os.sep)[-2] label = tf.argmax(tf.equal(class_names, label)) # 读取并处理图片 img = tf.io.read_file(file_path) img = tf.image.decode_jpeg(img, channels=3) # 根据图片格式调整为decode_png等 img = tf.image.resize(img, (256, 256)) return img, label # 收集所有图片路径 image_path = Path('D:/Desktop/tfmm/Personas/') file_paths = list(image_path.glob('*/*.jpg')) + list(image_path.glob('*/*.png')) # 构建TensorFlow数据集 ds = tf.data.Dataset.from_tensor_slices([str(p) for p in file_paths]) ds = ds.map(load_image, num_parallel_calls=tf.data.AUTOTUNE) # 划分训练/验证集 train_size = int(0.8 * len(file_paths)) train_ds = ds.take(train_size).shuffle(1000).batch(32) val_ds = ds.skip(train_size).batch(32)
4. 检测并转换文件名编码(可选)
如果系统UTF-8支持无法开启,可以先检测文件名的实际编码,再批量转换为UTF-8:
import chardet from pathlib import Path image_path = Path('D:/Desktop/tfmm/Personas/') for file in image_path.rglob('*'): if file.is_file(): # 检测文件名编码 raw_name = file.name.encode('raw_unicode_escape') detect_result = chardet.detect(raw_name) if detect_result['encoding'] and detect_result['encoding'] != 'utf-8': # 转换为UTF-8并重命名 new_name = raw_name.decode(detect_result['encoding']).encode('utf-8').decode('utf-8') file.rename(file.parent / new_name)
内容的提问来源于stack exchange,提问作者SeyyedMohammadAmin Mousavi
相关产品推荐
相关产品推荐

