You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决tf.keras.utils.image_dataset_from_directory的无效续字节错误?

解决Windows下TensorFlow处理非拉丁文件名的UnicodeDecodeError问题

问题根源在于Ubuntu默认使用UTF-8编码,而Windows系统默认编码(如GBK/CP1252)与UTF-8不兼容,导致TensorFlow的image_dataset_from_directory无法正确解析阿拉伯语/波斯语文件名。以下是几种无需重命名文件的解决方案:

1. 开启Windows系统UTF-8全局支持

这是最彻底的解决方式,让Windows和Linux保持一致的编码环境:

  • 打开「控制面板」→「区域」→「管理」标签页
  • 点击「更改系统区域设置」,勾选「Beta版:使用Unicode UTF-8提供全球语言支持」
  • 重启系统后,重新运行代码

2. 使用pathlib处理路径

用Python的pathlib模块替代字符串路径,它能更好地处理跨平台编码问题:

from pathlib import Path
import tensorflow as tf

image_path = Path('D:/Desktop/tfmm/Personas/')
train_ds = tf.keras.utils.image_dataset_from_directory(
    image_path,
    labels="inferred",
    label_mode="int",
    color_mode="rgb",
    batch_size=32,
    image_size=(256, 256),
    shuffle=True,
    seed=123,
    validation_split=0.2,
    subset='training',
    interpolation="bilinear",
    follow_links=False,
    crop_to_aspect_ratio=False
)

3. 自定义数据加载管道

如果上述方法无效,可以绕过image_dataset_from_directory,手动实现数据加载流程,完全控制编码处理:

import tensorflow as tf
from pathlib import Path
import os

def load_image(file_path):
    # 从父文件夹获取标签并转换为整数
    class_names = tf.constant([dir.name for dir in Path('D:/Desktop/tfmm/Personas/').iterdir() if dir.is_dir()])
    label = tf.strings.split(file_path, os.sep)[-2]
    label = tf.argmax(tf.equal(class_names, label))
    
    # 读取并处理图片
    img = tf.io.read_file(file_path)
    img = tf.image.decode_jpeg(img, channels=3)  # 根据图片格式调整为decode_png等
    img = tf.image.resize(img, (256, 256))
    return img, label

# 收集所有图片路径
image_path = Path('D:/Desktop/tfmm/Personas/')
file_paths = list(image_path.glob('*/*.jpg')) + list(image_path.glob('*/*.png'))

# 构建TensorFlow数据集
ds = tf.data.Dataset.from_tensor_slices([str(p) for p in file_paths])
ds = ds.map(load_image, num_parallel_calls=tf.data.AUTOTUNE)

# 划分训练/验证集
train_size = int(0.8 * len(file_paths))
train_ds = ds.take(train_size).shuffle(1000).batch(32)
val_ds = ds.skip(train_size).batch(32)

4. 检测并转换文件名编码(可选)

如果系统UTF-8支持无法开启,可以先检测文件名的实际编码,再批量转换为UTF-8:

import chardet
from pathlib import Path

image_path = Path('D:/Desktop/tfmm/Personas/')
for file in image_path.rglob('*'):
    if file.is_file():
        # 检测文件名编码
        raw_name = file.name.encode('raw_unicode_escape')
        detect_result = chardet.detect(raw_name)
        if detect_result['encoding'] and detect_result['encoding'] != 'utf-8':
            # 转换为UTF-8并重命名
            new_name = raw_name.decode(detect_result['encoding']).encode('utf-8').decode('utf-8')
            file.rename(file.parent / new_name)

内容的提问来源于stack exchange,提问作者SeyyedMohammadAmin Mousavi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 13:35:35