You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Transformer机器翻译UnicodeDecodeError及夏拉达/天城文处理问题求解

解决夏拉达文/天城文机器翻译中的Unicode解码与Tensor字节形式问题

问题现象

  1. 解码错误:加载数据集时触发错误:

UnicodeDecodeError: 'utf-8' codec can't decode bytes in position 15-16: unexpected end of data

  1. Tensor格式异常:tf_split_punct1、tf_split_punct2函数处理英文文本正常,但处理夏拉达文/天城文时返回字节形式的Tensor,示例:
tf.Tensor(b'\xf0\x91\x86\xa0\xf0\x91\x87\x80\xf0\x91\x86\xae\xf0\x91\x86\xb3\xf0\x91\x86\x81 \xf0\x91\x86\xa9\xf0\x91\x86\xa4\xf0\x91\x86\xb1\xf0\x91\x86\xb3\xf0\x91\x86\xa2\xf0\x91\x86\xa3\xf0\x91\x86\xb3\xf0\x91\x86\xa9\xf0\x91\x86\xb4 \xf0\x91\x86\xa8\xf0\x91\x86\x93\xf0\x91\x86\xae\xf0\x91\x86\xa2\xf0\x91\x87\x80\xf0\x91\x86\xae\xf0\x91\x86\xb5\xf0\x91\x86\xa0\xf0\x91\x86\xbc \xf0\x91\x86\xa8\xf0\x91\x86\xae\xf0\x91\x86\xa2\xf0\x91\x87\x80\xf0\x91\x86\xae\xf0\x91\x86\xbc\xf0\x91\x86\xb0\xf0\x91\x86\xb4\xf0\x91\x86\x9f\xf0\x91\x86\xb5\xf0\x91\x86\xa9\xf0\x91\x87\x80 \xf0\x91\x87\x86\xf0\x91\x87\x86 \xf0\x91\x86\xa4\xf0\x91\x86\xa9\xf0\x91\x86\xbe\xf0\x91\x86\xb1\xf0\x91\x87\x80\xf0\x91\x86\xa0\xf0\x91\x86\xb6 \xf0\x91\x86\xa0\xf0\x91\x86\xbc', shape=(), dtype=string)

问题根源分析

  1. 解码错误原因:
    • 数据集文件sha5.txt可能存在非UTF-8编码字节,或文件损坏导致字节缺失,errors='ignore'无法完全规避异常。
    • 代码遗漏random模块导入,random.shuffle(pairs)会触发NameError(用户未提及但属于潜在问题)。
  2. Tensor字节形式原因:
    • TensorFlow字符串Tensor默认以字节形式打印,但实际存储UTF-8文本,属于显示问题;但正则表达式未正确匹配夏拉达文/天城文的Unicode范围,可能导致文本处理不彻底。
    • 文本处理函数中未确保tf_text模块被正确导入,可能引发隐式错误。

解决方案

1. 修复Unicode解码错误

  • 确认文件编码:用文本编辑器(如Notepad++)查看sha5.txt实际编码,若为非UTF-8格式,修改open函数的encoding参数对应值。
  • 优化读取逻辑:逐行读取并跳过解码失败的行,避免一次性加载损坏字节块:
def load_data(fname):
    import random  # 补上遗漏的模块导入
    pairs = []
    with open(fname, "rb") as f:
        for line in f:
            try:
                # 尝试UTF-8解码,失败则跳过该行
                line_str = line.decode("utf-8").rstrip('\n')
                # 限制分割次数,避免句子含逗号导致拆分错误
                src, trgt = line_str.split(",", maxsplit=1)
                pairs.append((src, trgt))
            except UnicodeDecodeError:
                continue
    random.shuffle(pairs)
    source = [src for src, _ in pairs]
    target = [trgt for _, trgt in pairs]
    return (source, target)

2. 修正文本处理函数

  • 调整正则表达式,匹配夏拉达文(U+11800至U+1184F)和天城文(U+0900至U+097F)的正确Unicode范围,同时确保tf_text模块导入:
def tf_split_punct1(text):
    import tensorflow_text as tf_text  # 显式导入模块
    # 标准化UTF-8文本为NFKD格式
    text = tf_text.normalize_utf8(text, "NFKD")
    # 保留夏拉达文、空格和指定标点,移除其他字符
    text = tf.strings.regex_replace(text, r"[^\u11800-\u1184F\s.।॥,]", "")
    # 移除指定特殊字符
    text = tf.strings.regex_replace(text, r'[\u118C6\u118C5\u118D0-\u118D9]', '')
    # 去除首尾空格并添加标记
    text = tf.strings.strip(text)
    text = tf.strings.join(["[START]", text, "[END]"], separator=" ")
    return text

def tf_split_punct2(text):
    import tensorflow_text as tf_text
    text = tf_text.normalize_utf8(text, "NFKD")
    # 保留天城文、空格和指定标点
    text = tf.strings.regex_replace(text, r"[^\u0900-\u097F\s.।॥,]", "")
    text = tf.strings.regex_replace(text, r'[\u118C5॥]', '')
    # 移除天城文数字
    text = tf.strings.regex_replace(text, r'[\u0966-\u096F]', '')
    text = tf.strings.strip(text)
    text = tf.strings.join(["[START]", text, "[END]"], separator=" ")
    return text
  • 注:Tensor的字节形式显示是正常现象,若要查看实际文本,可使用tf.print(text)或text.numpy().decode('utf-8')转换为字符串。

3. 修正TextVectorization层的函数引用

代码中sourceTextProcessor错误引用了未定义的tf_lower_and_split_punct1,需改为实际定义的函数:

sourceTextProcessor = TextVectorization(
    standardize=tf_split_punct1, max_tokens=SOURCE_VOCAB_SIZE, split="whitespace"
)
targetTextProcessor = TextVectorization(
    standardize=tf_split_punct2, max_tokens=TARGET_VOCAB_SIZE, split="whitespace"
)

内容的提问来源于stack exchange,提问作者Khushbu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 09:49:51