Transformer机器翻译UnicodeDecodeError及夏拉达/天城文处理问题求解
解决夏拉达文/天城文机器翻译中的Unicode解码与Tensor字节形式问题
问题现象
- 解码错误:加载数据集时触发错误:
UnicodeDecodeError: 'utf-8' codec can't decode bytes in position 15-16: unexpected end of data
- Tensor格式异常:
tf_split_punct1、tf_split_punct2函数处理英文文本正常,但处理夏拉达文/天城文时返回字节形式的Tensor,示例:
tf.Tensor(b'\xf0\x91\x86\xa0\xf0\x91\x87\x80\xf0\x91\x86\xae\xf0\x91\x86\xb3\xf0\x91\x86\x81 \xf0\x91\x86\xa9\xf0\x91\x86\xa4\xf0\x91\x86\xb1\xf0\x91\x86\xb3\xf0\x91\x86\xa2\xf0\x91\x86\xa3\xf0\x91\x86\xb3\xf0\x91\x86\xa9\xf0\x91\x86\xb4 \xf0\x91\x86\xa8\xf0\x91\x86\x93\xf0\x91\x86\xae\xf0\x91\x86\xa2\xf0\x91\x87\x80\xf0\x91\x86\xae\xf0\x91\x86\xb5\xf0\x91\x86\xa0\xf0\x91\x86\xbc \xf0\x91\x86\xa8\xf0\x91\x86\xae\xf0\x91\x86\xa2\xf0\x91\x87\x80\xf0\x91\x86\xae\xf0\x91\x86\xbc\xf0\x91\x86\xb0\xf0\x91\x86\xb4\xf0\x91\x86\x9f\xf0\x91\x86\xb5\xf0\x91\x86\xa9\xf0\x91\x87\x80 \xf0\x91\x87\x86\xf0\x91\x87\x86 \xf0\x91\x86\xa4\xf0\x91\x86\xa9\xf0\x91\x86\xbe\xf0\x91\x86\xb1\xf0\x91\x87\x80\xf0\x91\x86\xa0\xf0\x91\x86\xb6 \xf0\x91\x86\xa0\xf0\x91\x86\xbc', shape=(), dtype=string)
问题根源分析
- 解码错误原因:
- 数据集文件
sha5.txt可能存在非UTF-8编码字节,或文件损坏导致字节缺失,errors='ignore'无法完全规避异常。 - 代码遗漏
random模块导入,random.shuffle(pairs)会触发NameError(用户未提及但属于潜在问题)。
- 数据集文件
- Tensor字节形式原因:
- TensorFlow字符串Tensor默认以字节形式打印,但实际存储UTF-8文本,属于显示问题;但正则表达式未正确匹配夏拉达文/天城文的Unicode范围,可能导致文本处理不彻底。
- 文本处理函数中未确保
tf_text模块被正确导入,可能引发隐式错误。
解决方案
1. 修复Unicode解码错误
- 确认文件编码:用文本编辑器(如Notepad++)查看
sha5.txt实际编码,若为非UTF-8格式,修改open函数的encoding参数对应值。 - 优化读取逻辑:逐行读取并跳过解码失败的行,避免一次性加载损坏字节块:
def load_data(fname): import random # 补上遗漏的模块导入 pairs = [] with open(fname, "rb") as f: for line in f: try: # 尝试UTF-8解码,失败则跳过该行 line_str = line.decode("utf-8").rstrip('\n') # 限制分割次数,避免句子含逗号导致拆分错误 src, trgt = line_str.split(",", maxsplit=1) pairs.append((src, trgt)) except UnicodeDecodeError: continue random.shuffle(pairs) source = [src for src, _ in pairs] target = [trgt for _, trgt in pairs] return (source, target)
2. 修正文本处理函数
- 调整正则表达式,匹配夏拉达文(
U+11800至U+1184F)和天城文(U+0900至U+097F)的正确Unicode范围,同时确保tf_text模块导入:
def tf_split_punct1(text): import tensorflow_text as tf_text # 显式导入模块 # 标准化UTF-8文本为NFKD格式 text = tf_text.normalize_utf8(text, "NFKD") # 保留夏拉达文、空格和指定标点,移除其他字符 text = tf.strings.regex_replace(text, r"[^\u11800-\u1184F\s.।॥,]", "") # 移除指定特殊字符 text = tf.strings.regex_replace(text, r'[\u118C6\u118C5\u118D0-\u118D9]', '') # 去除首尾空格并添加标记 text = tf.strings.strip(text) text = tf.strings.join(["[START]", text, "[END]"], separator=" ") return text def tf_split_punct2(text): import tensorflow_text as tf_text text = tf_text.normalize_utf8(text, "NFKD") # 保留天城文、空格和指定标点 text = tf.strings.regex_replace(text, r"[^\u0900-\u097F\s.।॥,]", "") text = tf.strings.regex_replace(text, r'[\u118C5॥]', '') # 移除天城文数字 text = tf.strings.regex_replace(text, r'[\u0966-\u096F]', '') text = tf.strings.strip(text) text = tf.strings.join(["[START]", text, "[END]"], separator=" ") return text
- 注:Tensor的字节形式显示是正常现象,若要查看实际文本,可使用
tf.print(text)或text.numpy().decode('utf-8')转换为字符串。
3. 修正TextVectorization层的函数引用
代码中sourceTextProcessor错误引用了未定义的tf_lower_and_split_punct1,需改为实际定义的函数:
sourceTextProcessor = TextVectorization( standardize=tf_split_punct1, max_tokens=SOURCE_VOCAB_SIZE, split="whitespace" ) targetTextProcessor = TextVectorization( standardize=tf_split_punct2, max_tokens=TARGET_VOCAB_SIZE, split="whitespace" )
内容的提问来源于stack exchange,提问作者Khushbu
相关产品推荐
相关产品推荐

