You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用TensorFlow训练低资源语-葡翻译模型遭遇维度不兼容错误

解决TensorFlow机器翻译模型中维度不兼容的错误

我正在尝试用TensorFlow训练低资源语言与葡萄牙语的机器翻译模型,但运行代码时触发了维度不兼容的错误:

PS C:\Users\myuser\PycharmProjects\teste> python .\tensorflow_model.py                    
2024-08-23 21:29:50.839647: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations.
To enable the following instructions: SSE SSE2 SSE3 SSE4.1 SSE4.2 AVX AVX2 AVX512F AVX512_VNNI AVX512_BF16 FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags.
Traceback (most recent call last):
  File ".\tensorflow_model.py", line 52, in <module>
    dataset = tf.data.Dataset.from_tensor_slices((src_tensor, tgt_tensor)).shuffle(BUFFER_SIZE)
  File "C:\Users\myuser\PycharmProjects\teste\.venv\lib\site-packages\tensorflow\python\data\ops\dataset_ops.py", line 831, in from_tensor_slices
    return from_tensor_slices_op._from_tensor_slices(tensors, name)
  File "C:\Users\myuser\PycharmProjects\teste\.venv\lib\site-packages\tensorflow\python\data\ops\from_tensor_slices_op.py", line 25, in _from_tensor_slices
    return _TensorSliceDataset(tensors, name=name)
  File "C:\Users\myuser\PycharmProjects\teste\.venv\lib\site-packages\tensorflow\python\data\ops\from_tensor_slices_op.py", line 45, in __init__
    batch_dim.assert_is_compatible_with(
  File "C:\Users\myuser\PycharmProjects\teste\.venv\lib\site-packages\tensorflow\python\framework\tensor_shape.py", line 300, in assert_is_compatible_with
    raise ValueError("Dimensions %s and %s are not compatible" %
ValueError: Dimensions 21 and 22 are not compatible

以下是我的实现代码:

import tensorflow as tf
import numpy as np
import re
import os

# Clean data
def preprocess_sentence(sentence):
    sentence = sentence.lower().strip()
    sentence = re.sub(r"([?.!,¿])", r" \1 ", sentence)
    sentence = re.sub(r'[" "]+', " ", sentence)
    sentence = re.sub(r"[^a-zA-Z?.!,¿]+", " ", sentence)
    sentence = sentence.strip()
    sentence = '<start> ' + sentence + ' <end>'
    return sentence

#Function to load data
def load_data(file_path_src, file_path_tgt):
    src_sentences = open(file_path_src, 'r', encoding='utf-8').read().strip().split('\n')
    tgt_sentences = open(file_path_tgt, 'r', encoding='utf-8').read().strip().split('\n')

    src_sentences = [preprocess_sentence(sentence) for sentence in src_sentences]
    tgt_sentences = [preprocess_sentence(sentence) for sentence in tgt_sentences]

    return src_sentences, tgt_sentences

#load data
src_sentences, tgt_sentences = load_data('src_language.txt', 'portuguese.txt')

#Tokenization
src_tokenizer = tf.keras.preprocessing.text.Tokenizer(filters='')
tgt_tokenizer = tf.keras.preprocessing.text.Tokenizer(filters='')

src_tokenizer.fit_on_texts(src_sentences)
tgt_tokenizer.fit_on_texts(tgt_sentences)

src_tensor = src_tokenizer.texts_to_sequences(src_sentences)
tgt_tensor = tgt_tokenizer.texts_to_sequences(tgt_sentences)

src_tensor = tf.keras.preprocessing.sequence.pad_sequences(src_tensor, padding='post')
tgt_tensor = tf.keras.preprocessing.sequence.pad_sequences(tgt_tensor, padding='post')

BUFFER_SIZE = len(src_tensor)

#Creating the Dataset
dataset = tf.data.Dataset.from_tensor_slices((src_tensor, tgt_tensor)).shuffle(BUFFER_SIZE) 

问题原因

错误提示的维度不兼容,本质是源语言张量(src_tensor)和目标语言张量(tgt_tensor)的样本数量不一致。tf.data.Dataset.from_tensor_slices要求两个输入张量的第一维度(样本数)必须完全相同,否则无法一一配对生成数据集。

修复步骤

  1. 检查平行语料的行数匹配
    打开src_language.txt和portuguese.txt,确认两个文件的行数完全一致。机器翻译的平行语料必须逐行对应,比如源语言第N行必须对应目标语言第N行。如果存在空行、行数差异,直接删除多余行或补充缺失的对应翻译。

  2. 在代码中添加行数校验逻辑
    修改load_data函数,提前发现行数不匹配问题,避免后续报错:

    def load_data(file_path_src, file_path_tgt):
        src_sentences = open(file_path_src, 'r', encoding='utf-8').read().strip().split('\n')
        tgt_sentences = open(file_path_tgt, 'r', encoding='utf-8').read().strip().split('\n')
        
        # 添加行数校验
        if len(src_sentences) != len(tgt_sentences):
            raise ValueError(f"源语言数据行数({len(src_sentences)})与目标语言数据行数({len(tgt_sentences)})不匹配")
    
        src_sentences = [preprocess_sentence(sentence) for sentence in src_sentences]
        tgt_sentences = [preprocess_sentence(sentence) for sentence in tgt_sentences]
    
        return src_sentences, tgt_sentences
    
  3. 验证张量形状
    在创建数据集前,打印两个张量的形状,确认第一维度一致:

    print("源语言张量形状:", src_tensor.shape)
    print("目标语言张量形状:", tgt_tensor.shape)
    

    正常输出应该类似(N, 21)和(N, 22),其中N是样本数量,后面的数字是各自的序列长度(该长度可以不同,不影响样本配对)。

补充说明

pad_sequences处理后的序列长度(第二维度)允许不同,模型训练时可通过按batch动态padding解决。但第一维度的样本数必须严格一致,这是平行语料的基本要求。

内容的提问来源于stack exchange,提问作者kim85

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 09:23:11