You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

训练解码器Transformer时遇AttributeError: Tensor无nested_row_splits属性

解决Transformer解码器训练中的AttributeError: 'Tensor' object has no attribute 'nested_row_splits'错误

错误现象

训练用于单词预测的解码器Transformer时,触发以下错误,更换损失函数无法解决:

AttributeError: 'Tensor' object has no attribute 'nested_row_splits'

相关代码

数据处理代码

BATCH_SIZE = 64
EPOCHS = 30
MAX_SEQUENCE_LENGTH = 394
VOCAB_SIZE = 15000
EMBED_DIM = 256
INTERMEDIATE_DIM = 512
NUM_HEADS = 8

def process_batch(ds):
    ds = tokenizer(ds)

    ## padd short senteces to max len using the [PAD] id
    ## add special tokens [START] and [END]

    ds_start_end_packer = StartEndPacker(
        sequence_length=MAX_SEQUENCE_LENGTH,
        start_value = tokenizer.token_to_id("[START]")
     )

    output = ds_start_end_packer(ds)

    return (output, ds)

def make_ds(seq):
    dataset = tf.data.Dataset.from_tensor_slices(seq)
    dataset = dataset.batch(BATCH_SIZE)
    dataset = dataset.map(process_batch, num_parallel_calls=tf.data.AUTOTUNE)
    dataset = dataset.shuffle(2048).prefetch(16).cache()

    return dataset

数据输出验证

train_ds = make_ds(train_seq)
val_ds = make_ds(val_seq)

for x,y in train_ds.take(1):
    print(f"inputs.shape: {x.shape}")
    print(f"features.shape: {y.shape}")

# 输出
inputs.shape: (64, 394)
features.shape: (64, None)

Transformer模型代码

decoder_inputs = Input(shape=(None,), dtype="int64", name="decoder_inputs")

x = TokenAndPositionEmbedding(
    vocabulary_size= VOCAB_SIZE,
    sequence_length = MAX_SEQUENCE_LENGTH,
    embedding_dim = EMBED_DIM,
    mask_zero =True
    )(decoder_inputs)

for _ in range(NUM_LAYERS):
    decoder_layer = TransformerDecoder(
        intermediate_dim = INTERMEDIATE_DIM, num_heads= NUM_HEADS
    )
    x = decoder_layer(x)

x = Dropout(0.5)(x)

decoder_ouput = Dense(VOCAB_SIZE)(x)
transformer = Model(inputs=decoder_inputs, outputs=decoder_ouput, name="transformer")

loss_fn = tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True)
perplexity = Perplexity(from_logits=True, mask_token_id = 0)

transformer.compile(optimizer = "adam", loss="mse", metrics=[perplexity])

transformer.fit(train_ds, epochs=EPOCHS,batch_size=128 ,verbose=2 , 
              validation_data=val_ds,callbacks=[model_checkpoint_callback])

完整错误堆栈

File "c:\users\hp\appdata\local\programs\python\python38\lib\site-packages\keras\engine\compile_utils.py", line 265, in __call__
    loss_value = loss_obj(y_t, y_p, sample_weight=sw)
File "c:\users\hp\appdata\local\programs\python\python38\lib\site-packages\keras\losses.py", line 152, in __call__
    losses = call_fn(y_true, y_pred)
File "c:\users\hp\appdata\local\programs\python\python38\lib\site-packages\keras\losses.py", line 272, in call  **
    return ag_fn(y_true, y_pred, **self._fn_kwargs)
File "c:\users\hp\appdata\local\programs\python\python38\lib\site-packages\keras\losses.py", line 2115, in _ragged_tensor_sparse_categorical_crossentropy
    return _ragged_tensor_apply_loss(fn, y_true, y_pred, y_pred_extra_dim=True)
File "c:\users\hp\appdata\local\programs\python\python38\lib\site-packages\keras\losses.py", line 1564, in _ragged_tensor_apply_loss
    nested_splits_list = [rt.nested_row_splits for rt in (y_true, y_pred)]
File "c:\users\hp\appdata\local\programs\python\python38\lib\site-packages\keras\losses.py", line 1564, in <listcomp>
    nested_splits_list = [rt.nested_row_splits for rt in (y_true, y_pred)]

AttributeError: 'Tensor' object has no attribute 'nested_row_splits'

错误原因

从数据输出可见,输入x是形状为(64, 394)的固定长度Tensor(已通过StartEndPacker填充),但标签y是形状为(64, None)的RaggedTensor(未做填充处理)。

损失函数在计算时,会检测到y_true是RaggedTensor,进而触发针对RaggedTensor的处理逻辑,尝试访问nested_row_splits属性,但模型输出y_pred是普通Tensor,没有该属性,因此报错。

此外,模型编译时配置的损失是mse(均方误差),但单词预测属于分类任务,应该使用定义好的loss_fn(SparseCategoricalCrossentropy),这也会加剧类型不匹配问题。

解决方案

方案1:统一输入和标签的格式,将标签转换为固定长度Tensor

修改process_batch函数,对标签也应用StartEndPacker进行填充,确保标签和输入长度一致:

def process_batch(ds):
    ds = tokenizer(ds)
    pad_id = tokenizer.token_to_id("[PAD]")
    end_id = tokenizer.token_to_id("[END]")
    
    # 处理输入:添加START,填充到MAX_SEQUENCE_LENGTH
    input_packer = StartEndPacker(
        sequence_length=MAX_SEQUENCE_LENGTH,
        start_value=tokenizer.token_to_id("[START]"),
        end_value=end_id,
        pad_value=pad_id
    )
    input_seq = input_packer(ds)
    
    # 处理标签:添加END,填充到MAX_SEQUENCE_LENGTH(单词预测任务中,目标是输入的下一个词)
    target_packer = StartEndPacker(
        sequence_length=MAX_SEQUENCE_LENGTH,
        end_value=end_id,
        pad_value=pad_id
    )
    target_seq = target_packer(ds)
    
    return (input_seq, target_seq)

方案2:将RaggedTensor标签转换为密集Tensor

如果不需要添加特殊标签,直接将RaggedTensor转换为固定长度的密集Tensor:

def process_batch(ds):
    ds = tokenizer(ds)
    ds_start_end_packer = StartEndPacker(
        sequence_length=MAX_SEQUENCE_LENGTH,
        start_value = tokenizer.token_to_id("[START]")
     )
    output = ds_start_end_packer(ds)
    
    # 将RaggedTensor标签转换为密集Tensor,用PAD填充到最大长度
    pad_id = tokenizer.token_to_id("[PAD]")
    target_seq = ds.to_tensor(default_value=pad_id, shape=(None, MAX_SEQUENCE_LENGTH))
    
    return (output, target_seq)

方案3:修正损失函数配置

模型编译时,使用定义好的分类损失函数,而非回归损失mse:

transformer.compile(optimizer = "adam", loss=loss_fn, metrics=[perplexity])

验证修改

重新生成数据集并检查形状:

for x,y in train_ds.take(1):
    print(f"inputs.shape: {x.shape}")
    print(f"features.shape: {y.shape}")
# 预期输出:
# inputs.shape: (64, 394)
# features.shape: (64, 394)

此时输入和标签均为固定长度Tensor,损失函数可以正常计算,不会触发RaggedTensor相关的错误逻辑。

内容的提问来源于stack exchange,提问作者mucho

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 01:00:59