You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Keras中使用TensorFlow的nce_loss?报错求助

解决Keras中调用TensorFlow NCE Loss的TypeError问题

我来帮你搞定这个问题!你遇到的TypeError核心原因是tf.nn.nce_loss需要几个你没传入的关键参数,而且Keras的自定义loss函数默认只能拿到y_true和y_pred两个参数,没法直接把权重、偏置这类外部参数传进去。下面我一步步给你拆解解决方案:

问题根源分析

tf.nn.nce_loss的必填参数远不止你传入的三个:

  • labels:目标标签(需要是int32/int64类型的2D张量,形状为[batch_size, 1])
  • inputs:模型的输出向量(通常是解码器输出的嵌入向量)
  • weights:目标词汇表的嵌入权重矩阵(形状[num_classes, embedding_size])
  • biases:每个目标类别的偏置(形状[num_classes])
  • num_sampled:负采样的样本数
  • num_classes:目标词汇表的总大小

而Keras默认的自定义loss函数只能接收y_true和y_pred,所以我们得用闭包封装或者自定义层的方式,把额外参数传入loss计算逻辑。

解决方案一:用闭包封装NCE Loss函数

这种方法最直接,通过外层函数把需要的参数传递给内部的loss计算函数,同时满足Keras对loss函数的参数要求。

修改后的代码示例

import tensorflow as tf
from keras.models import Sequential
from keras.layers import Dropout, Masking, Embedding, LSTM
from attention_decoder import AttentionDecoder

# 先定义目标端的关键参数
input_features = 10000  # 替换成你的输入词汇表大小
input_embed_dimension = 256
n_timesteps_in = 50
LSTM_Unitsize = 512
target_embed_dim = 256
target_vocab_size = 8000  # 替换成你的目标词汇表大小
num_sampled = 100

# 定义闭包函数,封装NCE需要的参数
def make_nce_loss(num_classes, embed_dim, num_sampled=100):
    # 初始化NCE需要的权重和偏置
    nce_weights = tf.Variable(
        tf.random.normal([num_classes, embed_dim]),
        trainable=True,
        name="nce_weights"
    )
    nce_biases = tf.Variable(
        tf.zeros([num_classes]),
        trainable=True,
        name="nce_biases"
    )
    
    def keras_nce_loss(y_true, y_pred):
        # 把Keras传入的标签转换成NCE要求的2D int64张量
        y_true = tf.expand_dims(tf.cast(y_true, tf.int64), axis=-1)
        # 计算NCE Loss
        return tf.nn.nce_loss(
            labels=y_true,
            inputs=y_pred,
            weights=nce_weights,
            biases=nce_biases,
            num_sampled=num_sampled,
            num_classes=num_classes
        )
    return keras_nce_loss

# 构建模型
model2 = Sequential()
model2.add(Embedding(input_features, input_embed_dimension, input_length=n_timesteps_in, mask_zero=True))
model2.add(Dropout(0.2))
model2.add(LSTM(LSTM_Unitsize, return_sequences=True, activation='relu'))
model2.add(Masking(mask_value=0.))
# 解码器输出维度要和目标嵌入维度一致
model2.add(AttentionDecoder(LSTM_Unitsize, target_embed_dim))

# 创建带参数的自定义NCE Loss
custom_nce_loss = make_nce_loss(target_vocab_size, target_embed_dim, num_sampled)

# 编译模型
model2.compile(loss=custom_nce_loss, optimizer='adam')

关键说明

  1. 闭包函数make_nce_loss负责初始化NCE需要的权重和偏置,并返回一个符合Keras要求的loss函数。
  2. 必须把y_true转换成int64类型的2D张量,这是tf.nn.nce_loss的强制要求。
  3. 解码器的输出维度要和目标嵌入维度一致,这样才能和NCE的权重矩阵做相似度计算。

解决方案二:自定义NCE Loss层(更符合Keras架构)

如果你的翻译任务是序列到序列(Seq2Seq)场景,用自定义层+Functional API会更灵活,能更好地处理序列输入输出:

import tensorflow as tf
from keras.models import Model
from keras.layers import Input, Dropout, Masking, Embedding, LSTM, Layer
from attention_decoder import AttentionDecoder

class NCELossLayer(Layer):
    def __init__(self, num_classes, num_sampled=100, **kwargs):
        self.num_classes = num_classes
        self.num_sampled = num_sampled
        super().__init__(**kwargs)
    
    def build(self, input_shape):
        # 在build阶段初始化NCE的权重和偏置
        embed_dim = input_shape[1][-1]
        self.nce_weights = self.add_weight(
            name='nce_weights',
            shape=(self.num_classes, embed_dim),
            initializer='glorot_uniform',
            trainable=True
        )
        self.nce_biases = self.add_weight(
            name='nce_biases',
            shape=(self.num_classes,),
            initializer='zeros',
            trainable=True
        )
        super().build(input_shape)
    
    def call(self, inputs):
        # inputs是一个列表:[目标标签, 解码器输出]
        y_true, y_pred = inputs
        # 转换标签格式
        y_true = tf.expand_dims(tf.cast(y_true, tf.int64), axis=-1)
        # 计算NCE Loss并添加到模型总loss中
        loss = tf.nn.nce_loss(
            labels=y_true,
            inputs=y_pred,
            weights=self.nce_weights,
            biases=self.nce_biases,
            num_sampled=self.num_sampled,
            num_classes=self.num_classes
        )
        self.add_loss(loss)
        # 返回dummy输出,满足Keras层的输出要求
        return y_pred

# 用Functional API构建Seq2Seq模型
input_seq = Input(shape=(50,))  # 输入序列长度
target_seq = Input(shape=(50,)) # 目标序列长度

x = Embedding(10000, 256, mask_zero=True)(input_seq)
x = Dropout(0.2)(x)
x = LSTM(512, return_sequences=True, activation='relu')(x)
x = Masking(mask_value=0.)(x)
decoder_out = AttentionDecoder(512, 256)(x)

# 添加自定义NCE Loss层
output = NCELossLayer(num_classes=8000, num_sampled=100)([target_seq, decoder_out])

model = Model(inputs=[input_seq, target_seq], outputs=output)
model.compile(optimizer='adam')

额外注意事项

  • 翻译任务的准确率指标acc在NCE Loss下不好直接计算,因为NCE是负采样近似,建议自定义Top-K准确率指标来评估模型效果。
  • 如果你已经有预训练的目标词嵌入,可以把nce_weights初始化为预训练的嵌入矩阵,提升模型收敛速度。

内容的提问来源于stack exchange,提问作者Hari Prasad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:34:44