You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras自定义Attention层接入BiGRU输出时报float()参数为tuple的TypeError

问题场景

使用Keras搭建序列模型,组网流程为:Input层接收定长为10的整数序列输入 → Embedding层生成嵌入向量 → BiGRU层提取序列上下文特征 → 自定义Attention层做特征加权聚合,运行代码时触发如下报错:
TypeError: float() argument must be a string or a number, not 'tuple'
已确认传入自定义Attention层的输入为BiGRU输出的合法Tensor类型,无法定位错误来源。

核心组网代码

u_in = L.Input(name="user", shape=(10,), dtype="int32")
# 嵌入层
u = u_in
u = self.embedding_layer(u)  
hidden_size= 128
u_gru = Bidirectional(GRU(hidden_size, return_sequences=True))(u)
u_att = Attention_layer()(u_gru)
print(u_att)

自定义Attention层实现代码

class Attention_layer(Layer):
    def __init__(self,
                 W_regularizer=None, b_regularizer=None,
                 W_constraint=None, b_constraint=None,
                 bias=True, **kwargs):
        self.supports_masking = True
        self.init = initializers.get('glorot_uniform')
        self.W_regularizer = regularizers.get(W_regularizer)
        self.b_regularizer = regularizers.get(b_regularizer)

        self.W_constraint = constraints.get(W_constraint)
        self.b_constraint = constraints.get(b_constraint)

        self.bias = bias
        super(Attention_layer, self).__init__(**kwargs)

    def build(self, input_shape):
        assert len(input_shape) == 3
        self.W = self.add_weight(name='att_weight',shape=(input_shape[-1], input_shape[-1],),
                                 initializer=self.init,
                                 regularizer=self.W_regularizer,
                                 constraint=self.W_constraint   )

        if self.bias:
            self.b = self.add_weight((input_shape[-1],),
                                     initializer='zero',
                                     name='{}_b'.format(self.name),
                                     regularizer=self.b_regularizer,
                                     constraint=self.b_constraint)

        super(Attention_layer, self).build(input_shape)
    def compute_mask(self, input, input_mask=None):
        # 不向后续层传递mask
        return None

    def call(self, x, mask=None):
        uit = K.dot(x, self.W)
        if self.bias:
            uit += self.b
        uit = K.tanh(uit)
        a = K.exp(uit)
        if mask is not None:
            a *= K.cast(mask, K.floatx())
        a /= K.cast(K.sum(a, axis=1, keepdims=True) + K.epsilon(), K.floatx())
        print(a)
        print(x)
        weighted_input = x * a
        print(weighted_input)
        return K.sum(weighted_input, axis=1)
    def compute_output_shape(self, input_shape):
        return (input_shape[0], input_shape[-1])

报错原因

存在两个核心问题,其中第一个直接触发该类型错误:

  • 自定义层权重初始化传参错误:定义偏置项self.b调用add_weight时,第一个位置参数直接传入了形状元组(input_shape[-1],),没有显式指定shape=参数名。在当前使用的Keras版本中,add_weight的第一个位置参数为权重名称(字符串类型),框架尝试将传入的元组转换为参数要求的数值/字符串类型时直接触发类型错误。同时此处初始化器传值错误,零初始化的合法标识为'zeros',而非'zero'。
  • 注意力分数计算逻辑错误:现有实现缺少注意力打分必需的上下文投影向量:计算得到tanh激活后的投影矩阵uit(形状为[batch_size, seq_len, hidden_dim*2],BiGRU为双向输出所以维度是2倍隐层大小)后,没有将其映射为每个时间步的单维注意力分数,直接对三维矩阵做指数运算和归一化,得到的权重a是三维矩阵,既不符合注意力权重的形状要求,也会在和二维mask做运算、后续加权求和时出现逻辑错误,即使解决类型错误也无法得到正确的注意力加权结果。

修复方案

按照以下步骤修改代码即可解决问题:

  1. 修正build方法中偏置项的初始化传参,补全缺失的上下文投影权重,修正初始化器名称:
def build(self, input_shape):
    assert len(input_shape) == 3
    self.W = self.add_weight(name='att_weight',
                            shape=(input_shape[-1], input_shape[-1],),
                            initializer=self.init,
                            regularizer=self.W_regularizer,
                            constraint=self.W_constraint)
    # 新增注意力打分用的上下文向量
    self.v = self.add_weight(name='att_context',
                            shape=(input_shape[-1], 1),
                            initializer=self.init,
                            regularizer=self.W_regularizer,
                            constraint=self.W_constraint)
    if self.bias:
        # 修正传参:显式指定shape参数,修正零初始化标识为zeros
        self.b = self.add_weight(name='{}_b'.format(self.name),
                                shape=(input_shape[-1],),
                                initializer='zeros',
                                regularizer=self.b_regularizer,
                                constraint=self.b_constraint)

    super(Attention_layer, self).build(input_shape)
  1. 修正call方法中的注意力分数计算逻辑,将三维投影矩阵映射为二维注意力分数,再做softmax归一化和加权求和:
def call(self, x, mask=None):
    uit = K.dot(x, self.W)
    if self.bias:
        uit += self.b
    uit = K.tanh(uit)
    # 投影得到单维注意力分数,形状为[batch_size, seq_len]
    e = K.squeeze(K.dot(uit, self.v), axis=-1)
    a = K.exp(e)
    # 处理mask
    if mask is not None:
        a *= K.cast(mask, K.floatx())
    # 按时间步维度做softmax归一化
    a /= K.cast(K.sum(a, axis=1, keepdims=True) + K.epsilon(), K.floatx())
    # 扩展权重维度适配输入形状[batch, seq_len, dim]
    a = K.expand_dims(a)
    weighted_input = x * a
    # 加权求和得到聚合后的特征
    return K.sum(weighted_input, axis=1)
  1. (可选)删除call方法中临时添加的print语句,避免图模式下打印无意义的张量占位信息干扰调试。

内容的提问来源于stack exchange,提问作者meimei

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.01 22:33:24