You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在表格数据分类的Keras模型中添加Attention层?

问题:如何在Keras表格分类模型中正确添加Attention层?

我想用表格数据集构建常规分类模型,尝试通过添加Attention层提升精度。现有Keras Sequential模型代码如下:

input_features_size = X_train.shape[1]

layers = [
    tf.keras.Input(shape = input_features_size),
    tf.keras.layers.Dense(64, activation = 'relu', name = 'first_layer'),
    tf.keras.layers.Dense(128, activation = 'relu', name = 'second_layer'),
    tf.keras.layers.BatchNormalization(axis = 1),
    tf.keras.layers.Dense(1, activation = 'sigmoid', name = 'output_layer')
]

metrics = [
    tf.keras.metrics.BinaryAccuracy(name = 'accuracy'),
    tf.keras.metrics.Precision(name = 'precision'),
    tf.keras.metrics.Recall(name = 'recall')
]

NUM_EPOCHS = 20

deep_learning_model = Sequential(layers = layers, name = 'DL_Classifier')
deep_learning_model.compile(
    loss = binary_crossentropy,
    optimizer = Adam(learning_rate = 1e-4),
    metrics = metrics
)

尝试添加tf.keras.layers.Attention()到layers列表时,出现错误:Attention layer must be called on a list of inputs, namely [query, value],请问该如何正确添加Attention层?


解决方案

原生的tf.keras.layers.Attention层需要接收两个输入张量(query和value),而Sequential模型是线性的单输入单输出结构,无法满足这种多输入要求。因此需要改用Keras函数式API来构建模型,这样可以灵活处理多输入、分支结构。

针对表格数据的扁平特征向量,我们可以使用自注意力机制——将同一层的输出同时作为query和value传入Attention层,让模型学习特征之间的关联权重,自动聚焦对分类更重要的特征。

修改后的完整代码

import tensorflow as tf
from tensorflow.keras.models import Model
from tensorflow.keras.layers import Input, Dense, BatchNormalization, Attention
from tensorflow.keras.optimizers import Adam
from tensorflow.keras.losses import binary_crossentropy

input_features_size = X_train.shape[1]

# 用函数式API构建模型
inputs = Input(shape=input_features_size)
x = Dense(64, activation='relu', name='first_layer')(inputs)
x = Dense(128, activation='relu', name='second_layer')(x)
x = BatchNormalization(axis=1)(x)

# 添加自注意力层:将当前特征向量同时作为query和value传入
attention_layer = Attention()
attended_output = attention_layer([x, x])

# 连接输出层
outputs = Dense(1, activation='sigmoid', name='output_layer')(attended_output)

# 定义完整模型
deep_learning_model = Model(inputs=inputs, outputs=outputs, name='DL_Classifier_With_Attention')

# 编译模型
metrics = [
    tf.keras.metrics.BinaryAccuracy(name='accuracy'),
    tf.keras.metrics.Precision(name='precision'),
    tf.keras.metrics.Recall(name='recall')
]
NUM_EPOCHS = 20

deep_learning_model.compile(
    loss=binary_crossentropy,
    optimizer=Adam(learning_rate=1e-4),
    metrics=metrics
)

# 查看模型结构
deep_learning_model.summary()

额外优化建议

  • 可以在Attention层前后添加tf.keras.layers.Dropout(rate=0.2),防止模型过拟合;
  • 若想调整注意力计算逻辑,可修改Attention层的use_scale参数(是否缩放注意力分数),或自定义注意力分数的计算方式。

内容的提问来源于stack exchange,提问作者JSVJ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 16:03:37