You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Attention提取100维特征中32维后输入LSTM的代码实现求助

问题描述

我有一个包含100个特征的数据集,已有数据部分预览图。
我希望通过Attention机制从这100个特征中提取32个特征,再输入到LSTM中,实现方案与arxiv平台编号为1902.11074的论文中的做法一致。
我想要复现论文给出的相同架构,但无法编写对应的实现代码,目前已完成的代码如下:

import tensorflow as tf
from sklearn.model_selection import train_test_split

X = data.iloc[:, 1:101].values
y = data.iloc[:, 0].values

# 拆分数据集
X_train, X_test, y_train, y_test = train_test_split(X, y, 
                                                    test_size=0.2,                                                                                                                                                             
                                                    random_state=2)

X_train = X_train.reshape(-1, 1, 100)
X_test = X_test.reshape(-1, 1, 100)

print("训练集和测试集维度: ", X_train.shape, X_test.shape, y_train.shape, y_test.shape)

# 搭建模型
model = tf.keras.Sequential()
model.add(tf.keras.layers.LSTM(128, return_sequences=True, input_shape=(1, 100)))
model.add(tf.keras.layers.Dropout(0.3))
model.add(tf.keras.layers.LSTM(32, return_sequences=False))
model.add(tf.keras.layers.Dropout(0.3))
model.add(tf.keras.layers.Dense(1, activation = 'linear'))

# 编译模型
optimizer = tf.keras.optimizers.Adam(learning_rate=0.01,
                                     beta_1=0.9,
                                     beta_2=0.999,
                                     epsilon=1e-7)
model.compile(loss='mean_absolute_error',
              optimizer=optimizer)

# 训练模型
history = model.fit(X_train, y_train, 
                    epochs=30,
                    batch_size=64,
                    verbose=1,
                    validation_split=0.2,
                    shuffle=True)

相关运行细节信息如下:

训练集和测试集维度:  (1890, 1, 100) (473, 1, 100) (1890,) (473,)

Model: "sequential"
_________________________________________________________________
Layer (type)                 Output Shape              Param #   
=================================================================
lstm (LSTM)                  (None, 1, 128)            117248    
_________________________________________________________________
dropout (Dropout)            (None, 1, 128)            0         
_________________________________________________________________
lstm_1 (LSTM)                (None, 32)                20608     
_________________________________________________________________
dropout_1 (Dropout)          (None, 32)                0         
_________________________________________________________________
dense (Dense)                (None, 1)                 33        
=================================================================
Total params: 137,889
Trainable params: 137,889
Non-trainable params: 0
_________________________________________________________________

内容的提问来源于stack exchange,提问作者Sparsh Gupta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 19:45:05