基于Keras LSTM的游戏物品序列预测及3D数组预处理构建
嘿,能上手LSTM做游戏物品序列预测,作为Keras新手已经很棒啦!我来一步步帮你搞定从Pandas DataFrame构建LSTM需要的3D数组这件事~
先明确LSTM的输入要求
LSTM层接收的输入是3D张量,格式为:(样本数, 时间步长, 特征数)
- 样本数:每个独立的序列(这里就是每个
gameId+side组合,对应同一场游戏里同一阵营的物品购买序列) - 时间步长:序列的长度(你要预测前十件,所以我们统一把序列长度固定为10)
- 特征数:每个时间步的特征维度(比如编码后的物品ID是1维,one-hot编码后是物品种类数维)
你的DataFrame已经按gameId、side、timestamp排好序了,这给我们省了不少事,直接分组就行!
步骤1:分组提取物品序列
首先把同一个gameId+side的物品ID聚合成一个序列:
import pandas as pd # 假设你的DataFrame叫df grouped = df.groupby(['gameId', 'side'])['itemId'].apply(list).reset_index(name='item_sequence')
这一步会生成一个新的DataFrame,每一行对应一个gameId+side的物品购买序列。
步骤2:编码离散的ItemId
LSTM不能直接处理离散的物品ID,得把它转换成数值特征,这里给你两种常用方案:
方案A:LabelEncoder(适合搭配Embedding层)
把每个唯一的ItemId映射成一个整数索引,后续可以用Keras的Embedding层把它转换成低维向量,适合物品种类多的场景:
from sklearn.preprocessing import LabelEncoder le = LabelEncoder() # 先给原始DataFrame的itemId做编码 df['itemId_encoded'] = le.fit_transform(df['itemId']) # 再分组提取编码后的序列 grouped_encoded = df.groupby(['gameId', 'side'])['itemId_encoded'].apply(list).reset_index(name='encoded_sequence')
方案B:OneHotEncoder(适合直接输入LSTM)
把每个ItemId转换成独热向量,每个时间步的特征数等于物品的总种类数,适合物品种类较少的场景:
from sklearn.preprocessing import OneHotEncoder import numpy as np # 把itemId转换成2D数组(OneHotEncoder要求输入是2D) item_ids = df['itemId'].values.reshape(-1, 1) ohe = OneHotEncoder(sparse_output=False) item_onehot = ohe.fit_transform(item_ids) # 把独热编码的结果合并回原DataFrame onehot_cols = ohe.get_feature_names_out(['itemId']) df = pd.concat([df, pd.DataFrame(item_onehot, columns=onehot_cols)], axis=1) # 分组提取独热序列,同时统一长度为10 def extract_onehot_seq(group): seq = group[onehot_cols].values # 不足10个物品就补0,超过10个就截断前10个 if len(seq) < 10: seq = np.pad(seq, ((0, 10 - len(seq)), (0, 0)), mode='constant') else: seq = seq[:10] return seq grouped_onehot = df.groupby(['gameId', 'side']).apply(extract_onehot_seq).tolist()
步骤3:构建3D输入数组
现在把处理好的序列转换成LSTM需要的3D格式:
用LabelEncoder编码的情况
from tensorflow.keras.preprocessing.sequence import pad_sequences import numpy as np # 提取所有编码后的序列,统一长度为10(不足补0,超过截断) sequences = grouped_encoded['encoded_sequence'].tolist() padded_sequences = pad_sequences(sequences, maxlen=10, padding='post', truncating='post') # 转换成3D数组:(样本数, 10, 1),最后一维是特征数(这里每个时间步是1个整数) X = np.expand_dims(padded_sequences, axis=-1)
用OneHotEncoder编码的情况
直接把分组后的列表转换成numpy数组就行,已经是3D格式了:
X = np.array(grouped_onehot) # 形状是:(样本数, 10, 物品种类数)
进阶:生成序列预测的输入&目标
如果你的目标是用前面的物品预测后面的物品(比如用前1-9个物品预测第2-10个),需要拆分输入和标签:
# 假设用LabelEncoder的padded_sequences(形状:(样本数,10)) X_input = padded_sequences[:, :-1] # 输入:前9个物品,形状(样本数,9) y_target = padded_sequences[:, 1:] # 目标:后9个物品,形状(样本数,9) # 转换成LSTM需要的3D输入 X_input = np.expand_dims(X_input, axis=-1) # 如果目标要做分类(每个时间步预测ItemId),可以把y_target转换成独热编码 y_target_onehot = ohe.transform(y_target.reshape(-1,1)).reshape(y_target.shape[0], y_target.shape[1], -1)
新手小贴士
- 填充方式:
pad_sequences里的padding='post'是在序列末尾补0,padding='pre'是在开头补0,根据你的场景选(一般序列预测用post更合理); - 物品序列不足10的情况:如果有的
gameId+side购买的物品少于10件,你可以选择填充,或者直接过滤掉这些样本,取决于你的数据集大小; - Embedding层的使用:如果用LabelEncoder,后续在Keras模型里加一层
Embedding(input_dim=len(le.classes_), output_dim=32),把整数索引转换成32维向量,比独热编码高效得多。
内容的提问来源于stack exchange,提问作者Dwayne Hart
相关产品推荐
相关产品推荐

