如何将pandas数据框中不同类型变量映射到自动编码器对应输入层
问题背景
我有25个混合类型变量:部分为二值变量,部分为连续变量,绝大多数是需要做嵌入(embedding)的高基数类别因子。
我搭建了一个接收多输入的深度学习自编码器(autoencoder)模型,模型结构如下:
autoencoder.summary() Model: "claims_ae" __________________________________________________________________________________________________ Layer (type) Output Shape Param # Connected to ================================================================================================== input_provid (InputLayer) [(None, 1)] 0 __________________________________________________________________________________________________ input_pos_code (InputLayer) [(None, 1)] 0 __________________________________________________________________________________________________ input_prindiag (InputLayer) [(None, 1)] 0 __________________________________________________________________________________________________ input_billtype2 (InputLayer) [(None, 1)] 0 __________________________________________________________________________________________________ input_lob (InputLayer) [(None, 1)] 0 __________________________________________________________________________________________________ input_ppg_code (InputLayer) [(None, 1)] 0 __________________________________________________________________________________________________ input_segment (InputLayer) [(None, 1)] 0 __________________________________________________________________________________________________ input_dofr (InputLayer) [(None, 1)] 0 __________________________________________________________________________________________________ embedding_8 (Embedding) (None, 1, 70) 337820 input_provid[0][0] __________________________________________________________________________________________________ embedding_9 (Embedding) (None, 1, 6) 168 input_pos_code[0][0] __________________________________________________________________________________________________ embedding_10 (Embedding) (None, 1, 77) 447524 input_prindiag[0][0] __________________________________________________________________________________________________ embedding_11 (Embedding) (None, 1, 5) 95 input_billtype2[0][0] __________________________________________________________________________________________________ embedding_12 (Embedding) (None, 1, 3) 24 input_lob[0][0] __________________________________________________________________________________________________ embedding_13 (Embedding) (None, 1, 10) 930 input_ppg_code[0][0] __________________________________________________________________________________________________ embedding_14 (Embedding) (None, 1, 2) 8 input_segment[0][0] __________________________________________________________________________________________________ embedding_15 (Embedding) (None, 1, 3) 21 input_dofr[0][0] __________________________________________________________________________________________________ input_number_features (InputLay [(None, 4)] 0 __________________________________________________________________________________________________ input_binary_features (InputLay [(None, 12)] 0 __________________________________________________________________________________________________ reshape_8 (Reshape) (None, 70) 0 embedding_8[0][0] __________________________________________________________________________________________________ reshape_9 (Reshape) (None, 6) 0 embedding_9[0][0] __________________________________________________________________________________________________ reshape_10 (Reshape) (None, 77) 0 embedding_10[0][0] __________________________________________________________________________________________________ reshape_11 (Reshape) (None, 5) 0 embedding_11[0][0] __________________________________________________________________________________________________ reshape_12 (Reshape) (None, 3) 0 embedding_12[0][0] __________________________________________________________________________________________________ reshape_13 (Reshape) (None, 10) 0 embedding_13[0][0] __________________________________________________________________________________________________ reshape_14 (Reshape) (None, 2) 0 embedding_14[0][0] __________________________________________________________________________________________________ reshape_15 (Reshape) (None, 3) 0 embedding_15[0][0] __________________________________________________________________________________________________ concatenate_1 (Concatenate) (None, 192) 0 input_number_features[0][0] input_binary_features[0][0] reshape_8[0][0] reshape_9[0][0] reshape_10[0][0] reshape_11[0][0] reshape_12[0][0] reshape_13[0][0] reshape_14[0][0] reshape_15[0][0] __________________________________________________________________________________________________ dense_8 (Dense) (None, 16) 3088 concatenate_1[0][0] __________________________________________________________________________________________________ dense_9 (Dense) (None, 8) 136 dense_8[0][0] __________________________________________________________________________________________________ dense_10 (Dense) (None, 4) 36 dense_9[0][0] __________________________________________________________________________________________________ dense_11 (Dense) (None, 2) 10 dense_10[0][0] __________________________________________________________________________________________________ dense_12 (Dense) (None, 4) 12 dense_11[0][0] __________________________________________________________________________________________________ dense_13 (Dense) (None, 8) 40 dense_12[0][0] __________________________________________________________________________________________________ dense_14 (Dense) (None, 16) 144 dense_13[0][0] __________________________________________________________________________________________________ dense_15 (Dense) (None, 192) 3264 dense_14[0][0] ================================================================================================== Total params: 793,320 Trainable params: 793,320 Non-trainable params: 0
我的问题是,如何将pandas数据框中的变量映射到模型对应的输入层?例如将每个高基数类别因子传入正确的嵌入层。我初步的想法是在训练前按输入层要求的顺序调整数据集,或是建立变量到输入层的映射关系(例如provID -> input_provid)。
解决方案
优先选择建立特征到输入层的映射关系,比手动调整数据集列顺序的容错率高得多,Keras的多输入模型本身就支持传入字典匹配输入层,操作很方便:
- 预处理阶段先对所有高基数类别特征做LabelEncoder编码,把原始的字符串/离散值转成从0开始的连续整数,每个特征单独训练一个编码器,推理时复用训练好的编码器,避免出现编码不一致的问题。
- 定义特征与输入层的映射规则:
- 8个需要走嵌入层的单值类别特征,逐一对应到同名的输入层
- 4个连续特征统一整理成二维数组,对应
input_number_features输入层 - 12个二值特征统一整理成二维数组,对应
input_binary_features输入层
- 构造输入字典,代码示例如下:
import numpy as np import pandas as pd from sklearn.preprocessing import LabelEncoder # 1. 定义映射关系,替换成你自己的特征名 cat_feat_map = { "provID": "input_provid", "pos_code": "input_pos_code", "prindiag": "input_prindiag", "billtype2": "input_billtype2", "lob": "input_lob", "ppg_code": "input_ppg_code", "segment": "input_segment", "dofr": "input_dofr" } num_feats = ["连续特征1", "连续特征2", "连续特征3", "连续特征4"] binary_feats = ["二值特征1", "二值特征2", ..., "二值特征12"] # 2. 编码类别特征(这里仅做示例,实际要把编码器存下来用于推理) encoders = {} for feat in cat_feat_map.keys(): le = LabelEncoder() df[feat] = le.fit_transform(df[feat]) encoders[feat] = le # 3. 构造模型输入字典 model_input = {} # 处理单个类别特征,扩充一维匹配输入层(None,1)的形状要求 for feat_name, input_name in cat_feat_map.items(): model_input[input_name] = np.expand_dims(df[feat_name].values, axis=-1) # 处理连续和二值特征 model_input["input_number_features"] = df[num_feats].values model_input["input_binary_features"] = df[binary_feats].values
- 训练和推理直接用上述字典即可:
# 训练示例,自编码器的输入和输出一致 autoencoder.fit(model_input, model_input, epochs=50, batch_size=256, validation_split=0.2)
这种方式的优势很明显:不需要调整数据集列顺序,后续新增/删除特征只要修改映射配置即可,不会出现特征和输入层错位的问题,代码可读性和可维护性都更高。
注意:如果有未在训练集出现过的未知类别,建议在编码时预留一个专门的索引对应未知类别,避免推理时报错。
内容的提问来源于stack exchange,提问作者Stu
相关产品推荐
相关产品推荐

