如何在保留时序结构的前提下从张量时间序列数据集提取关键特征?
针对时间序列张量的特征选择方案
你的数据是**(样本数, 序列长度, 特征数)**的三维张量,直接用Sklearn的SelectKBest、RFE这类方法需要拍平数据,会丢失时序结构,下面是几种不用破坏时序结构就能筛选特征的方案:
1. 基于时序统计量的映射筛选
先对每个特征的时间序列提取统计量(均值、方差、极值、时序差分统计等),把每个原始特征转换成一组统计特征,再用传统特征选择工具筛选,最后映射回原始特征维度。这种方法既保留了特征的时序信息,又能复用Sklearn的成熟工具。
示例代码:
import numpy as np from sklearn.feature_selection import SelectKBest, f_classif # 将列表转成三维张量 X = np.array(X_list) # shape: (100, 10, 5) y = np.array(y) # 提取每个特征的时序统计量:均值+标准差 X_stats = np.hstack([ X.mean(axis=1), # 每个特征的时序均值,shape(100,5) X.std(axis=1) # 每个特征的时序标准差,shape(100,5) ]) # 最终shape(100,10),每两列对应一个原始特征的统计值 # 用SelectKBest筛选和标签相关性最高的统计特征 selector = SelectKBest(f_classif, k=4) selector.fit(X_stats, y) # 映射回原始特征 selected_feats = set() for idx in np.where(selector.get_support())[0]: original_feat_idx = idx // 2 # 每两个统计特征对应一个原始特征 selected_feats.add(original_feat_idx) print("筛选出的原始特征索引:", sorted(selected_feats))
2. 基于深度学习注意力机制的自动筛选
用带注意力层的时序模型(比如LSTM+注意力),让模型自动学习每个特征在时序中的重要性,通过注意力权重的均值来判断特征的全局相关性,直接处理三维张量,完全保留时序结构。
示例代码:
import tensorflow as tf from tensorflow.keras.layers import Input, LSTM, Dense, Layer from tensorflow.keras.models import Model # 自定义特征注意力层,输出特征重要性 class FeatureAttention(Layer): def __init__(self): super().__init__() self.dense = Dense(1, activation='tanh') def call(self, inputs): # inputs shape: (batch_size, seq_len, n_features) attn_weights = self.dense(inputs) # 生成每个时序步的特征权重 attn_weights = tf.nn.softmax(attn_weights, axis=1) # 时序维度归一化 # 对时序维度取平均,得到每个特征的全局权重 global_feat_weights = tf.reduce_mean(attn_weights, axis=0) self.add_weight(name='feat_importance', shape=(n_features,), trainable=False) self.set_weights([global_feat_weights]) return inputs * attn_weights # 加权后的时序特征 # 构建模型 inputs = Input(shape=(n_sequences, n_features)) attended_feats = FeatureAttention()(inputs) lstm_out = LSTM(32)(attended_feats) outputs = Dense(1, activation='sigmoid')(lstm_out) model = Model(inputs=inputs, outputs=outputs) model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy']) # 训练模型 model.fit(X, y, epochs=10, batch_size=8) # 获取特征重要性并筛选top3 feat_importance = model.get_layer('feature_attention').get_weights()[0].flatten() top_feats = np.argsort(feat_importance)[::-1][:3] print("筛选出的原始特征索引:", top_feats)
3. 时序专用库的特征选择工具
用专门处理时间序列的库,比如sktime或tsfresh,它们支持直接处理三维时序数据,内置特征选择逻辑:
- sktime:可以先将时序数据转换为表格化特征,再结合分类器的特征重要性筛选原始特征
- tsfresh:自动提取大量时序特征,同时筛选出和标签相关的特征,可映射回原始特征维度
sktime示例代码:
from sktime.transformations.panel.reduce import Tabularizer from sktime.feature_selection import FeatureSelector from sklearn.ensemble import RandomForestClassifier # 将时序张量转换为表格化特征(每个时序步作为一列) tabularizer = Tabularizer() X_tab = tabularizer.fit_transform(X) # 用随机森林的特征重要性筛选特征 selector = FeatureSelector( estimator=RandomForestClassifier(), transform="feature_importance", k=5 ) selector.fit(X_tab, y) # 分组计算原始特征的平均重要性 feat_importance = selector.estimator_.feature_importances_ grouped_importance = [np.mean(feat_importance[i*n_sequences : (i+1)*n_sequences]) for i in range(n_features)] # 筛选top3特征 top_feats = np.argsort(grouped_importance)[::-1][:3] print("筛选出的原始特征索引:", top_feats)
内容的提问来源于stack exchange,提问作者Pitone
相关产品推荐
相关产品推荐

