使用Sklearn FeatureHasher编码后模型无法预测新数据问题排查
问题排查与修复方案
核心问题分析
训练阶段的致命错误
每个IP单独初始化并拟合FeatureHasher
遍历每个src_ip时都新建FeatureHasher实例并调用fit_transform,会导致相同特征值在不同IP的哈希映射完全不一致。比如IP1的org=Hetzner Online GmbH和IP2的同一个org会被哈希到不同特征位置,模型根本学不到稳定的特征映射,训练集准确率高只是巧合(比如数据按IP分组,测试集和训练集的特征哈希映射刚好重叠)。特征维度逻辑混乱
每个列用n_features=20的哈希器,9个列拼接后总维度应为9*20=180,但每个IP单独拟合哈希器的操作,让相同特征的哈希结果完全不兼容,模型学到的是无意义的特征分布。
预测阶段的错误
- 哈希器未初始化
代码注释掉了hasher = FeatureHasher(...),直接调用hasher.transform会抛出未定义错误,根本无法运行。 - 哈希器参数不匹配
即使初始化,注释里的n_features=11和训练时的n_features=20不一致,特征维度完全不匹配,模型无法接收输入。 - 特征输入格式错误
hasher.transform([feature_value])中,feature_value已经是[str(new_data[column])],再套一层列表会变成二维列表,不符合input_type='string'的要求(应传入字符串的一维可迭代对象)。
修复后的训练代码
import numpy as np from sklearn.feature_extraction import FeatureHasher from sklearn.model_selection import train_test_split columns_to_hash = ['port_cluster', 'org', 'asn', 'protocol', 'event_type', 'os', 'country_name', 'city_name', 'class'] attribute_weights = { 'port_cluster': 0.7, "protocol": 0.4, "city_name": 0.4, 'org': 0.5, 'asn': 0.5, 'event_type': 0.6, 'os': 0.3, 'country_name': 0.4, 'class': 0.3, } # 为每个特征列创建固定的FeatureHasher,全局拟合一次 column_hashers = {} for col in columns_to_hash: hasher = FeatureHasher(n_features=20, input_type='string') # 用整个数据集的该列拟合哈希器 hasher.fit(df[col].astype(str).values.reshape(-1, 1)) column_hashers[col] = hasher hashed_feature_list = [] # 遍历所有行处理特征(无需按IP分组,哈希器已全局拟合) for idx, row in df.iterrows(): row_hashed = [] for col in columns_to_hash: # 转换为字符串并处理单个值 feature_val = str(row[col]) # 用预拟合的哈希器转换 hashed = column_hashers[col].transform([[feature_val]]).toarray() # 应用权重 weighted_hashed = hashed * attribute_weights[col] row_hashed.append(weighted_hashed) # 拼接当前行的所有特征 row_features = np.concatenate(row_hashed, axis=1) hashed_feature_list.append(row_features) # 合并所有特征 all_hashed_features = np.concatenate(hashed_feature_list, axis=0) X = all_hashed_features y = df['cluster_label'] X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42) # 后续训练模型的代码(例如 model.fit(X_train, y_train))
修复后的预测代码
# 注意:需保留训练阶段的column_hashers和attribute_weights,或保存后加载 # 此处假设已从训练环境中获取/加载这两个变量 new_data = { 'src_ip': '65.21.234.90', 'asn': 24940, 'country_name': 'Finland', 'city_name': 'Helsinki', 'open_ports': ['5060/sip', '2000/ikettle'], 'protocol': 'UDP', 'ip_rep': None, 'first_seen': '2023-08-22T16:56:45.733Z', 'last_time': '2023-08-22T16:56:45.733Z', 'class': 'A', 'event_type': 'sip', 'event_data': { 'request_line': 'OPTIONS sip:100@165.22.44.143 SIP/2.0', 'uri': 'sip:100@165.22.44.143', 'version': 'SIP/2.0', 'method': 'OPTIONS' }, 'link': None, 'os': 'None', 'org': 'Hetzner Online GmbH', 'port_cluster': -1 } new_data_hashed = [] for col in columns_to_hash: if col in new_data: feature_val = str(new_data[col]) # 使用训练时预拟合的哈希器转换 hashed = column_hashers[col].transform([[feature_val]]).toarray() weighted_hashed = hashed * attribute_weights[col] new_data_hashed.append(weighted_hashed) new_data_features = np.concatenate(new_data_hashed, axis=1) # 预测 predicted_label = model.predict(new_data_features) print(f"Predicted Cluster Label: {predicted_label[0]}")
关键优化点说明
- 全局拟合哈希器:每个特征列使用同一个预拟合的FeatureHasher,确保训练和预测时相同特征值的哈希映射完全一致。
- 统一特征维度:每个列固定20维,9个列拼接后总维度为180,训练和预测的输入维度完全匹配。
- 修正输入格式:单个特征值用
[[feature_val]]传入transform,符合FeatureHasher对单个样本的输入要求。 - 取消按IP分组处理:哈希器已经全局拟合,无需按IP单独处理,保证所有样本的特征映射统一。
内容的提问来源于stack exchange,提问作者Aditya Bhat
相关产品推荐
相关产品推荐

