You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Sklearn FeatureHasher编码后模型无法预测新数据问题排查

问题排查与修复方案

核心问题分析

训练阶段的致命错误

  1. 每个IP单独初始化并拟合FeatureHasher
    遍历每个src_ip时都新建FeatureHasher实例并调用fit_transform,会导致相同特征值在不同IP的哈希映射完全不一致。比如IP1的org=Hetzner Online GmbH和IP2的同一个org会被哈希到不同特征位置,模型根本学不到稳定的特征映射,训练集准确率高只是巧合(比如数据按IP分组,测试集和训练集的特征哈希映射刚好重叠)。

  2. 特征维度逻辑混乱
    每个列用n_features=20的哈希器,9个列拼接后总维度应为9*20=180,但每个IP单独拟合哈希器的操作,让相同特征的哈希结果完全不兼容,模型学到的是无意义的特征分布。

预测阶段的错误

  1. 哈希器未初始化
    代码注释掉了hasher = FeatureHasher(...),直接调用hasher.transform会抛出未定义错误,根本无法运行。
  2. 哈希器参数不匹配
    即使初始化,注释里的n_features=11和训练时的n_features=20不一致,特征维度完全不匹配,模型无法接收输入。
  3. 特征输入格式错误
    hasher.transform([feature_value])中,feature_value已经是[str(new_data[column])],再套一层列表会变成二维列表,不符合input_type='string'的要求(应传入字符串的一维可迭代对象)。

修复后的训练代码

import numpy as np
from sklearn.feature_extraction import FeatureHasher
from sklearn.model_selection import train_test_split

columns_to_hash = ['port_cluster', 'org', 'asn', 'protocol', 'event_type', 'os', 'country_name', 'city_name', 'class']

attribute_weights = {
    'port_cluster': 0.7,
    "protocol": 0.4,
    "city_name": 0.4,
    'org': 0.5,
    'asn': 0.5,
    'event_type': 0.6,
    'os': 0.3,
    'country_name': 0.4,
    'class': 0.3,
}

# 为每个特征列创建固定的FeatureHasher,全局拟合一次
column_hashers = {}
for col in columns_to_hash:
    hasher = FeatureHasher(n_features=20, input_type='string')
    # 用整个数据集的该列拟合哈希器
    hasher.fit(df[col].astype(str).values.reshape(-1, 1))
    column_hashers[col] = hasher

hashed_feature_list = []
# 遍历所有行处理特征(无需按IP分组,哈希器已全局拟合)
for idx, row in df.iterrows():
    row_hashed = []
    for col in columns_to_hash:
        # 转换为字符串并处理单个值
        feature_val = str(row[col])
        # 用预拟合的哈希器转换
        hashed = column_hashers[col].transform([[feature_val]]).toarray()
        # 应用权重
        weighted_hashed = hashed * attribute_weights[col]
        row_hashed.append(weighted_hashed)
    # 拼接当前行的所有特征
    row_features = np.concatenate(row_hashed, axis=1)
    hashed_feature_list.append(row_features)

# 合并所有特征
all_hashed_features = np.concatenate(hashed_feature_list, axis=0)

X = all_hashed_features
y = df['cluster_label']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# 后续训练模型的代码(例如 model.fit(X_train, y_train))

修复后的预测代码

# 注意:需保留训练阶段的column_hashers和attribute_weights,或保存后加载
# 此处假设已从训练环境中获取/加载这两个变量

new_data = {
    'src_ip': '65.21.234.90',
    'asn': 24940,
    'country_name': 'Finland',
    'city_name': 'Helsinki',
    'open_ports': ['5060/sip', '2000/ikettle'],
    'protocol': 'UDP',
    'ip_rep': None,
    'first_seen': '2023-08-22T16:56:45.733Z',
    'last_time': '2023-08-22T16:56:45.733Z',
    'class': 'A',
    'event_type': 'sip',
    'event_data': {
        'request_line': 'OPTIONS sip:100@165.22.44.143 SIP/2.0',
        'uri': 'sip:100@165.22.44.143',
        'version': 'SIP/2.0',
        'method': 'OPTIONS'
    },
    'link': None,
    'os': 'None',
    'org': 'Hetzner Online GmbH',
    'port_cluster': -1
}

new_data_hashed = []
for col in columns_to_hash:
    if col in new_data:
        feature_val = str(new_data[col])
        # 使用训练时预拟合的哈希器转换
        hashed = column_hashers[col].transform([[feature_val]]).toarray()
        weighted_hashed = hashed * attribute_weights[col]
        new_data_hashed.append(weighted_hashed)

new_data_features = np.concatenate(new_data_hashed, axis=1)

# 预测
predicted_label = model.predict(new_data_features)
print(f"Predicted Cluster Label: {predicted_label[0]}")

关键优化点说明

  1. 全局拟合哈希器:每个特征列使用同一个预拟合的FeatureHasher,确保训练和预测时相同特征值的哈希映射完全一致。
  2. 统一特征维度:每个列固定20维,9个列拼接后总维度为180,训练和预测的输入维度完全匹配。
  3. 修正输入格式:单个特征值用[[feature_val]]传入transform,符合FeatureHasher对单个样本的输入要求。
  4. 取消按IP分组处理:哈希器已经全局拟合,无需按IP单独处理,保证所有样本的特征映射统一。

内容的提问来源于stack exchange,提问作者Aditya Bhat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 06:23:14