Label Encoding编码索引未按原标签频次排序的原因与解决方法
问题详情
原始标签分布:
df['Label'].value_counts() ------------------------------------ Benign 4401366 DDoS attacks-LOIC-HTTP 576191 FTP-BruteForce 193360 SSH-Bruteforce 187589 DoS attacks-GoldenEye 41508 DoS attacks-Slowloris 10990 Name: Label, dtype: int64
使用Label Encoding编码:
from sklearn.preprocessing import LabelEncoder label_encoder = LabelEncoder() label_encoder.fit(df.Label) df['Label']= label_encoder.transform(df.Label)
实际编码后结果:
df['Label'].value_counts() ------------------------------ 0 4380628 1 576191 4 193354 5 187589 2 41508 3 10990 Name: Label, dtype: int64
期望编码后结果:
df['Label'].value_counts() ------------------------------ 0 4380628 1 576191 2 193354 3 187589 4 41508 5 10990 Name: Label, dtype: int64
问:出现该问题的原因是什么,该如何解决?
原因分析
LabelEncoder默认按标签的字符串字典序分配编码值,而非按标签的出现次数排序。
原始标签的字符串字典序为:Benign → 0DDoS attacks-LOIC-HTTP →1DoS attacks-GoldenEye →2DoS attacks-Slowloris →3FTP-BruteForce →4SSH-Bruteforce →5
这就导致原本计数第三的FTP-BruteForce被编码为4,计数第四的SSH-Bruteforce被编码为5,和你期望的「按计数从大到小分配编码」不符。
解决方法
以下两种方法均可实现你想要的编码顺序:
方法1:手动构建编码映射字典
直接按期望的标签顺序创建映射,用map方法替换原标签:
# 按计数从大到小定义期望的标签顺序 desired_label_order = [ 'Benign', 'DDoS attacks-LOIC-HTTP', 'FTP-BruteForce', 'SSH-Bruteforce', 'DoS attacks-GoldenEye', 'DoS attacks-Slowloris' ] # 构建编码映射字典 label_mapping = {label: idx for idx, label in enumerate(desired_label_order)} # 替换原标签列 df['Label'] = df['Label'].map(label_mapping)
方法2:使用pandas分类类型指定顺序
先将标签列转为有序分类类型,再提取编码:
import pandas as pd desired_label_order = [ 'Benign', 'DDoS attacks-LOIC-HTTP', 'FTP-BruteForce', 'SSH-Bruteforce', 'DoS attacks-GoldenEye', 'DoS attacks-Slowloris' ] # 将标签转为有序分类,指定顺序 df['Label'] = pd.Categorical(df['Label'], categories=desired_label_order, ordered=True) # 提取分类编码 df['Label'] = df['Label'].cat.codes
内容的提问来源于stack exchange,提问作者Dead
相关产品推荐
相关产品推荐

