You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用pd.get_dummies()生成包含全词汇表的独热编码?

如何用Pandas生成包含指定完整词汇表的独热编码?

你遇到的问题是pd.get_dummies默认只会对输入数据中实际存在的类别生成编码列,而不会自动包含你指定的vocab里的所有类别。另外,你传入的columns参数用法有误——这个参数是针对DataFrame的列名,而非用来指定词汇表的。

你的原始代码与问题复现

原始数据定义:

vocab = ['a', 'b', 'c', 'd', 'e', 'f', 'g', 'h', 'i', 'j']
list1 = ['a', 'b', 'c', 'd', 'e']
list2 = ['f', 'g', 'h', 'i', 'j']

你尝试的编码代码:

import pandas as pd
encoding1 = pd.get_dummies(data= list1, columns= vocab)
encoding2 = pd.get_dummies(data= list2, columns= vocab)

期望输出:

encoding1 =
   a  b  c  d  e  f  g  h  i  j
0  1  0  0  0  0  0  0  0  0  0
1  0  1  0  0  0  0  0  0  0  0
2  0  0  1  0  0  0  0  0  0  0
3  0  0  0  1  0  0  0  0  0  0
4  0  0  0  0  1  0  0  0  0  0

encoding2 =
   a  b  c  d  e  f  g  h  i  j
0  0  0  0  0  1  0  0  0  0
1  0  0  0  0  0  1  0  0  0
2  0  0  0  0  0  0  1  0  0
3  0  0  0  0  0  0  0  1  0
4  0  0  0  0  0  0  0  0  1

实际输出:

encoding1 =
   a  b  c  d  e
0  1  0  0  0  0
1  0  1  0  0  0
2  0  0  1  0  0
3  0  0  0  1  0
4  0  0  0  0  1

encoding2 =
   f  g  h  i  j
0  1  0  0  0  0
1  0  1  0  0  0
2  0  0  1  0  0
3  0  0  0  1  0
4  0  0  0  0  1

解决方案:用pd.Categorical指定完整词汇表

要让get_dummies生成包含vocab所有类别的编码,关键是先将输入列表转换为分类类型(Categorical),并明确指定categories为你的完整词汇表。这样即使输入数据中没有某些类别,编码结果也会保留这些列并填充0。

完整代码示例

import pandas as pd

vocab = ['a', 'b', 'c', 'd', 'e', 'f', 'g', 'h', 'i', 'j']
list1 = ['a', 'b', 'c', 'd', 'e']
list2 = ['f', 'g', 'h', 'i', 'j']

# 将列表转换为Categorical类型,指定完整词汇表作为所有可能类别
cat_list1 = pd.Categorical(list1, categories=vocab)
encoding1 = pd.get_dummies(cat_list1)

cat_list2 = pd.Categorical(list2, categories=vocab)
encoding2 = pd.get_dummies(cat_list2)

# 打印结果
print("encoding1 =")
print(encoding1)
print("\nencoding2 =")
print(encoding2)

输出结果

运行后会得到你期望的包含所有vocab列的独热编码,不存在的类别列值为0,完全符合你的需求。

补充说明

  • pd.Categorical的作用是定义数据的所有可能类别,强制约束数据的分类范围,这是让get_dummies生成完整编码的核心。
  • 如果你的数据是DataFrame格式,也可以对指定列做同样的转换:
    df1 = pd.DataFrame({'value': list1})
    df1['value'] = pd.Categorical(df1['value'], categories=vocab)
    encoding1 = pd.get_dummies(df1['value'])
    

内容的提问来源于stack exchange,提问作者ritsj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:39:40