如何对数据集列做One Hot Encoding?pd.get_dummies报错求助
问题描述
原始数据集:
bread milk butter jam nutella cheese chips 0 bread NaN butter jam nutella NaN NaN 1 NaN NaN butter jam nutella NaN chips 2 NaN milk NaN NaN NaN cheese NaN 3 bread milk butter jam nutella cheese chips 4 bread milk NaN NaN nutella NaN NaN 5 bread milk butter jam NaN cheese chips 6 bread milk NaN NaN nutella NaN NaN 7 bread NaN butter NaN NaN cheese NaN 8 bread NaN butter jam nutella NaN NaN 9 NaN milk butter jam NaN cheese NaN 10 bread NaN NaN jam nutella cheese chips 11 bread milk butter jam nutella NaN NaN 12 bread NaN butter NaN nutella cheese NaN 13 bread NaN butter jam nutella cheese chips 14 bread milk butter jam nutella cheese chips 15 NaN milk butter jam nutella cheese NaN 16 NaN milk NaN jam nutella cheese NaN 17 bread milk butter jam nutella cheese chips 18 bread NaN butter jam nutella cheese NaN 19 bread milk butter NaN nutella cheese NaN 20 NaN milk NaN NaN NaN NaN chips
需求:对每列进行One Hot Encoding,生成如下格式的全量0/1数据集:
| bread | milk | butter | jam | nutella | cheese | chips |
|---|---|---|---|---|---|---|
| 1 | 0 | 1 | 1 | 1 | 0 | 0 |
| 0 | 0 | 1 | 1 | 1 | 0 | 1 |
尝试代码:
pd.get_dummies(book_data, columns = ['bread', 'milk','butter','jam', 'nutella','cheese','chips'])
出现错误:
KeyError: "['bread', 'cheese'] not in index"
解决方案
你用错了方法——pd.get_dummies是给分类变量生成新的独热编码列的工具,但你的需求是直接把每一列转换成「是否存在该商品」的0/1值,而非对分类值编码。另外报错的核心原因是原始数据的列名可能带有前后空格,导致代码找不到对应列。
正确实现代码:
import pandas as pd # 假设原始数据已读取为DataFrame对象df # 第一步:去除列名前后的空格,解决KeyError问题 df.columns = df.columns.str.strip() # 第二步:将每列转换为0/1:非空值(存在商品)为1,空值(NaN)为0 one_hot_df = df.notna().astype(int) # 查看结果 print(one_hot_df)
这段代码完成两个关键操作:
- 清理列名,消除因空格导致的列名匹配失败
- 通过
notna()判断单元格是否存在商品,将布尔结果转为整数,直接得到目标格式的0/1数据集
内容的提问来源于stack exchange,提问作者Addy
相关产品推荐
相关产品推荐

