如何用Python的pd.get_dummies()基于字符串关键词创建哑变量
问题
我有如下格式的DataFrame:
PROMOTED_PRODUCT__CREATIVE ROAS 0 Simple Green 1 Gal. Concentrated 0.027573 1 Simple Green 1 Gal. Concentrated 0.082969 2 Simple Green 1 Gal. Concentrated 0.056278 3 Simple Green 32 oz Concentrated 0.037286 4 Simple Green 32 oz Concentrated 0.355841 5 Simple Green 32 oz Concentrated 0.355853 6 Simple Green 16 oz Concentrated 0.355923 7 Simple Green 16 oz Concentrated 0.355749 8 Simple Green 16 oz Concentrated 0.355810
我希望基于PROMOTED_PRODUCT__CREATIVE列字符串中的属性(如'1 gal'、'32 oz'、'16 oz'等)创建如下格式的哑变量:
1_gal 32_oz 16_oz 0 1 0 0 1 1 0 0 2 1 0 0 3 0 1 0 4 0 1 0 5 0 1 0 ...
请问能否通过pd.get_dummies()快速实现该需求?恳请各位提供帮助,非常感谢!
解决方案
可以通过pd.get_dummies()实现,核心是先从原列提取目标属性,再生成哑变量,步骤如下:
提取容量属性
用正则表达式匹配字符串中的数字+单位部分,统一格式(空格换下划线、移除点号):import pandas as pd import re # 假设原数据集为df df['size'] = df['PROMOTED_PRODUCT__CREATIVE'].apply( lambda x: re.search(r'\d+ (Gal\.|oz)', x).group().replace(' ', '_').replace('.', '') )执行后
size列会得到1_Gal、32_oz、16_oz这类标准化值。生成哑变量
直接对size列调用pd.get_dummies():dummy_vars = pd.get_dummies(df['size'])输出结果就是你需要的哑变量矩阵。
合并到原数据集(可选)
如果需要把哑变量和原数据整合,用pd.concat():df_final = pd.concat([df, dummy_vars], axis=1)
内容的提问来源于stack exchange,提问作者mexicanRmy
相关产品推荐
相关产品推荐

