使用Pandas从字符串变量生成哑变量时出现重复变量的问题
重复哑变量的原因与解决办法
问题原因
你遇到的重复哑变量问题,核心是类别字符串存在隐形空格差异:
替换括号后,crops列的内容是类似maize, cassava的格式(逗号后带空格),直接用,作为分隔符调用str.get_dummies()时,会把 maize(前面带空格)和maize判定为两个不同类别,最终生成重复的哑变量列(比如 maize和maize)。
解决办法
有两种简单的修复方式:
方式一:提前清理空格
在替换括号后,统一把逗号+空格的组合替换成纯逗号,确保所有类别字符串无前置空格:
import pandas as pd data = { 'id': [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20], 'crops': ['[maize]', '[maize, cassava]', '[beans, cassava, potato]', '[beans, potato]', '[beans, cassava, maize, potato]', '[beans]', '[cassava, maize, potato]', '[beans, maize]', '[cassava, maize, potato]', '[cassava]', '[beans, cassava, potato]', '[maize, potato]', '[beans, maize, potato]', '[beans, cassava, maize, potato]', '[potato]', '[cassava, potato]', '[beans]', '[maize]', '[potato]', '[cassava]'], } df = pd.DataFrame(data) # 替换括号,同时清理逗号后的空格 df['crops'] = df['crops'].str.replace(r'[\[\]]', '', regex=True) df['crops'] = df['crops'].str.replace(', ', ',', regex=False) res = df.join(df.pop('crops').str.get_dummies(',')) res
方式二:用正则作为分隔符
直接在get_dummies()中使用正则表达式',\s*'作为分隔符,匹配逗号及后面任意数量的空格,避免空格干扰:
import pandas as pd data = { 'id': [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20], 'crops': ['[maize]', '[maize, cassava]', '[beans, cassava, potato]', '[beans, potato]', '[beans, cassava, maize, potato]', '[beans]', '[cassava, maize, potato]', '[beans, maize]', '[cassava, maize, potato]', '[cassava]', '[beans, cassava, potato]', '[maize, potato]', '[beans, maize, potato]', '[beans, cassava, maize, potato]', '[potato]', '[cassava, potato]', '[beans]', '[maize]', '[potato]', '[cassava]'], } df = pd.DataFrame(data) df['crops'] = df['crops'].str.replace('[', '', regex=False) df['crops'] = df['crops'].str.replace(']', '', regex=False) # 用正则匹配逗号+任意空格作为分隔符 res = df.join(df.pop('crops').str.get_dummies(',\s*')) res
两种方式都能确保类别字符串统一,不会生成重复的哑变量列。
内容的提问来源于stack exchange,提问作者Stephen Okiya
相关产品推荐
相关产品推荐

