将分类特征转换为数值时触发ValueError: A given column is not a column of the dataframe的问题排查
解决ColumnTransformer使用时的ValueError: A given column is not a column of the dataframe问题
嘿,这个问题我太熟了!你在设置ColumnTransformer的时候犯了个很容易踩的小坑:
你写了这段代码:
cat = X["Country"] transformer = ct([("one_hot", one_hot, cat)],remainder="passthrough")
这里的cat变量存的是X["Country"]返回的整列数据内容(比如第一个值是"Tunisia"),但ColumnTransformer要求的第三个参数是列的标识信息——要么是列名(字符串,比如"Country"),要么是列的索引位置(整数,比如这里是1,因为X里第一列是Age,第二列是Country)。
当你把整列数据传进去时,scikit-learn会错误地把数据里的第一个值"Tunisia"当成列名去DataFrame里查找,而你的DataFrame根本没有叫这个名字的列,所以就抛出了KeyError: 'Tunisia',进而引发了后续的ValueError。
修正后的完整代码
只需要把列的指定方式改成列名(或索引)就行,这里用列名更直观:
import numpy as np import pandas as pd import matplotlib.pyplot as plt from sklearn.compose import ColumnTransformer as ct from sklearn.model_selection import train_test_split as tts from sklearn.preprocessing import OneHotEncoder data = pd.read_csv("advertising.csv") X = data[["Age","Country"]] y = data[["Clicked on Ad"]] one_hot = OneHotEncoder() # 直接传入列名列表,而不是列数据 transformer = ct([("one_hot", one_hot, ["Country"])], remainder="passthrough") transformed_X = transformer.fit_transform(X) print(transformed_X)
这里用列表["Country"]是更规范的写法,如果你需要对多列做独热编码,直接在列表里加列名就行;当然也可以传单个字符串"Country",效果是一样的。
另一种可行写法(用列索引)
如果你的DataFrame列顺序固定不变,也可以用列的索引位置来指定:
# Country是X的第二列,索引从0开始所以是1 transformer = ct([("one_hot", one_hot, [1])], remainder="passthrough")
这种写法在列名很长或者有重复列名的时候会更方便。
内容的提问来源于stack exchange,提问作者cagatay.e.sahin
相关产品推荐
相关产品推荐

