多元线性回归非数值因变量编码报错:KeyError找不到指定列
问题描述
我正在开展多元线性回归任务,用3个自变量(2个非数值型、1个数值型)预测8个因变量(5个数值型、3个非数值型),需要编码处理非数值数据。但代码执行时出现KeyError,提示找不到'Wall Insulation'等3个非数值因变量列,已确认CSV文件包含对应列名,无拼写、空格或大小写问题。
报错信息
Traceback (most recent call last): File "D:\EMJMD-SMACCs\Thesis\Thesis Work\Recommendation Engine\Code\venv\lib\site-packages\pandas\core\indexes\base.py", line 3802, in get_loc return self._engine.get_loc(casted_key) File "pandas_libs\index.pyx", line 138, in pandas._libs.index.IndexEngine.get_loc File "pandas_libs\index.pyx", line 165, in pandas._libs.index.IndexEngine.get_loc File "pandas_libs\hashtable_class_helper.pxi", line 5745, in pandas._libs.hashtable.PyObjectHashTable.get_item File "pandas_libs\hashtable_class_helper.pxi", line 5753, in pandas._libs.hashtable.PyObjectHashTable.get_item KeyError: 'Wall Insulation' The above exception was the direct cause of the following exception: Traceback (most recent call last): File "D:\EMJMD-SMACCs\Thesis\Thesis Work\Recommendation Engine\Code\exppppp.py", line 20, in data['Wall Insulation Encoded'] = label_enc.fit_transform(data['Wall Insulation']) File "D:\EMJMD-SMACCs\Thesis\Thesis Work\Recommendation Engine\Code\venv\lib\site-packages\pandas\core\frame.py", line 3807, in __getitem__ indexer = self.columns.get_loc(key) File "D:\EMJMD-SMACCs\Thesis\Thesis Work\Recommendation Engine\Code\venv\lib\site-packages\pandas\core\indexes\base.py", line 3804, in get_loc raise KeyError(key) from err KeyError: 'Wall Insulation'
错误原因
- 列被提前删除:你把
Wall Insulation、Roof Insulation、Window Glazing加入了enc_cat列表,执行data.drop(enc_cat, axis=1, inplace=True)时已经把这些因变量列删掉了,后面再调用data['Wall Insulation']自然找不到对应的列。 - LabelEncoder复用问题:用同一个
LabelEncoder实例编码三个不同的分类变量,解码时会出现类别映射混乱的问题,因为每个变量的类别集合是独立的。 - 模型选型不当:用线性回归预测离散的分类变量不合适,线性回归输出连续值,强行转成整数解码会导致逻辑误差。
解决步骤
- 拆分分类列:把自变量的分类列和因变量的分类列分开处理,只删除自变量的原始分类列,保留因变量列直到编码完成。
- 独立使用LabelEncoder:为每个因变量单独创建
LabelEncoder实例,避免类别映射交叉污染。 - 调整数据处理顺序:先完成因变量的编码,再清理不需要的原始列(仅清理自变量的分类列)。
修正后的代码
import pandas as pd from sklearn.linear_model import LinearRegression from sklearn.preprocessing import OneHotEncoder, LabelEncoder # 加载数据,建议用原始字符串避免转义问题 data = pd.read_csv(r'D:\EMJMD-SMACCs\Thesis\Thesis Work\Recommendation Engine\Experiment.csv') # 拆分自变量分类列和因变量分类列 cat_features = ['Building type', 'Building climate'] # 自变量里的非数值型列 cat_targets = ['Wall Insulation', 'Roof Insulation', 'Window Glazing'] # 因变量里的非数值型列 # 编码自变量的分类列,用get_feature_names_out自动生成规范列名 enc = OneHotEncoder(sparse_output=False) enc_df = pd.DataFrame(enc.fit_transform(data[cat_features]), columns=enc.get_feature_names_out(cat_features)) data = pd.concat([data, enc_df], axis=1) # 只删除自变量的原始分类列 data.drop(cat_features, axis=1, inplace=True) # 为每个因变量单独创建LabelEncoder # 编码Wall Insulation le_wall = LabelEncoder() data['Wall Insulation Encoded'] = le_wall.fit_transform(data['Wall Insulation']) # 编码Roof Insulation le_roof = LabelEncoder() data['Roof Insulation Encoded'] = le_roof.fit_transform(data['Roof Insulation']) # 编码Window Glazing le_window = LabelEncoder() data['Window Glazing Encoded'] = le_window.fit_transform(data['Window Glazing']) # 准备训练集 X = data[['Building area'] + list(enc_df.columns)] y = data[['Wall U value', 'Roof U value', 'Wall Insulation Encoded', 'Wall Insulation thickness', 'Roof Insulation Encoded', 'Roof insulation thickness', 'Window U value', 'Window Glazing Encoded']] model = LinearRegression().fit(X, y) # 用户输入处理 building_type = input("Enter the building type (e.g. Single family house): ") building_climate = input("Enter the building climate (e.g. Mediterranean): ") building_area = float(input("Enter the building area (in square meters): ")) # 编码用户输入的自变量 enc_input = enc.transform([[building_type, building_climate]]).toarray() enc_input_df = pd.DataFrame(enc_input, columns=enc.get_feature_names_out(cat_features)) # 拼接输入并预测 user_input = pd.concat([pd.DataFrame({'Building area': [building_area]}), enc_input_df], axis=1) predicted_enc = model.predict(user_input) # 解码分类变量(注意:线性回归输出连续值,取整为临时方案,建议后续改用分类模型) predicted_wall_insulation = le_wall.inverse_transform([round(predicted_enc[0][2])])[0] predicted_roof_insulation = le_roof.inverse_transform([round(predicted_enc[0][4])])[0] predicted_window_glazing = le_window.inverse_transform([round(predicted_enc[0][7])])[0] # 打印结果 print("Predicted values for the given input:") print(f"Wall U value: {predicted_enc[0][0]:.3f}") print(f"Roof U value: {predicted_enc[0][1]:.3f}") print(f"Wall Insulation thickness: {predicted_enc[0][3]:.3f}") print(f"Roof insulation thickness: {predicted_enc[0][5]:.3f}") print(f"Window U value: {predicted_enc[0][6]:.3f}") print(f"Wall Insulation: {predicted_wall_insulation}") print(f"Roof Insulation: {predicted_roof_insulation}") print(f"Window Glazing: {predicted_window_glazing}")
额外注意
- 模型适配问题:线性回归不适合预测分类变量,建议针对3个非数值因变量单独训练分类模型(如逻辑回归、随机森林分类器),或者使用支持多任务学习的框架拆分回归和分类任务。
- 路径处理:Windows系统下的文件路径建议用原始字符串(前缀加
r),避免转义字符(如\T、\R)被解析成特殊字符。 - 编码可靠性:
OneHotEncoder的get_feature_names_out()方法能自动生成带前缀的列名,比手动拼接更不易出错。
内容的提问来源于stack exchange,提问作者Talha Siddiqui
相关产品推荐
相关产品推荐

