You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多元线性回归非数值因变量编码报错:KeyError找不到指定列

问题描述

我正在开展多元线性回归任务,用3个自变量(2个非数值型、1个数值型)预测8个因变量(5个数值型、3个非数值型),需要编码处理非数值数据。但代码执行时出现KeyError,提示找不到'Wall Insulation'等3个非数值因变量列,已确认CSV文件包含对应列名,无拼写、空格或大小写问题。

报错信息

Traceback (most recent call last):
File "D:\EMJMD-SMACCs\Thesis\Thesis Work\Recommendation Engine\Code\venv\lib\site-packages\pandas\core\indexes\base.py", line 3802, in get_loc
return self._engine.get_loc(casted_key)
File "pandas_libs\index.pyx", line 138, in pandas._libs.index.IndexEngine.get_loc
File "pandas_libs\index.pyx", line 165, in pandas._libs.index.IndexEngine.get_loc
File "pandas_libs\hashtable_class_helper.pxi", line 5745, in pandas._libs.hashtable.PyObjectHashTable.get_item
File "pandas_libs\hashtable_class_helper.pxi", line 5753, in pandas._libs.hashtable.PyObjectHashTable.get_item
KeyError: 'Wall Insulation'

The above exception was the direct cause of the following exception:

Traceback (most recent call last):
File "D:\EMJMD-SMACCs\Thesis\Thesis Work\Recommendation Engine\Code\exppppp.py", line 20, in 
data['Wall Insulation Encoded'] = label_enc.fit_transform(data['Wall Insulation'])
File "D:\EMJMD-SMACCs\Thesis\Thesis Work\Recommendation Engine\Code\venv\lib\site-packages\pandas\core\frame.py", line 3807, in __getitem__
indexer = self.columns.get_loc(key)
File "D:\EMJMD-SMACCs\Thesis\Thesis Work\Recommendation Engine\Code\venv\lib\site-packages\pandas\core\indexes\base.py", line 3804, in get_loc
raise KeyError(key) from err
KeyError: 'Wall Insulation'
错误原因
  1. 列被提前删除:你把Wall Insulation、Roof Insulation、Window Glazing加入了enc_cat列表,执行data.drop(enc_cat, axis=1, inplace=True)时已经把这些因变量列删掉了,后面再调用data['Wall Insulation']自然找不到对应的列。
  2. LabelEncoder复用问题:用同一个LabelEncoder实例编码三个不同的分类变量,解码时会出现类别映射混乱的问题,因为每个变量的类别集合是独立的。
  3. 模型选型不当:用线性回归预测离散的分类变量不合适,线性回归输出连续值,强行转成整数解码会导致逻辑误差。
解决步骤
  1. 拆分分类列:把自变量的分类列和因变量的分类列分开处理,只删除自变量的原始分类列,保留因变量列直到编码完成。
  2. 独立使用LabelEncoder:为每个因变量单独创建LabelEncoder实例,避免类别映射交叉污染。
  3. 调整数据处理顺序:先完成因变量的编码,再清理不需要的原始列(仅清理自变量的分类列)。
修正后的代码
import pandas as pd
from sklearn.linear_model import LinearRegression
from sklearn.preprocessing import OneHotEncoder, LabelEncoder

# 加载数据,建议用原始字符串避免转义问题
data = pd.read_csv(r'D:\EMJMD-SMACCs\Thesis\Thesis Work\Recommendation Engine\Experiment.csv')

# 拆分自变量分类列和因变量分类列
cat_features = ['Building type', 'Building climate']  # 自变量里的非数值型列
cat_targets = ['Wall Insulation', 'Roof Insulation', 'Window Glazing']  # 因变量里的非数值型列

# 编码自变量的分类列,用get_feature_names_out自动生成规范列名
enc = OneHotEncoder(sparse_output=False)
enc_df = pd.DataFrame(enc.fit_transform(data[cat_features]), columns=enc.get_feature_names_out(cat_features))
data = pd.concat([data, enc_df], axis=1)
# 只删除自变量的原始分类列
data.drop(cat_features, axis=1, inplace=True)

# 为每个因变量单独创建LabelEncoder
# 编码Wall Insulation
le_wall = LabelEncoder()
data['Wall Insulation Encoded'] = le_wall.fit_transform(data['Wall Insulation'])
# 编码Roof Insulation
le_roof = LabelEncoder()
data['Roof Insulation Encoded'] = le_roof.fit_transform(data['Roof Insulation'])
# 编码Window Glazing
le_window = LabelEncoder()
data['Window Glazing Encoded'] = le_window.fit_transform(data['Window Glazing'])

# 准备训练集
X = data[['Building area'] + list(enc_df.columns)]
y = data[['Wall U value', 'Roof U value', 'Wall Insulation Encoded', 
          'Wall Insulation thickness', 'Roof Insulation Encoded', 
          'Roof insulation thickness', 'Window U value', 'Window Glazing Encoded']]
model = LinearRegression().fit(X, y)

# 用户输入处理
building_type = input("Enter the building type (e.g. Single family house): ")
building_climate = input("Enter the building climate (e.g. Mediterranean): ")
building_area = float(input("Enter the building area (in square meters): "))

# 编码用户输入的自变量
enc_input = enc.transform([[building_type, building_climate]]).toarray()
enc_input_df = pd.DataFrame(enc_input, columns=enc.get_feature_names_out(cat_features))

# 拼接输入并预测
user_input = pd.concat([pd.DataFrame({'Building area': [building_area]}), enc_input_df], axis=1)
predicted_enc = model.predict(user_input)

# 解码分类变量(注意:线性回归输出连续值,取整为临时方案,建议后续改用分类模型)
predicted_wall_insulation = le_wall.inverse_transform([round(predicted_enc[0][2])])[0]
predicted_roof_insulation = le_roof.inverse_transform([round(predicted_enc[0][4])])[0]
predicted_window_glazing = le_window.inverse_transform([round(predicted_enc[0][7])])[0]

# 打印结果
print("Predicted values for the given input:")
print(f"Wall U value: {predicted_enc[0][0]:.3f}")
print(f"Roof U value: {predicted_enc[0][1]:.3f}")
print(f"Wall Insulation thickness: {predicted_enc[0][3]:.3f}")
print(f"Roof insulation thickness: {predicted_enc[0][5]:.3f}")
print(f"Window U value: {predicted_enc[0][6]:.3f}")
print(f"Wall Insulation: {predicted_wall_insulation}")
print(f"Roof Insulation: {predicted_roof_insulation}")
print(f"Window Glazing: {predicted_window_glazing}")
额外注意
  • 模型适配问题:线性回归不适合预测分类变量,建议针对3个非数值因变量单独训练分类模型(如逻辑回归、随机森林分类器),或者使用支持多任务学习的框架拆分回归和分类任务。
  • 路径处理:Windows系统下的文件路径建议用原始字符串(前缀加r),避免转义字符(如\T、\R)被解析成特殊字符。
  • 编码可靠性:OneHotEncoder的get_feature_names_out()方法能自动生成带前缀的列名,比手动拼接更不易出错。

内容的提问来源于stack exchange,提问作者Talha Siddiqui

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 11:54:55