You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何避免LabelEncoder将y列误添加到X的numpy数组中

问题描述

我写了如下代码,但是执行标签编码器的最后一行:
X = MultiColumnLabelEncoder(columns = ['newlyConst','balcony', 'cellar', 'lift', 'garden', ]).fit_transform(df)
时,会把y列(租金列rent)添加到X的numpy数组中。
我不确定怎么用其他方式指定需要编码的列来避免这个问题,比如直接指定X的np数组和对应列而不是通过df操作,但是我这么做的时候会报索引错误。
非常感谢大家的帮助!

更新

我已经用更简洁的方案替换了冗长的标签编码器,替换方案如下:

df = df.astype({"newlyConst" :int, "balcony" : int, "cellar" : int, "lift" : int, "garden":int})

完整代码

import numpy as np
import matplotlib.pyplot as plt
import pandas as pd
from sklearn import linear_model


df = pd.read_csv('immo_data.csv')


df.drop(columns=['telekomTvOffer', 'telekomHybridUploadSpeed', 'pricetrend',
        'telekomUploadSpeed', 'scoutId', 'noParkSpaces', 'yearConstructedRange',
        'houseNumber', 'interiorQual', 'petsAllowed', 'street', 'streetPlain', 'baseRentRange',
        'geo_plz','geo_bln', 'geo_krs','thermalChar', 'floor','numberOfFloors', 'noRoomsRange', 'livingSpaceRange',
        'regio3', 'description', 'facilities', 'hasKitchen','heatingCosts', 'energyEfficiencyClass',
        'lastRefurbish', 'electricityBasePrice', 'electricityKwhPrice','date','condition', 'typeOfFlat','serviceCharge'
        ,'heatingType','firingTypes', 'yearConstructed'], axis=1, inplace = True)

df_head=df.head(250)

df_nan_count=df.isna().sum()
# 由于'firingTypes', 'yearConstructed', 'condition', 'typeOfFlat'列的NaN占比超过40-50%,已删除这些列

df.dropna(inplace=True)

df3=df.count()

df=df[['regio1', 'newlyConst', 'balcony', 'picturecount', 'cellar', 'livingSpace', 
       'lift','noRooms', 'garden', 'baseRent', 'totalRent']]


dfcount = df.nunique()


## 回归分析
X = df.iloc[:, :-1].values
y = df.iloc[:, -1].values



# 编码分类数据
from sklearn.preprocessing import LabelEncoder
from sklearn.pipeline import Pipeline
le = LabelEncoder()



class MultiColumnLabelEncoder:
    def __init__(self,columns = None):
        self.columns = columns # 待编码的列名数组

    def fit(self,X,y=None):
        return self # 此处不需要相关逻辑
   
        '''
        用LabelEncoder()转换X中self.columns指定的列,若未指定列,则转换X中所有列
       '''

    def transform(self,X):
      
        output = X.copy()
        if self.columns is not None:
            for col in self.columns:
                output[col] = LabelEncoder().fit_transform(output[col])
        else:
            for colname,col in output.iteritems():
                output[colname] = LabelEncoder().fit_transform(col)
        return output

    def fit_transform(self,X,y=None):
        return self.fit(X,y).transform(X)

X = MultiColumnLabelEncoder(columns = ['newlyConst','balcony', 'cellar', 'lift', 'garden', ]).fit_transform(df)



# 编码分类数据
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder
ct = ColumnTransformer(transformers=[('encoder', OneHotEncoder(), [0])], remainder='passthrough')
X = np.array(ct.fit_transform(X))

示例df数据

regio1  newlyConst  balcony  picturecount  cellar  livingSpace  lift  noRooms  garden  totalRent
0   Nordrhein_Westfalen           0        0             6       1        86.00     0      4.0       1     840.00
2               Sachsen           1        1             8       1        83.80     1      3.0       0    1300.00
4                Bremen           0        1            19       0        84.97     0      3.0       0     903.00
6               Sachsen           0        0             9       1        62.00     0      2.0       1     380.00
7                Bremen           0        1             5       1        60.30     0      3.0       0     584.25
8     Baden_Württemberg           0        0             5       1        53.00     0      2.0       0     690.00
10              Sachsen           0        1            11       1        40.20     0      2.0       0     307.00
11              Sachsen           0        0             9       1        80.00     0      3.0       1     555.00
12      Rheinland_Pfalz           0        0             4       0       100.00     0      4.0       1     920.00
13  Nordrhein_Westfalen           0        0             3       0       123.44     0      4.0       0    1150.00

问题原因及解决方案

根本原因

你出现y列被包含进X的核心原因是:调用MultiColumnLabelEncoder.fit_transform时传入的是完整的df(包含最后一列租金标签totalRent),而不是你之前拆分好的仅含特征的df.iloc[:, :-1],编码器会完整保留输入的所有列,自然会把y列也带进去。
另外你直接传入numpy数组报错的原因是:自定义的MultiColumnLabelEncoder是通过列名来定位需要编码的列的,numpy数组没有列名属性,无法匹配你指定的列名字符串参数,就会触发索引错误。

修复方案

方案1:修改编码器调用参数

只传入特征部分的DataFrame即可,把编码器调用行修改为:

# 仅传入除最后一列y之外的特征数据
X = MultiColumnLabelEncoder(columns = ['newlyConst','balcony', 'cellar', 'lift', 'garden']).fit_transform(df.iloc[:, :-1])

方案2:使用你更新的astype转换方案(更推荐)

你需要编码的这几列本身就是0/1二值分类,直接转int类型即可,不需要自定义编码器,逻辑更简洁:

# 先转换指定列的类型
df = df.astype({"newlyConst" :int, "balcony" : int, "cellar" : int, "lift" : int, "garden":int})
# 再拆分特征和标签
X = df.iloc[:, :-1]
y = df.iloc[:, -1]

后续再走OneHot编码逻辑即可,不会再出现y列混入X的问题。

内容的提问来源于stack exchange,提问作者JamesArthur

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 14:57:03