Pandas对LotShape列Lambda编码异常及assign报错问题求助
问题排查与解决方案
第一种尝试的问题分析
- 函数参数与逻辑错误:
- 你写的
irreg函数接收的是整个df['LotShape']Series,但map是对df['LotShape']的每个元素逐个处理,函数应该接收单个元素而非整个Series。 - 条件判断
if 'IR1' or 'IR2' or 'IR3' in df['LotShape']逻辑完全错误:'IR1'是字符串常量,本身为真,导致这个条件永远成立,所有结果都会返回2。正确逻辑是判断当前元素是否在['IR1','IR2','IR3']列表中。
- 你写的
修正后的第一种写法:
def irreg(x): if x in ['IR1', 'IR2', 'IR3']: return 2 else: return 1 df['LotShape_encoded'] = df['LotShape'].apply(irreg) # 或者用map更简洁 # df['LotShape_encoded'] = df['LotShape'].map(lambda x: 2 if x in ['IR1','IR2','IR3'] else 1) # 验证结果 print(df.LotShape_encoded.value_counts())
第二种尝试的问题分析
- 循环逻辑无效:遍历
df['LotShape']的元素i时,修改i的值不会对原Series产生任何影响,最终返回的还是原始列,完全没实现编码。 - assign用法错误:
df.assign()需要传入关键字参数定义新列,比如assign(LotShape_encoded=lambda x: ...),你直接把lambda作为位置参数传入,导致参数数量不匹配报错。
更简洁的Pandas原生写法
不需要自定义函数,用Pandas内置方法更高效:
方法1:使用replace
df['LotShape_encoded'] = df['LotShape'].replace({'Reg':1, 'IR1':2, 'IR2':2, 'IR3':2})
方法2:使用np.where
import numpy as np df['LotShape_encoded'] = np.where(df['LotShape'].isin(['IR1','IR2','IR3']), 2, 1)
方法3:使用apply(简化版)
df['LotShape_encoded'] = df['LotShape'].apply(lambda x: 2 if x.startswith('IR') else 1)
以上几种方法都能正确实现编码逻辑,验证后结果会和原始数据对应:1对应925条,2对应484+41+10=535条。
内容的提问来源于stack exchange,提问作者Jennifer Crosby
相关产品推荐
相关产品推荐

