Pandas使用assign+apply+lambda多列条件赋值报错的解决方法
问题:DataFrame按匹配项填充新列时的语法错误及修复方案
原始DataFrame
labItemsNameRef label 0 FBS decrease 1 FBS decrease 2 FBS increase 3 HbA1c decrease 4 Creatinine changeless ... ... ... 123901 FBS decrease 123902 HbA1c increase 123903 Micro Creatinine changeless 123904 DTX ก่อนอาหาร increase 123905 Urine Creatinine changeless
报错代码
df = df.assign( FBS = lambda df: df.apply(lambda x: x['label'] if x['labItemsNameRef'] == 'FBS'), HbA1c = lambda df: df.apply(lambda x: x['label'] if x['labItemsNameRef'] == 'HbA1c'), DTX = lambda df: df.apply(lambda x: x['label'] if x['labItemsNameRef'] == 'DTX'), BUN = lambda df: df.apply(lambda x: x['label'] if x['labItemsNameRef'] == 'BUN'), Creatinine = lambda df: df.apply(lambda x: x['label'] if x['labItemsNameRef'] == 'Creatinine'))
错误信息
FBX = lambda df: df.apply(lambda x: x['label'] if x['labItemsNameRef'] == 'FBX'), ^ SyntaxError: invalid syntax
期望输出
labItemsNameRef label FBS HbA1c Creatinine BUN DTX 0 FBS decrease decrease NaN NaN NaN NaN 1 FBS decrease decrease NaN NaN NaN NaN 2 FBS increase increase NaN NaN NaN NaN 3 HbA1c decrease NaN decrease NaN NaN NaN 4 Creatinine changeless NaN NaN changeless NaN NaN ... ... ... ... ... ... ... ... 123901 FBS decrease decrease NaN NaN NaN NaN 123902 HbA1c increase NaN increase NaN NaN NaN 123903 Micro Creatinine changeless NaN NaN NaN NaN NaN 123904 DTX ก่อนอาหาร increase NaN NaN NaN NaN NaN 123905 Urine Creatinine changeless NaN NaN NaN NaN NaN
问题原因
报错核心是Python三元表达式不完整,if分支必须搭配else分支;另外原代码漏掉了apply的axis=1参数(默认按列处理,不符合逐行判断需求)。同时,逐行apply的方式对12万行的大DataFrame效率极低,不推荐。
正确实现方案
方案1:用np.where(简洁高效,适合指定列)
import numpy as np # 定义需要生成的目标列 target_cols = ['FBS', 'HbA1c', 'DTX', 'BUN', 'Creatinine'] # 循环生成每一列 for col in target_cols: df[col] = np.where(df['labItemsNameRef'] == col, df['label'], np.nan)
方案2:用pivot(批量处理,适合动态列)
如果需要自动匹配所有可能的labItemsNameRef值,再筛选目标列:
# 创建临时列用于透视 df['temp'] = df['label'] # 透视生成对应列,缺失值自动填充NaN pivot_df = df.pivot(index=df.index, columns='labItemsNameRef', values='temp') # 只保留需要的目标列,不存在的列自动补NaN target_cols = ['FBS', 'HbA1c', 'DTX', 'BUN', 'Creatinine'] pivot_df = pivot_df.reindex(columns=target_cols) # 合并回原DataFrame df = df.join(pivot_df) # 删除临时列 df.drop('temp', axis=1, inplace=True)
修复原代码版本(仅作语法修正,不推荐)
如果一定要沿用assign+apply的写法,需补全三元表达式并添加axis=1:
import numpy as np df = df.assign( FBS=lambda df: df.apply(lambda x: x['label'] if x['labItemsNameRef'] == 'FBS' else np.nan, axis=1), HbA1c=lambda df: df.apply(lambda x: x['label'] if x['labItemsNameRef'] == 'HbA1c' else np.nan, axis=1), DTX=lambda df: df.apply(lambda x: x['label'] if x['labItemsNameRef'] == 'DTX' else np.nan, axis=1), BUN=lambda df: df.apply(lambda x: x['label'] if x['labItemsNameRef'] == 'BUN' else np.nan, axis=1), Creatinine=lambda df: df.apply(lambda x: x['label'] if x['labItemsNameRef'] == 'Creatinine' else np.nan, axis=1) )
内容的提问来源于stack exchange,提问作者Suphakit Suphapinyo
相关产品推荐
相关产品推荐

