You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas使用assign+apply+lambda多列条件赋值报错的解决方法

问题:DataFrame按匹配项填充新列时的语法错误及修复方案

原始DataFrame

labItemsNameRef     label
0       FBS                 decrease
1       FBS                 decrease
2       FBS                 increase
3       HbA1c               decrease
4       Creatinine          changeless
...    ...                  ...
123901  FBS                 decrease
123902  HbA1c               increase
123903  Micro Creatinine    changeless
123904  DTX ก่อนอาหาร       increase
123905  Urine Creatinine    changeless

报错代码

df = df.assign(
     FBS = lambda df: df.apply(lambda x: x['label'] if x['labItemsNameRef'] == 'FBS'),
     HbA1c = lambda df: df.apply(lambda x: x['label'] if x['labItemsNameRef'] == 'HbA1c'),
     DTX = lambda df: df.apply(lambda x: x['label'] if x['labItemsNameRef'] == 'DTX'),
     BUN = lambda df: df.apply(lambda x: x['label'] if x['labItemsNameRef'] == 'BUN'),
     Creatinine = lambda df: df.apply(lambda x: x['label'] if x['labItemsNameRef'] == 'Creatinine'))

错误信息

FBX = lambda df: df.apply(lambda x: x['label'] if x['labItemsNameRef'] == 'FBX'),
                                                                                   ^
SyntaxError: invalid syntax

期望输出

labItemsNameRef  label       FBS      HbA1c    Creatinine BUN DTX
0      FBS              decrease    decrease NaN      NaN        NaN    NaN
1      FBS              decrease    decrease NaN      NaN        NaN    NaN
2      FBS              increase    increase NaN      NaN        NaN    NaN
3      HbA1c            decrease    NaN      decrease NaN        NaN    NaN
4      Creatinine       changeless  NaN      NaN      changeless NaN    NaN
...     ...               ...       ...      ...      ...   ... ...
123901 FBS              decrease    decrease NaN      NaN        NaN    NaN
123902 HbA1c            increase    NaN      increase NaN        NaN    NaN
123903 Micro Creatinine changeless  NaN      NaN      NaN        NaN    NaN
123904 DTX ก่อนอาหาร     increase    NaN      NaN      NaN        NaN    NaN
123905 Urine Creatinine changeless  NaN      NaN      NaN        NaN    NaN

问题原因

报错核心是Python三元表达式不完整,if分支必须搭配else分支;另外原代码漏掉了apply的axis=1参数(默认按列处理,不符合逐行判断需求)。同时,逐行apply的方式对12万行的大DataFrame效率极低,不推荐。

正确实现方案

方案1:用np.where(简洁高效,适合指定列)

import numpy as np

# 定义需要生成的目标列
target_cols = ['FBS', 'HbA1c', 'DTX', 'BUN', 'Creatinine']

# 循环生成每一列
for col in target_cols:
    df[col] = np.where(df['labItemsNameRef'] == col, df['label'], np.nan)

方案2:用pivot(批量处理,适合动态列)

如果需要自动匹配所有可能的labItemsNameRef值,再筛选目标列:

# 创建临时列用于透视
df['temp'] = df['label']
# 透视生成对应列,缺失值自动填充NaN
pivot_df = df.pivot(index=df.index, columns='labItemsNameRef', values='temp')
# 只保留需要的目标列,不存在的列自动补NaN
target_cols = ['FBS', 'HbA1c', 'DTX', 'BUN', 'Creatinine']
pivot_df = pivot_df.reindex(columns=target_cols)
# 合并回原DataFrame
df = df.join(pivot_df)
# 删除临时列
df.drop('temp', axis=1, inplace=True)

修复原代码版本(仅作语法修正,不推荐)

如果一定要沿用assign+apply的写法,需补全三元表达式并添加axis=1:

import numpy as np

df = df.assign(
    FBS=lambda df: df.apply(lambda x: x['label'] if x['labItemsNameRef'] == 'FBS' else np.nan, axis=1),
    HbA1c=lambda df: df.apply(lambda x: x['label'] if x['labItemsNameRef'] == 'HbA1c' else np.nan, axis=1),
    DTX=lambda df: df.apply(lambda x: x['label'] if x['labItemsNameRef'] == 'DTX' else np.nan, axis=1),
    BUN=lambda df: df.apply(lambda x: x['label'] if x['labItemsNameRef'] == 'BUN' else np.nan, axis=1),
    Creatinine=lambda df: df.apply(lambda x: x['label'] if x['labItemsNameRef'] == 'Creatinine' else np.nan, axis=1)
)

内容的提问来源于stack exchange,提问作者Suphakit Suphapinyo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 05:15:27