You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pandas apply新增条件列触发SpecificationError错误解决

错误原因

触发SpecificationError: Function names must be unique if there is no new column names assigned的核心原因是clin.apply(surv, axis=1)的参数传错:

  • pandas.DataFrame.apply()的第一个位置参数必须是可调用的处理函数,用来逐行/逐列计算生成结果
  • 你传入的surv是提前通过列表推导式生成好的结果值列表,不是可调用函数,pandas会误把列表元素识别为要调用的函数名,在找不到对应新列名配置、且识别到的"函数名"重复时就会抛出该报错。
正确实现方案

优先用pandas原生向量化运算实现,性能比逐行apply高几个量级,完全匹配赋值规则:

import numpy as np

# 先将目标字段转为要求的float32类型
clin["days_to_death"] = clin["days_to_death"].astype(np.float32)

# 按规则赋值:days_to_death >= 730(2*365)或为空则标记1,其余标记0
clin["SURV"] = np.where(
    (clin["days_to_death"] >= 2 * 365) | (clin["days_to_death"].isna()),
    1,
    0
)

如果你已经提前生成了和clin行数完全一致的surv结果列表,不需要调用apply,直接赋值即可:

clin["SURV"] = surv

如果一定要用逐行apply的写法(不推荐,大数据量下速度极慢),需要传入合法的判断函数:

import pandas as pd

def calc_surv(row):
    d = row["days_to_death"]
    return 1 if pd.isna(d) or d >= 2*365 else 0

clin["SURV"] = clin.apply(calc_surv, axis=1)

内容的提问来源于stack exchange,提问作者melolilili

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.02 05:24:52