如何基于条件对DataFrame列应用自定义函数生成新列/新DataFrame?
问题描述
我想在条件满足时应用函数,返回整数与选定向量的乘积并追加到DataFrame右侧,但始终得到异常结果。我是Python新手,求帮忙。
示例数据与代码
# 示例数据 data = {'Condition': ['C','D','A','C'], 'Values':[10,20,30,40], 'Ftype':[3,3,3,3] } df = pd.DataFrame(data) # 随字母变化的时间序列配置表 profile = {'profile_ID': ['A','B','C','D','E','F'], 'Period1':[0.1,0.0,0.1,0.0,0.2,0.1], 'Period2':[0.0,0.0,0.0,0.0,0.0,0.0], 'Period3':[0.1,0.1,0.1,0.1,0.1,0.1], 'Period4':[0.2,0.2,0.2,0.2,0.2,0.2]} profile = pd.DataFrame(profile) # 自定义函数 def Custom_function (Ftype): if Ftype > 2: return_array = df['Values'] * profile return return_array # 生成新DataFrame dfx=df['Ftype'].apply(Custom_function)
期望输出(仅展示第一行示例)
| Condition | Values | Ftype | Period1 | Period2 | Period3 | Period4 |
|---|---|---|---|---|---|---|
| C | 10 | 3 | 1 | 0 | 1 | 2 |
实际异常输出
0 0 1 2 3 Period1 Period2 Period3 ... 1 0 1 2 3 Period1 Period2 Period3 ... 2 0 1 2 3 Period1 Period2 Period3 ... 3 0 1 2 3 Period1 Period2 Period3 ... Name: Ftype, dtype: object
我试过用groupby匹配df和profile中的对应字母,但担心效率问题,现在卡壳了。
问题分析与解决方案
你的代码有几个核心问题,逐一修正就能得到想要的结果:
1. 核心问题点
- 函数里直接用
df['Values'] * profile,会让整列Values和整个profile表做广播运算,完全没关联每行的Condition和profile_ID。 - 函数直接引用全局的
df,而不是针对当前行处理,导致apply逐行执行时重复计算整表。
2. 正确实现步骤
步骤1:把profile转成字典,快速匹配
将profile表转成以profile_ID为键的字典,查找效率远高于groupby:
# 转换为字典,key是字母ID,value是对应的Period数值 profile_dict = profile.set_index('profile_ID').to_dict('index')
步骤2:修改自定义函数,逐行处理
让函数接收整行数据,这样能拿到当前行的Condition、Values,匹配对应profile后计算乘积:
def custom_function(row): if row['Ftype'] > 2: # 匹配当前行Condition对应的配置项 p = profile_dict.get(row['Condition'], {}) # 计算Values与各Period的乘积 return pd.Series({ 'Period1': row['Values'] * p.get('Period1', 0), 'Period2': row['Values'] * p.get('Period2', 0), 'Period3': row['Values'] * p.get('Period3', 0), 'Period4': row['Values'] * p.get('Period4', 0) }) # 条件不满足时返回全0 return pd.Series([0,0,0,0], index=['Period1','Period2','Period3','Period4'])
步骤3:生成结果并合并到原表
用apply逐行处理,再合并结果:
# 生成Period列数据 period_df = df.apply(custom_function, axis=1) # 合并到原DataFrame result_df = pd.concat([df, period_df], axis=1)
最终输出结果
运行后result_df的格式如下:
| Condition | Values | Ftype | Period1 | Period2 | Period3 | Period4 |
|---|---|---|---|---|---|---|
| C | 10 | 3 | 1.0 | 0.0 | 1.0 | 2.0 |
| D | 20 | 3 | 0.0 | 0.0 | 2.0 | 4.0 |
| A | 30 | 3 | 3.0 | 0.0 | 3.0 | 6.0 |
| C | 40 | 3 | 4.0 | 0.0 | 4.0 | 8.0 |
额外说明
- 字典查找的时间复杂度是O(1),比groupby循环效率高很多,完全不用担心性能问题。
- 如果后续新增Period列,只要保持profile表的格式,修改函数时可以动态生成列,不用手动写死。
内容的提问来源于stack exchange,提问作者Pgmr_Beginner
相关产品推荐
相关产品推荐

