应用返回字典的自定义函数时,如何高效为DataFrame新增列?
问题:扩展DataFrame的高效实现方式?
现有一个DataFrame对象df,自定义函数tFunc定义如下:
import random def tFunc(row): if (random.random()>0.5): info={'A': 'a', 'B': 'b', 'C': 'c', 'D': 'd'} else: info={'A': 'a', 'B': 'b', 'C': 'c'} return info # Workaround # for key in info: # row[key]=info[key] # return row df.apply(tFunc, axis=1, result_type='expand')
该函数返回字典info,希望扩展现有df,请问是否存在比上述注释中的临时解决方案更优、更快的实现方式?
高效实现方案
方案一:直接用join合并扩展结果
你当前写的df.apply(tFunc, axis=1, result_type='expand')本身就能把字典展开成新的DataFrame,直接用join和原DataFrame合并即可,比注释里的逐行修改方法高效得多:
df = df.join(df.apply(tFunc, axis=1, result_type='expand'))
- 优势:代码简洁,
result_type='expand'会自动处理字典中缺失的键(比如部分行没有D键时,对应位置填充NaN),无需手动循环处理每一行。
方案二:预生成结果再拼接(性能最优)
如果DataFrame行数较多,apply的逐行处理会有性能瓶颈,建议预先生成所有行的结果列表,再转成DataFrame后拼接:
import pandas as pd import random def tFunc(row): if random.random() > 0.5: return {'A': 'a', 'B': 'b', 'C': 'c', 'D': 'd'} else: return {'A': 'a', 'B': 'b', 'C': 'c'} # 预生成所有行的结果 results = [tFunc(None) for _ in range(len(df))] # 转换为DataFrame new_cols = pd.DataFrame(results) # 合并到原DataFrame df = pd.concat([df, new_cols], axis=1)
- 优势:避免了
apply的逐行循环开销,列表推导生成结果的速度远快于apply,适合大数据量场景。
补充说明
- 如果原DataFrame已存在
A、B等列,concat或join会自动给重复列名添加后缀(如A_x、A_y);若需要覆盖原列,可直接赋值:df[new_cols.columns] = new_cols。 - 注释里的逐行修改
row的方法不仅性能差,还可能因为row是原DataFrame的视图/副本导致意外问题,不推荐使用。
内容的提问来源于stack exchange,提问作者Dr.PB
相关产品推荐
相关产品推荐

