如何在Pandas DataFrame的多列上使用apply方法?
问题解决:批量处理Pandas多列的lambda函数报错
原DataFrame如下:
| ID | Col1 | Col2 |
|---|---|---|
| 1 | 0 | 0 |
| 2 | 3 | 1 |
| 3 | 1 | 6 |
你尝试批量处理Col1和Col2列的代码如下:
column_list = ['Col1', 'Col2'] pivot_df[column_list] = pivot_df[column_list].apply(lambda x: 1 if x>1 else x)
触发的错误信息:
ValueError: The truth value of a Series is ambiguous. Use a.empty, a.bool(), a.item(), a.any() or a.all().
错误原因
DataFrame.apply() 默认按列处理(axis=0),此时lambda里的x是一整列的Series,x>1会生成由布尔值组成的Series,if语句无法直接判断整个Series的真假,因此报错。
正确解决方案
方案1:用applymap()做元素级处理
applymap()是针对DataFrame的元素级操作,会遍历每个单元格应用lambda函数:
import pandas as pd # 构造原DataFrame pivot_df = pd.DataFrame({'ID': [1,2,3], 'Col1': [0,3,1], 'Col2': [0,1,6]}) column_list = ['Col1', 'Col2'] # 批量处理目标列 pivot_df[column_list] = pivot_df[column_list].applymap(lambda x: 1 if x>1 else x) print(pivot_df)
执行后输出结果:
ID Col1 Col2 0 1 0 0 1 2 1 1 2 3 1 1
方案2:用向量化np.where()(效率更高)
对于数值型列,numpy的向量化操作比循环类方法效率更高,适合大数据量场景:
import pandas as pd import numpy as np pivot_df = pd.DataFrame({'ID': [1,2,3], 'Col1': [0,3,1], 'Col2': [0,1,6]}) column_list = ['Col1', 'Col2'] pivot_df[column_list] = np.where(pivot_df[column_list] > 1, 1, pivot_df[column_list]) print(pivot_df)
输出结果与方案1完全一致。
内容的提问来源于stack exchange,提问作者rshfq
相关产品推荐
相关产品推荐

