如何在Pandas DataFrame中实现多列条件判断生成结果列
实现Pandas DataFrame多条件判断生成结果列
先看你的DataFrame定义:
import pandas as pd import numpy as np a = {'Col1': [0,1,1,0], 'Col2': [0,1,0,1]} df = pd.DataFrame(data = a)
方法1:用np.select处理多条件(最清晰)
np.select可以一次性定义多组条件与对应结果,比嵌套np.where更直观:
# 定义条件列表 conditions = [ (df['Col1'] == 1) & (df['Col2'] == 1), (df['Col1'] == 1) & (df['Col2'] == 0), (df['Col1'] == 0) & (df['Col2'] == 1), (df['Col1'] == 0) & (df['Col2'] == 0) ] # 对应结果列表 results = [ 'col1 + col2', 'col1', 'col2', 'other' ] # 生成result列 df['result'] = np.select(conditions, results)
这种方式条件与结果一一对应,后续新增条件也方便维护。
方法2:用apply逐行处理(适合列数多的场景)
如果后续要增加更多列,不想写大量固定条件判断,可以用apply遍历每行,动态获取符合条件的列名:
def get_result(row): # 筛选值为1的列名并转为小写 matched_cols = [col.lower() for col in row.index if row[col] == 1] if len(matched_cols) == 0: return 'other' elif len(matched_cols) == 2: return ' + '.join(matched_cols) else: return matched_cols[0] df['result'] = df.apply(get_result, axis=1)
这个方法扩展性强,哪怕新增Col3、Col4,只要逻辑是“值为1的列名拼接,全0返回other”,都不用修改代码。
方法3:嵌套np.where(基础实现)
如果一定要用np.where,可以通过多层嵌套实现需求:
df['result'] = np.where( (df['Col1'] == 0) & (df['Col2'] == 0), 'other', np.where( (df['Col1'] == 1) & (df['Col2'] == 1), 'col1 + col2', np.where(df['Col1'] == 1, 'col1', 'col2') ) )
不过嵌套层数过多时可读性会下降,不如前两种方法实用。
运行任意一种方法后,最终DataFrame结果如下:
Col1 Col2 result 0 0 0 other 1 1 1 col1 + col2 2 1 0 col1 3 0 1 col2
内容的提问来源于stack exchange,提问作者wysouf
相关产品推荐
相关产品推荐

