如何为含布尔值列的DataFrame生成标记True列的新列并统计?
Pandas 布尔列标记与统计方案
1. 生成标记True列名的新列
方法一:直观写法(适合小数据集)
假设你的DataFrame叫df,布尔列是bol_1、bol_2、bol_3,一行代码就能生成目标列:
df['criteria'] = df.apply(lambda row: [col for col in df.columns if row[col]], axis=1) # 可选:把单元素列表转成字符串,和示例格式一致 df['criteria'] = df['criteria'].apply(lambda x: x[0] if len(x) == 1 else x)
方法二:高效写法(适合大数据集)
用矩阵乘法代替apply,速度更快:
# 把True对应列名,False对应空串,拼接后拆分 df['criteria'] = df.dot(df.columns + ',').str.rstrip(',').str.split(',') # 同样处理单元素转字符串 df['criteria'] = df['criteria'].apply(lambda x: x[0] if len(x) == 1 else x if x else None)
2. 筛选有至少一个True的行
直接用any按行判断:
# 提取所有存在True值的行 has_true_rows = df[df[['bol_1', 'bol_2', 'bol_3']].any(axis=1)]
3. 基础统计操作
统计bol_1作为唯一True值的行数
两种方式任选:
# 方式一:直接用原布尔列判断 count_bol1_only = len(df[(df['bol_1']) & (~df['bol_2']) & (~df['bol_3'])]) # 方式二:用生成的criteria列判断(如果已经转成字符串) count_bol1_only = df['criteria'].eq('bol_1').sum()
批量统计各列作为唯一True值的行数
循环处理所有布尔列,一次出结果:
bool_cols = ['bol_1', 'bol_2', 'bol_3'] for col in bool_cols: # 其他列必须全为False other_cols = [c for c in bool_cols if c != col] count = (df[col] & ~df[other_cols].any(axis=1)).sum() print(f"{col}作为唯一True值的行数:{count}")
统计各True组合的出现次数
如果criteria是列表,转成元组后用value_counts统计:
df['criteria_tuple'] = df['criteria'].apply(tuple) print(df['criteria_tuple'].value_counts())
内容的提问来源于stack exchange,提问作者An old man in the sea.
相关产品推荐
相关产品推荐

