Python:值为1时拼接列名到新列,解决空列表统计问题
解决方法:生成符合要求的合并条件列并支持统计
示例数据
| alcoholism | diabites | handicapped | hypertensive | 目标新列 |
|---|---|---|---|---|
| 1 | 0 | 1 | 0 | alcoholism, handicapped |
| 0 | 1 | 0 | 1 | diabites, hypertensive |
| 0 | 1 | 0 | 0 | diabites |
需求
- 当上述任意列值为1时,新列保留对应列名,用逗号分隔
- 所有列值均为0时,新列返回
no condition
现有问题
之前尝试的代码返回结果是列表格式(带方括号,如[handicapped, alcoholism]),全0时为空列表[],无法直接使用value_counts()统计或绘图。
修正后的代码
方法一:逐行处理(直观易读)
先确保列名和数据中的完全匹配,再定义函数生成目标字符串:
# 匹配数据中的真实列名 problems = ['alcoholism', 'diabites', 'handicapped', 'hypertensive'] def get_conditions(row): # 筛选出值为1的列名 matched = [col for col in problems if row[col] == 1] return ', '.join(matched) if matched else 'no condition' # 生成新列 df['sp_name'] = df.apply(get_conditions, axis=1)
方法二:矩阵运算(高效适合大数据集)
利用dot方法批量拼接字符串,效率比逐行apply更高:
problems = ['alcoholism', 'diabites', 'handicapped', 'hypertensive'] # 生成匹配的列名字符串,自动用逗号分隔 cond_str = df[problems].eq(1).dot([f"{col}, " for col in problems]).str.rstrip(', ') # 把空字符串替换为指定内容 df['sp_name'] = cond_str.replace('', 'no condition')
效果验证
生成的sp_name列是纯字符串格式:
- 符合条件时:
alcoholism, handicapped - 全0时:
no condition
可以直接执行df['sp_name'].value_counts()统计各类别数量,也能正常用于后续绘图操作。
内容的提问来源于stack exchange,提问作者AHMED_HELMY
相关产品推荐
相关产品推荐

