如何在Pandas中实现指定的DataFrame列内容拼接需求?
解决方案
首先构造示例DataFrame:
import pandas as pd data = { 'tray': [0, 0, 1, 0, 0], 'bag': [1, 1, 1, 0, 0], 'ball': [1, 0, 0, 1, 0] } df = pd.DataFrame(data)
方法一:逐行筛选拼接(直观易读)
通过apply按行遍历,筛选值为1的列名并拼接:
df['Presence'] = df.apply( lambda row: ','.join([col for col, val in row.items() if val == 1]) or 'No Presence', axis=1 )
逻辑说明:
- 对每行遍历列名和对应值,收集值为1的列名
- 用逗号连接收集到的列名,若结果为空(全0行),则返回
No Presence
方法二:向量运算(高效适合大数据)
利用Pandas的向量运算实现,性能优于逐行遍历:
# 用dot乘积拼接列名,再处理末尾逗号和空值 df['Presence'] = df.dot(df.columns + ',').str.rstrip(',').replace('', 'No Presence')
逻辑说明:
df.dot(df.columns + ',')将每行中值为1的列名加逗号后累加str.rstrip(',')移除末尾多余的逗号replace('', 'No Presence')把全0行产生的空字符串替换为指定文本
最终得到的DataFrame如下:
| tray | bag | ball | Presence |
|---|---|---|---|
| 0 | 1 | 1 | bag,ball |
| 0 | 1 | 0 | bag |
| 1 | 1 | 0 | tray,bag |
| 0 | 0 | 1 | ball |
| 0 | 0 | 0 | No Presence |
内容的提问来源于stack exchange,提问作者Scope
相关产品推荐
相关产品推荐

