如何为Pandas DataFrame每行拼接值为1的列名生成新列?
为Pandas DataFrame生成Category列:拼接值为1的列名
嘿,我来帮你搞定这个需求!你需要给每行数据新增一个Category列,把该行中值为1的列名拼接起来,没有符合条件的就填NaN,对吧?先看看你的原始数据:
首先是创建DataFrame的代码:
import pandas as pd from pandas import DataFrame df = DataFrame({'A':['Cat had a nap','Dog had puppies','Did you see a Donkey','kitten got angry','puppy was cute'], 'Cat':[1,0,0,1,0], 'Dog':[0,1,0,0,1]})
原始数据长这样:
| A | Cat | Dog | |
|---|---|---|---|
| 0 | Cat had a nap | 1 | 0 |
| 1 | Dog had puppies | 0 | 1 |
| 2 | Did you see a Donkey | 0 | 0 |
| 3 | kitten got angry | 1 | 0 |
| 4 | puppy was cute | 0 | 1 |
下面是两种高效的解决方案:
方法一:用dot方法(推荐,效率更高)
这种方法利用Pandas的向量化操作,比逐行遍历快很多,适合大数据量:
# 指定要检查的目标列(这里是Cat和Dog) target_cols = ['Cat', 'Dog'] # 生成Category列 df['Category'] = df[target_cols].dot(target_cols + ',').str.rstrip(',').replace('', pd.NA)
步骤解释:
df[target_cols].dot(target_cols + ','):将每行中值为1的列名后面加上逗号拼接,比如第一行得到"Cat,".str.rstrip(','):去掉末尾多余的逗号,变成"Cat".replace('', pd.NA):把空字符串(对应没有值为1的行)替换成NaN
方法二:用apply逐行处理(更直观)
如果你更喜欢逻辑清晰的逐行处理,这个方法也很实用:
target_cols = ['Cat', 'Dog'] def get_category(row): # 收集当前行中值为1的列名 matching_cols = [col for col in target_cols if row[col] == 1] # 有符合条件的就拼接,没有就返回NaN return ','.join(matching_cols) if matching_cols else pd.NA # 应用函数到每行 df['Category'] = df.apply(get_category, axis=1)
最终结果
两种方法都会得到相同的结果:
print(df)
输出:
A Cat Dog Category 0 Cat had a nap 1 0 Cat 1 Dog had puppies 0 1 Dog 2 Did you see a Donkey 0 0 NaN 3 kitten got angry 1 0 Cat 4 puppy was cute 0 1 Dog
内容的提问来源于stack exchange,提问作者Sharvari Gc
相关产品推荐
相关产品推荐

