You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas DataFrame中基于列值为1的条件创建含对应列名的新列

Solution: Create Column with Names of Columns Where Value is 1

Let's start with your sample DataFrame to work through the solution:

import pandas as pd

df = pd.DataFrame({
    'id': [1,2,3,4], 
    'attr1': [1,1,0,0], 
    'attr2': [0,1,1,0], 
    'attr3': [1,1,1,0], 
    'attr4': [1,1,1,1]
})

Method 1: Simple Row-Wise apply (Readable for Small Data)

This approach uses row-wise apply to collect column names where the value equals 1. We'll exclude the id column since it doesn't use 0/1 values (adjust this if you want to include it!).

# Define which columns to check (exclude 'id')
target_columns = df.columns.drop('id')

# Add new column with list of matching attribute names
df['selected_attrs'] = df[target_columns].apply(
    lambda row: list(row[row == 1].index), 
    axis=1  # Apply function across each row
)

Output:

id  attr1  attr2  attr3  attr4          selected_attrs
0   1      1      0      1      1  [attr1, attr3, attr4]
1   2      1      1      1      1  [attr1, attr2, attr3, attr4]
2   3      0      1      1      1      [attr2, attr3, attr4]
3   4      0      0      0      1                  [attr4]

Method 2: Comma-Separated String Instead of List

If you want a human-readable string instead of a list, modify the lambda to join the names:

df['selected_attrs_str'] = df[target_columns].apply(
    lambda row: ', '.join(row[row == 1].index), 
    axis=1
)

Output:

id  attr1  attr2  attr3  attr4          selected_attrs         selected_attrs_str
0   1      1      0      1      1  [attr1, attr3, attr4]  attr1, attr3, attr4
1   2      1      1      1      1  [attr1, attr2, attr3, attr4]  attr1, attr2, attr3, attr4
2   3      0      1      1      1      [attr2, attr3, attr4]  attr2, attr3, attr4
3   4      0      0      0      1                  [attr4]  attr4

Method 3: Vectorized Approach (Faster for Large Datasets)

For bigger DataFrames, apply can be slow. Use this vectorized method with dot for better performance:

# Multiply row values (0/1) by column names, then clean up trailing punctuation
df['selected_attrs_fast'] = df[target_columns].dot(target_columns + ', ').str.rstrip(', ')

This works by "multiplying" each 1 in the row with its column name (plus a comma), then stripping the trailing comma and space.

Quick Notes:

  • If you want to include the id column in the check, just remove the target_columns line and use df directly in the apply/dot functions.
  • All methods preserve your original DataFrame structure while adding the new column.

内容的提问来源于stack exchange,提问作者saurabh kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:13:29