如何在Pandas DataFrame中基于列值为1的条件创建含对应列名的新列
Solution: Create Column with Names of Columns Where Value is 1
Let's start with your sample DataFrame to work through the solution:
import pandas as pd df = pd.DataFrame({ 'id': [1,2,3,4], 'attr1': [1,1,0,0], 'attr2': [0,1,1,0], 'attr3': [1,1,1,0], 'attr4': [1,1,1,1] })
Method 1: Simple Row-Wise apply (Readable for Small Data)
This approach uses row-wise apply to collect column names where the value equals 1. We'll exclude the id column since it doesn't use 0/1 values (adjust this if you want to include it!).
# Define which columns to check (exclude 'id') target_columns = df.columns.drop('id') # Add new column with list of matching attribute names df['selected_attrs'] = df[target_columns].apply( lambda row: list(row[row == 1].index), axis=1 # Apply function across each row )
Output:
id attr1 attr2 attr3 attr4 selected_attrs 0 1 1 0 1 1 [attr1, attr3, attr4] 1 2 1 1 1 1 [attr1, attr2, attr3, attr4] 2 3 0 1 1 1 [attr2, attr3, attr4] 3 4 0 0 0 1 [attr4]
Method 2: Comma-Separated String Instead of List
If you want a human-readable string instead of a list, modify the lambda to join the names:
df['selected_attrs_str'] = df[target_columns].apply( lambda row: ', '.join(row[row == 1].index), axis=1 )
Output:
id attr1 attr2 attr3 attr4 selected_attrs selected_attrs_str 0 1 1 0 1 1 [attr1, attr3, attr4] attr1, attr3, attr4 1 2 1 1 1 1 [attr1, attr2, attr3, attr4] attr1, attr2, attr3, attr4 2 3 0 1 1 1 [attr2, attr3, attr4] attr2, attr3, attr4 3 4 0 0 0 1 [attr4] attr4
Method 3: Vectorized Approach (Faster for Large Datasets)
For bigger DataFrames, apply can be slow. Use this vectorized method with dot for better performance:
# Multiply row values (0/1) by column names, then clean up trailing punctuation df['selected_attrs_fast'] = df[target_columns].dot(target_columns + ', ').str.rstrip(', ')
This works by "multiplying" each 1 in the row with its column name (plus a comma), then stripping the trailing comma and space.
Quick Notes:
- If you want to include the
idcolumn in the check, just remove thetarget_columnsline and usedfdirectly in the apply/dot functions. - All methods preserve your original DataFrame structure while adding the new column.
内容的提问来源于stack exchange,提问作者saurabh kumar
相关产品推荐
相关产品推荐

