Pandas根据列中值为1的条件获取对应列名并填充至指定列
Here are a couple of straightforward, pandas-native ways to get the result you want:
Method 1: Use idxmax()
Since each row has exactly one column with a value of 1, idxmax(axis=1) will directly return the column name where that 1 is located. We just need to target the relevant status columns (ignoring the empty choice column first):
import pandas as pd # Original DataFrame sample = pd.DataFrame({'choice': [None, None, None], 'Active': [0, 0, 1], 'Completed': [1, 0, 0], 'Cancelled': [0, 1, 0]}) # Overwrite the 'choice' column with the matching status name sample['choice'] = sample[['Active', 'Completed', 'Cancelled']].idxmax(axis=1) # If you only want the 'choice' column in the final output: sample = sample[['choice']]
This will give you exactly the DataFrame you're looking for:
choice 0 Completed 1 Cancelled 2 Active
Method 2: Dot product trick
Another concise approach uses the dot product between the status columns and their names. Since only one value per row is 1, the dot product will resolve to the corresponding column name:
status_columns = ['Active', 'Completed', 'Cancelled'] sample['choice'] = sample[status_columns].dot(status_columns)
This works just as well as the first method, and both are way more efficient than looping through rows manually (they use pandas' vectorized operations under the hood).
内容的提问来源于stack exchange,提问作者Nikko

