在R中按行计算变量集中值的加权出现次数并选取最大值
Alright, let's walk through exactly how to solve this problem with pandas. I'll break it down into simple, actionable steps with example code so you can adapt it to your specific df8 data frame.
First, let's assume your df8 has columns representing the distinct values you're tracking, and each row holds the count of how many times that value appears. For example:
import pandas as pd # Sample df8 (replace this with your actual data) df8 = pd.DataFrame({ 'value_A': [2, 5, 1], 'value_B': [3, 1, 4], 'value_C': [1, 2, 5] }) # Your weight vector V = (0.25, 0.25, 0.5)
Important: Make sure the order of columns in df8 matches the order of weights in V! If they don't, reorder your columns first (e.g., df8 = df8[['value_A', 'value_B', 'value_C']] to align with V's sequence).
We'll create a new data frame where each count is multiplied by its matching weight from V. Pandas makes this straightforward with element-wise multiplication:
# Calculate weighted values for each value per row weighted_values = df8 * V
Now, for each row, we need to pick the value (column name) that has the largest weighted result. Use idxmax(axis=1) to grab the column name of the maximum value in each row:
# Add a new column to df8 with the selected value df8['selected_value'] = weighted_values.idxmax(axis=1)
If multiple values in a row have the same maximum weighted value, idxmax will only return the first one. To capture all tied values, use a custom apply function:
def get_all_max_values(row): max_weight = row.max() # Return all column names where the weighted value equals the max return ', '.join(row[row == max_weight].index) # Add a column with all tied values (if any) df8['selected_values_all'] = weighted_values.apply(get_all_max_values, axis=1)
Running the full code with our sample df8 will give you something like this:
| value_A | value_B | value_C | selected_value | selected_values_all |
|---|---|---|---|---|
| 2 | 3 | 1 | value_B | value_B |
| 5 | 1 | 2 | value_A | value_A |
| 1 | 4 | 5 | value_C | value_C |
And the weighted values data frame will show the calculated values:
| value_A | value_B | value_C |
|---|---|---|
| 0.5 | 0.75 | 0.5 |
| 1.25 | 0.25 | 1.0 |
| 0.25 | 1.0 | 2.5 |
内容的提问来源于stack exchange,提问作者Avi

