如何在DataFrame指定多列中提取最频繁值并新增列存储
Here's how you can add column F to your DataFrame, where F holds the most frequent character from columns A-E for each row:
First, let's recreate your sample DataFrame:
import pandas as pd # Create the original DataFrame data = { 'S': [1, 2, 3, 4], 'A': ['N', 'N', 'Y', 'Y'], 'B': ['N', 'Y', 'N', 'N'], 'C': ['N', 'Y', 'Y', 'Y'], 'D': ['N', 'N', 'N', 'Y'], 'E': ['N', 'N', 'N', 'Y'] } df = pd.DataFrame(data)
Next, we'll use apply() with a lambda function to calculate the most frequent value across columns A-E for each row. We use value_counts().idxmax() which gives us the value with the highest count:
# Add column F with the most frequent character from A-E df['F'] = df[['A', 'B', 'C', 'D', 'E']].apply(lambda row: row.value_counts().idxmax(), axis=1)
Alternatively, you can use pandas' mode() method, which directly returns the most frequent value(s) for each row. Since we have an odd number of columns (5), there will always be a single most frequent value, so we can take the first result:
# Alternative approach using mode() df['F'] = df[['A', 'B', 'C', 'D', 'E']].mode(axis=1)[0]
Either way, the resulting DataFrame will match your expected output:
S A B C D E F 0 1 N N N N N N 1 2 N Y Y N N N 2 3 Y N Y N N N 3 4 Y N Y Y Y Y
Explanation:
df[['A', 'B', 'C', 'D', 'E']]selects only the columns we want to analyze.apply(..., axis=1)runs the function across each row instead of columns.row.value_counts()counts occurrences of each character in the row, sorted from most to least frequent.idxmax()picks the value with the highest count (the first one in the sorted counts).- The
mode()method is a more concise way to get the most frequent value, and since we have 5 columns (no ties possible), taking the first column of the result gives us the correct value.
内容的提问来源于stack exchange,提问作者shome

