Pandas中基于Top N值对多列进行分桶处理的实现疑问
Hey there! Let's tackle this problem step by step. You want to process multiple columns in a DataFrame, keep only the Top N values for each column (replacing others with "other"), and create new columns with these transformed values. The good news is you don't need to juggle both row and column references with apply—since each column can be processed independently, we can keep things straightforward.
Step 1: Example Setup
First, let's create a sample DataFrame to work with (we'll use categorical columns here, since that's a common use case for this task):
import pandas as pd import numpy as np # Sample categorical DataFrame df = pd.DataFrame({ 'Fruit': ['apple', 'banana', 'apple', 'orange', 'banana', 'banana', 'grape'], 'Animal': ['cat', 'dog', 'dog', 'cat', 'bird', 'dog', 'bird'] })
Step 2: Define Top N and Process Columns
Let's say we want to keep the Top 2 most frequent values in each column. Here's how to do it efficiently:
n = 2 # Define your Top N value # Loop through each column to create transformed new columns for col in df.columns: # Get the Top N most frequent values for the current column top_values = df[col].value_counts().head(n).index # Create a new column: keep value if it's in Top N, else "other" df[f'{col}_top{n}'] = np.where(df[col].isin(top_values), df[col], 'other')
Step 3: View the Result
Running the code above will modify your DataFrame to include the new transformed columns. Here's what the output looks like:
| Fruit | Animal | Fruit_top2 | Animal_top2 | |
|---|---|---|---|---|
| 0 | apple | cat | apple | cat |
| 1 | banana | dog | banana | dog |
| 2 | apple | dog | apple | dog |
| 3 | orange | cat | other | cat |
| 4 | banana | bird | banana | other |
| 5 | banana | dog | banana | dog |
| 6 | grape | bird | other | other |
Handling Numeric Columns
If your columns are numeric and you want the Top N largest values (instead of most frequent), adjust how you get top_values:
# Sample numeric DataFrame df_num = pd.DataFrame({ 'Score': [10, 20, 10, 30, 20, 20, 5], 'Age': [5, 3, 3, 5, 1, 3, 1] }) n = 2 for col in df_num.columns: # Get the Top N largest unique values for the current column top_values = df_num[col].nlargest(n).unique() df_num[f'{col}_top{n}'] = np.where(df_num[col].isin(top_values), df_num[col].astype(str), 'other')
Using apply (If You Prefer)
If you want to use apply to process columns in one go, you can define a helper function and apply it across columns:
def process_column(col, n): top_vals = col.value_counts().head(n).index return col.apply(lambda x: x if x in top_vals else 'other') # Apply the function to all columns and merge results back top_n_df = df.apply(process_column, n=2) top_n_df.columns = [f'{col}_top2' for col in df.columns] df = pd.concat([df, top_n_df], axis=1)
This achieves the same result as the loop approach—choose whichever feels more readable to you!
内容的提问来源于stack exchange,提问作者trystuff

