You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas中基于Top N值对多列进行分桶处理的实现疑问

Hey there! Let's tackle this problem step by step. You want to process multiple columns in a DataFrame, keep only the Top N values for each column (replacing others with "other"), and create new columns with these transformed values. The good news is you don't need to juggle both row and column references with apply—since each column can be processed independently, we can keep things straightforward.

Step 1: Example Setup

First, let's create a sample DataFrame to work with (we'll use categorical columns here, since that's a common use case for this task):

import pandas as pd
import numpy as np

# Sample categorical DataFrame
df = pd.DataFrame({
    'Fruit': ['apple', 'banana', 'apple', 'orange', 'banana', 'banana', 'grape'],
    'Animal': ['cat', 'dog', 'dog', 'cat', 'bird', 'dog', 'bird']
})

Step 2: Define Top N and Process Columns

Let's say we want to keep the Top 2 most frequent values in each column. Here's how to do it efficiently:

n = 2  # Define your Top N value

# Loop through each column to create transformed new columns
for col in df.columns:
    # Get the Top N most frequent values for the current column
    top_values = df[col].value_counts().head(n).index
    # Create a new column: keep value if it's in Top N, else "other"
    df[f'{col}_top{n}'] = np.where(df[col].isin(top_values), df[col], 'other')

Step 3: View the Result

Running the code above will modify your DataFrame to include the new transformed columns. Here's what the output looks like:

FruitAnimalFruit_top2Animal_top2
0applecatapplecat
1bananadogbananadog
2appledogappledog
3orangecatothercat
4bananabirdbananaother
5bananadogbananadog
6grapebirdotherother

Handling Numeric Columns

If your columns are numeric and you want the Top N largest values (instead of most frequent), adjust how you get top_values:

# Sample numeric DataFrame
df_num = pd.DataFrame({
    'Score': [10, 20, 10, 30, 20, 20, 5],
    'Age': [5, 3, 3, 5, 1, 3, 1]
})

n = 2
for col in df_num.columns:
    # Get the Top N largest unique values for the current column
    top_values = df_num[col].nlargest(n).unique()
    df_num[f'{col}_top{n}'] = np.where(df_num[col].isin(top_values), df_num[col].astype(str), 'other')

Using apply (If You Prefer)

If you want to use apply to process columns in one go, you can define a helper function and apply it across columns:

def process_column(col, n):
    top_vals = col.value_counts().head(n).index
    return col.apply(lambda x: x if x in top_vals else 'other')

# Apply the function to all columns and merge results back
top_n_df = df.apply(process_column, n=2)
top_n_df.columns = [f'{col}_top2' for col in df.columns]
df = pd.concat([df, top_n_df], axis=1)

This achieves the same result as the loop approach—choose whichever feels more readable to you!

内容的提问来源于stack exchange,提问作者trystuff

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:06:49