You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中将超过80分位数的值替换为80分位数

How to Cap Values at the 80th Percentile in a Pandas DataFrame

Got it, let's walk through solving this problem step by step. You want to replace any values in the salary column that are above the 80th percentile with the 80th percentile value itself. Here's exactly how to implement this cleanly in pandas:

Step 1: Set Up Your DataFrame

First, let's recreate the original DataFrame you provided:

import pandas as pd
import numpy as np

# Original data
data = {
    'Id': [1, 2, 3, 4, 5, 6, 7, 8, 9, 10],
    'gender': ['m', 'm', 'm', 'f', 'm', 'f', 'm', 'f', 'm', 'f'],
    'age': [27, 26, 57, 27, 57, 29, 47, 27, 37, 43],
    'salary': [100, 100000, 180, 150, 200, 100, 130, 140, 100, 2000]
}

df = pd.DataFrame(data)

Step 2: Calculate the 80th Percentile

Next, compute the 80th percentile for the salary column. Pandas' quantile() method uses linear interpolation by default, which matches the 560 value in your expected output:

# Compute 80th percentile for salary
salary_80th = df['salary'].quantile(0.8)
# Output: 560.0

Step 3: Replace Values Above the Percentile

You have two straightforward ways to do this:

Option 1: Use clip() (Simplest Method)

The clip() method is designed exactly for this kind of value truncation. It lets you set an upper (and/or lower) bound for your column:

# Cap salary values at the 80th percentile
df['salary'] = df['salary'].clip(upper=salary_80th)

Option 2: Use np.where() (More Flexible)

If you want explicit conditional logic, numpy.where() works great—it replaces values where the condition is true with your target value:

# Replace values above the 80th percentile with the percentile value
df['salary'] = np.where(df['salary'] > salary_80th, salary_80th, df['salary'])

Step 4: Verify the Result

After running either method, your DataFrame will match the expected output:

Idgenderagesalary
1m27100
2m26560
3m57180
4f27150
5m57200
6f29100
7m47130
8f27140
9m37100
10f43560

Bonus: Apply to All Numeric Columns

If you ever need to cap values for all numeric columns (like age and salary), you can loop through numeric columns:

# Get all numeric columns
numeric_cols = df.select_dtypes(include=['int64', 'float64']).columns

# Cap each numeric column at its 80th percentile
for col in numeric_cols:
    col_80th = df[col].quantile(0.8)
    df[col] = df[col].clip(upper=col_80th)

内容的提问来源于stack exchange,提问作者Danish

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 10:12:38