如何在Pandas中将超过80分位数的值替换为80分位数
Got it, let's walk through solving this problem step by step. You want to replace any values in the salary column that are above the 80th percentile with the 80th percentile value itself. Here's exactly how to implement this cleanly in pandas:
Step 1: Set Up Your DataFrame
First, let's recreate the original DataFrame you provided:
import pandas as pd import numpy as np # Original data data = { 'Id': [1, 2, 3, 4, 5, 6, 7, 8, 9, 10], 'gender': ['m', 'm', 'm', 'f', 'm', 'f', 'm', 'f', 'm', 'f'], 'age': [27, 26, 57, 27, 57, 29, 47, 27, 37, 43], 'salary': [100, 100000, 180, 150, 200, 100, 130, 140, 100, 2000] } df = pd.DataFrame(data)
Step 2: Calculate the 80th Percentile
Next, compute the 80th percentile for the salary column. Pandas' quantile() method uses linear interpolation by default, which matches the 560 value in your expected output:
# Compute 80th percentile for salary salary_80th = df['salary'].quantile(0.8) # Output: 560.0
Step 3: Replace Values Above the Percentile
You have two straightforward ways to do this:
Option 1: Use clip() (Simplest Method)
The clip() method is designed exactly for this kind of value truncation. It lets you set an upper (and/or lower) bound for your column:
# Cap salary values at the 80th percentile df['salary'] = df['salary'].clip(upper=salary_80th)
Option 2: Use np.where() (More Flexible)
If you want explicit conditional logic, numpy.where() works great—it replaces values where the condition is true with your target value:
# Replace values above the 80th percentile with the percentile value df['salary'] = np.where(df['salary'] > salary_80th, salary_80th, df['salary'])
Step 4: Verify the Result
After running either method, your DataFrame will match the expected output:
| Id | gender | age | salary |
|---|---|---|---|
| 1 | m | 27 | 100 |
| 2 | m | 26 | 560 |
| 3 | m | 57 | 180 |
| 4 | f | 27 | 150 |
| 5 | m | 57 | 200 |
| 6 | f | 29 | 100 |
| 7 | m | 47 | 130 |
| 8 | f | 27 | 140 |
| 9 | m | 37 | 100 |
| 10 | f | 43 | 560 |
Bonus: Apply to All Numeric Columns
If you ever need to cap values for all numeric columns (like age and salary), you can loop through numeric columns:
# Get all numeric columns numeric_cols = df.select_dtypes(include=['int64', 'float64']).columns # Cap each numeric column at its 80th percentile for col in numeric_cols: col_80th = df[col].quantile(0.8) df[col] = df[col].clip(upper=col_80th)
内容的提问来源于stack exchange,提问作者Danish

