Pandas分类对比:示例DataFrame创建与分位数分段实操
Creating a Sample DataFrame and Applying Quantile Binning with Pandas
Here's a straightforward walkthrough of generating a sample DataFrame and using quantile-based binning to categorize values in a column:
1. Generate the Initial DataFrame
First, we'll create a 10-row, 2-column DataFrame with random values using NumPy and Pandas:
import pandas as pd import numpy as np df = pd.DataFrame(np.random.randn(10, 2), columns=list('AB'))
When run, this produces a DataFrame like this (your values will vary since we're using random data):
A B 0 0.459759 0.152645 1 0.183613 0.756527 2 -1.836027 0.032433 3 0.264336 0.170171 4 -0.276347 0.208389 5 0.677709 0.725274 6 -0.547858 0.376683 7 -0.994759 -0.750373 8 0.556593 1.282167 9 -1.444533 0.589768
2. Apply Quantile Binning
Next, we'll use pd.qcut() to split the values in column A into 4 quantile-based bins. The duplicates="drop" parameter ensures any duplicate bin edges are removed automatically to avoid errors:
df['A_rank'] = pd.qcut(df['A'], [0, 0.25, 0.5, 0.75, 1], duplicates="drop")
This adds a new column A_rank that labels each row with the quantile bin it falls into:
A B A_rank 0 0.459759 0.152645 (0.411, 0.678] 1 0.183613 0.756527 (-0.0464, 0.411] 2 -1.836027 0.032433 (-1.837, -0.883] 3 0.264336 0.170171 (-0.0464, 0.411] 4 -0.276347 0.208389 (-0.883, -0.0464] 5 0.677709 0.725274 (0.411, 0.678] 6 -0.547858 0.376683 (-0.883, -0.0464] 7 -0.994759 -0.750373 (-1.837, -0.883] 8 0.556593 1.282167 (0.411, 0.678] 9 -1.444533 0.589768 (-1.837, -0.883]
Quick Notes:
pd.qcut()divides data into bins such that each bin has roughly the same number of observations, which is perfect for quantile-based ranking.- The list
[0, 0.25, 0.5, 0.75, 1]defines our thresholds (25th, 50th, 75th percentiles). duplicates="drop"is handy when your dataset has repeated values that would cause overlapping bin edges.
内容的提问来源于stack exchange,提问作者SankMa
相关产品推荐
相关产品推荐

