You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何自然使用pd.cut对数据分箱并获取指定分箱数据?

How to get indices of data in the 4th equal-width bin from a pandas Series without manually setting labels?

I have a pandas Series with 100 float values, and I want to split it into 10 equal-width bins, then get the indices of the data in the 4th bin. Here's the code I tried:

import pandas as pd; import numpy as np
np.random.seed(1)
s = pd.Series(np.random.randn(100))
cut = pd.cut(s, bins=10, labels=range(10))
fourth_bin = s[cut == 4]
fourth_bin

Output:

9   -0.249370
12  -0.322417
13  -0.384054
16  -0.172428
26  -0.122890
28  -0.267888
31  -0.396754
40  -0.191836
51  -0.352250
53  -0.349343
54  -0.208894
63  -0.298093
65  -0.075572
71  -0.504466
76  -0.306204
80  -0.222328
81  -0.200758
92  -0.375285
96  -0.343854
dtype: float64

This approach feels a bit clunky and not very natural. I want to avoid manually setting labels and just use pd.cut(s, bins=10) directly. I tried something like s[s in pd.cut(s, bins=10).categories[4]] but it doesn't work. Is there a more intuitive way to do this?


Great question! Let's break down two clean, intuitive ways to achieve this without manually defining labels for pd.cut:

1. Use the codes attribute (simplest approach)

When you call pd.cut without specifying labels, it returns a Categorical object. The codes property of this object gives a 0-based integer array where each value maps to the bin index of the corresponding element in your Series. This is exactly like your manually set range(10) labels, but generated automatically.

To get the same bin you targeted with cut == 4 in your original code, just filter where codes == 4:

import pandas as pd; import numpy as np
np.random.seed(1)
s = pd.Series(np.random.randn(100))

# Generate bins without custom labels
cut_bins = pd.cut(s, bins=10)

# Filter elements in the 5th bin (0-based index 4, matching your original code)
fourth_bin = s[cut_bins.codes == 4]

If you actually meant the 4th bin when counting from 0 (i.e., the first bin is index 0), just change the condition to cut_bins.codes == 3.

2. Filter using the bin's interval bounds

If you want to work directly with the actual interval of the target bin, you can extract its lower and upper bounds and use Series.between() to filter elements. This is useful if you want to explicitly see which range you're targeting.

# Get the target bin interval (0-based index 4, same as your original code)
target_bin = cut_bins.categories[4]

# Extract lower and upper bounds
lower_bound = target_bin.left
upper_bound = target_bin.right

# Filter elements within the interval (matches pd.cut's default right-closed behavior)
fourth_bin = s[s.between(lower_bound, upper_bound, inclusive='both')]

Note: pd.cut uses right-closed intervals by default (except for the first bin), so inclusive='both' aligns with that behavior. If you modified the right parameter in pd.cut, adjust inclusive accordingly (e.g., inclusive='right').

Why your original attempt failed

The expression s in pd.cut(s, bins=10).categories[4] tries to check if the entire Series is contained within a single interval, which isn't a valid vectorized operation. You need element-wise checks, which either the codes attribute or between() handles seamlessly.


内容的提问来源于stack exchange,提问作者Vim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:49:44