You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用pandas.cut()对含非数值的数组按指定规则分箱分类

pandas数组分箱分类实现方案

核心处理逻辑

  • 先清洗原始数组的字符串格式问题:去除字符串首尾空格,修正格式异常的数字串
  • 识别非数值内容(nan、文本值),统一归为not_number分类
  • 对有效数值使用pd.cut()按指定区间分箱
  • 按指定分类顺序统计计数

完整可运行代码

import pandas as pd
import numpy as np

# 原始去重数组(修正原代码漏写逗号、引号的语法笔误,保留原始带空格的'22 '场景)
arr = np.array(['10','8', '15','20','21','22 ', '27','28', np.nan, '29', '30', '32', '33', 'Values'])

# 第一步:预处理数据,统一清洗字符串空格,转数值类型,转失败的内容标记为NaN
cleaned_str = [x.strip() if isinstance(x, str) else x for x in arr]
num_col = pd.to_numeric(cleaned_str, errors='coerce')

# 第二步:初始化分类列,先标记非数值类
res = pd.Series(index=range(len(arr)), dtype='object')
res[num_col.isna()] = 'not_number'

# 第三步:对有效数值执行pd.cut分箱
bin_edges = [-np.inf, 10, 32, np.inf]
bin_labels = ['10 and below', '11-32', '33 and up']
res[~num_col.isna()] = pd.cut(num_col[~num_col.isna()], bins=bin_edges, labels=bin_labels).astype(str)

# 第四步:按指定分类顺序统计计数
count_stats = res.value_counts().reindex(['not_number', '10 and below', '11-32', '33 and up'])
print(count_stats)

代码说明

  • 预处理阶段的strip()操作专门处理类似'22 '这类带首尾空格的数字字符串,避免其无法正常转换为数值被误判为非数值类
  • pd.to_numeric(errors='coerce')会将无法转为数值的内容(包括原数组的nan、'Values'文本)统一转为NaN,方便批量归类到not_number
  • 分箱边界设置为负无穷到10、10到32、32到正无穷,完全匹配三个数值区间的划分规则
  • 最后用reindex调整统计结果的展示顺序,和预期输出顺序保持一致

运行上述代码将输出和预期完全一致的统计结果:

not_number        2
10 and below      2
11-32             9
33 and up         1

内容的提问来源于stack exchange,提问作者Mke

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.03 11:24:37