You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过分箱将DataFrame的float列转为categorical列(cost列问题)

问题解决:DataFrame的cost列分箱转换错误修正

你在处理cost列时的核心问题是:给pd.cut传入了字符串类型的分箱规则(['=>1000', '<1000']),但pd.cut需要基于数值边界划分区间,无法直接识别字符串描述的规则。以下是修正后的实现方案:

修正后的完整代码

import pandas as pd

def PreprocessDataframe(df):
    # numbers列分箱:合并冗余边界,确保区间连续且覆盖目标范围
    df['numbers'] = pd.cut(
        df['numbers'],
        bins=[17, 24, 44, 65, 92],  # 左边界设为17,配合include_lowest确保18被纳入第一个区间
        labels=['18-24', '25-44', '45-65', '66-92'],
        include_lowest=True
    )
    
    # cost列分箱:基于数值边界实现>=1000和<1000的划分
    df['cost'] = pd.cut(
        df['cost'],
        bins=[-float('inf'), 1000, float('inf')],  # 用-inf/inf覆盖所有可能的数值
        labels=['<1000', '>=1000'],
        include_lowest=True
    )
    return df

# 测试示例DataFrame
sample_df = pd.DataFrame({
    'cost': [360.0, 120.0, 2000.0],
    'numbers': [23.0, 35.0, 49.0]
})

processed_df = PreprocessDataframe(sample_df)
print(processed_df)

关键细节说明

  • numbers列优化:原代码的bins=[18,24,25,44,45,65,66,92]会产生24-25的空区间,合并相邻边界后更简洁,同时保证所有目标数值都能落入对应区间。
  • cost列修正逻辑:
    • 用[-float('inf'), 1000, float('inf')]作为数值边界,覆盖所有可能的cost取值(包括极小/极大值)
    • 通过labels参数直接映射为你需要的分类文本
    • include_lowest=True确保等于1000的值被划入>=1000区间(若需将1000归到<1000,可将bins调整为[-inf, 999.999, inf])

运行输出结果

cost  numbers
0    <1000    18-24
1    <1000    25-44
2  >=1000    45-65

内容的提问来源于stack exchange,提问作者Deepak M

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 19:09:32