You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何对数值范围重编码?解决重编码值被覆盖问题

数值范围重编码问题解决方法

你的代码问题在于分步赋值时,最后一步的df2['col1'] >= -0.075会匹配到之前已经被改成1、2、3的整数(这些数值显然都大于-0.075),导致所有值最终被覆盖为4。以下是几种可行的解决方案:

方法1:嵌套np.where一次性赋值

把所有判断逻辑嵌套在一个np.where语句中,一次性完成重编码,避免分步覆盖:

import numpy as np

df2['col1'] = np.where(df2['col1'] < -1.27, 1,
                      np.where((df2['col1'] >= -1.27) & (df2['col1'] < -0.74), 2,
                               np.where((df2['col1'] >= -0.74) & (df2['col1'] < -0.075), 3, 4)))

逻辑是从左到右依次判断:满足第一个条件返回1,不满足则进入下一层判断,以此类推,最后所有不满足前面条件的返回4。

方法2:用pandas.cut分箱(最简洁)

pandas的cut函数专门用于数值分箱重编码,代码更易维护:

import pandas as pd

# 定义分箱边界,左闭右开,最后一个区间包含所有大于等于-0.075的数值
bins = [-float('inf'), -1.27, -0.74, -0.075, float('inf')]
# 对应每个区间的标签
labels = [1, 2, 3, 4]

# 执行分箱并映射标签,include_lowest确保最左区间包含最小值
df2['col1'] = pd.cut(df2['col1'], bins=bins, labels=labels, include_lowest=True)
# 如需转为整数类型,添加astype(int)
df2['col1'] = df2['col1'].astype(int)

方法3:用np.select多条件映射

通过条件列表和结果列表对应,清晰定义每个区间的映射规则:

import numpy as np

# 按顺序定义判断条件
conditions = [
    df2['col1'] < -1.27,
    (df2['col1'] >= -1.27) & (df2['col1'] < -0.74),
    (df2['col1'] >= -0.74) & (df2['col1'] < -0.075)
]
# 对应每个条件的结果
choices = [1, 2, 3]

# 匹配条件返回对应值,无匹配则返回默认值4
df2['col1'] = np.select(conditions, choices, default=4)

内容的提问来源于stack exchange,提问作者Log On

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 03:01:07