在Pandas中将一列数值映射为另一列分类值的实现方法
给Pandas DataFrame的数值列添加分类列
根据你的需求,我们可以通过几种简洁的方式实现基于score列的数值范围新增names分类列,以下是两种常用方案:
方案一:使用pd.cut()(固定区间分类首选)
pd.cut()可以直接将数值列按指定区间划分成分类,完美匹配你的场景:
import pandas as pd # 构造原始DataFrame df = pd.DataFrame({'score': [16.0, 49.0, 55.0, 65.0, 77.0, 89.0, 98.0]}) # 定义区间边界和对应分类标签 bins = [-float('inf'), 55, 77, 89, float('inf')] labels = ['low', 'avg', 'high', 'very high'] # 新增分类列 df['names'] = pd.cut(df['score'], bins=bins, labels=labels, include_lowest=True) print(df)
输出结果:
| score | names |
|---|---|
| 16.0 | low |
| 49.0 | low |
| 55.0 | low |
| 65.0 | avg |
| 77.0 | avg |
| 89.0 | high |
| 98.0 | very high |
说明:include_lowest=True确保左区间包含边界值(比如55会被分到第一个区间)。
方案二:使用np.select()(复杂自定义条件首选)
如果你的分类规则需要更灵活的逻辑(比如非连续区间),可以用np.select()实现:
import pandas as pd import numpy as np df = pd.DataFrame({'score': [16.0, 49.0, 55.0, 65.0, 77.0, 89.0, 98.0]}) # 定义每个分类的判断条件和对应标签 conditions = [ df['score'] <= 55, (df['score'] > 55) & (df['score'] <= 77), (df['score'] > 77) & (df['score'] <= 89), df['score'] > 89 ] choices = ['low', 'avg', 'high', 'very high'] # 新增分类列 df['names'] = np.select(conditions, choices) print(df)
该方案会得到和上面完全一致的结果,优势是可以自由调整每个分类的判断逻辑,适配更复杂的场景。
内容的提问来源于stack exchange,提问作者User277883
相关产品推荐
相关产品推荐

