You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用含大小于条件的字典映射Pandas DataFrame列?

问题:能否用区间描述字典通过map方法生成Pandas新列?

我创建了如下Pandas DataFrame:

ds = {'col1':[1,2,2,3,4,5,5,6,7,8]}
df = pd.DataFrame(data=ds)

该DataFrame内容如下:

col1
0     1
1     2
2     2
3     3
4     4
5     5
6     5
7     6
8     7
9     8

我通过自定义函数并使用apply方法生成了新列newCol:

def criteria(row):
    if((row['col1'] > 0) & (row['col1'] <= 2)):
        return "A"
    elif((row['col1'] > 2) & (row['col1'] <= 3)):
        return "B"
    else:
        return "C"
    
df['newCol'] = df.apply(criteria, axis=1)

生成后的DataFrame如下:

col1 newCol
0     1      A
1     2      A
2     2      A
3     3      B
4     4      C
5     5      C
6     5      C
7     6      C
8     7      C
9     8      C

我能否创建如下格式的字典:

dict = {
        '0 <= 2' : "A",
        '2 <= 3' : "B",
        'Else' : "C"
        }

并通过以下方式应用到DataFrame:

df['newCol'] = df['col1'].map(dict)

请问是否可行?


解答

这种方式不可行,原因如下:

  • map()方法的核心逻辑是将Series中的每个值与字典的键做精确匹配,你定义的字典键是字符串形式的区间描述(如'0 <= 2'),而col1中的值是整数类型,两者完全无法匹配,最终newCol列会全部填充为NaN。

推荐替代方案:使用pd.cut()

pd.cut()是Pandas专门为数值区间分箱设计的函数,效率远高于apply,完全契合你的需求:

import pandas as pd

# 定义区间边界和对应标签
bins = [-float('inf'), 2, 3, float('inf')]
labels = ['A', 'B', 'C']

# 生成新列
df['newCol'] = pd.cut(df['col1'], bins=bins, labels=labels, include_lowest=True)
  • bins参数定义区间:(-∞,2]对应A,(2,3]对应B,(3,∞)对应C
  • include_lowest=True确保左边界的数值(如2)被正确划分到第一个区间

备选方案:离散值字典映射(仅适用于有限离散值场景)

如果一定要用字典映射,需要将col1中每个可能的离散值作为字典的键,剩余值用fillna填充:

mapping_dict = {1: 'A', 2: 'A', 3: 'B'}
df['newCol'] = df['col1'].map(mapping_dict).fillna('C')

内容的提问来源于stack exchange,提问作者Giampaolo Levorato

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 19:46:11