You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按指定顺序将DataFrame的locale列转为有序分类变量?

解决有序分类变量自定义编码问题

问题背景

我有一个X_train DataFrame,其中locale列的唯一值为['Regional', 'Local', 'National'],想要将该列转换为有序分类变量,指定顺序为Local=0、Regional=1、National=2。但当前用factorize的实现没效果,预期所有值为National时输出2,也可以尝试支持自定义顺序的LabelEncoder(如果有这个功能的话)。

原代码

print(X_train['locale'][:10])
cat = pd.Categorical(X_train['locale'], categories = ['Local', 'Regional', 'National'])
codes, uniques = pd.factorize(cat)
print(codes[:10])

样本数据(X_train前5行)

{'id': {0: 0, 1: 1, 2: 2, 3: 3, 4: 4},
 'date': {0: Timestamp('2013-01-01 00:00:00'),
  1: Timestamp('2013-01-01 00:00:00'),
  2: Timestamp('2013-01-01 00:00:00'),
  3: Timestamp('2013-01-01 00:00:00'),
  4: Timestamp('2013-01-01 00:00:00')},
 'store_nbr': {0: '1', 1: '1', 2: '1', 3: '1', 4: '1'},
 'family': {0: 'AUTOMOTIVE',
  1: 'BABY CARE',
  2: 'BEAUTY',
  3: 'BEVERAGES',
  4: 'BOOKS'},
 'sales': {0: 0.0, 1: 0.0, 2: 0.0, 3: 0.0, 4: 0.0},
 'onpromotion': {0: 0, 1: 0, 2: 0, 3: 0, 4: 0},
 'city': {0: 'Quito', 1: 'Quito', 2: 'Quito', 3: 'Quito', 4: 'Quito'},
 'state': {0: 'Pichincha',
  1: 'Pichincha',
  2: 'Pichincha',
  3: 'Pichincha',
  4: 'Pichincha'},
 'store_type': {0: 'D', 1: 'D', 2: 'D', 3: 'D', 4: 'D'},
 'cluster': {0: '13', 1: '13', 2: '13', 3: '13', 4: '13'},
 'dcoilwtico': {0: nan, 1: nan, 2: nan, 3: nan, 4: nan},
 'transactions': {0: nan, 1: nan, 2: nan, 3: nan, 4: nan},
 'holiday_type': {0: 'Holiday',
  1: 'Holiday',
  2: 'Holiday',
  3: 'Holiday',
  4: 'Holiday'},
 'locale': {0: 'National',
  1: 'National',
  2: 'National',
  3: 'National',
  4: 'National'},
 'locale_name': {0: 'Ecuador',
  1: 'Ecuador',
  2: 'Ecuador',
  3: 'Ecuador',
  4: 'Ecuador'},
 'description': {0: 'Primer dia del ano',
  1: 'Primer dia del ano',
  2: 'Primer dia del ano',
  3: 'Primer dia del ano',
  4: 'Primer dia del ano'},
 'transferred': {0: False, 1: False, 2: False, 3: False, 4: False},
 'year': {0: '2013', 1: '2013', 2: '2013', 3: '2013', 4: '2013'},
 'month': {0: '1', 1: '1', 2: '1', 3: '1', 4: '1'},
 'week': {0: '1', 1: '1', 2: '1', 3: '1', 4: '1'},
 'quarter': {0: '1', 1: '1', 2: '1', 3: '1', 4: '1'},
 'day_of_week': {0: 'Tuesday',
  1: 'Tuesday',
  2: 'Tuesday',
  3: 'Tuesday',
  4: 'Tuesday'}}

问题原因

pd.factorize处理Categorical对象时,是按数据中出现的顺序编码,而非你定义的分类顺序,所以达不到预期效果。


解决方案

方案1:直接用Categorical的codes属性(最推荐)

创建有序分类时标记ordered=True,直接取内置编码即可:

# 创建有序分类,指定自定义顺序
cat = pd.Categorical(X_train['locale'], categories=['Local', 'Regional', 'National'], ordered=True)
# 获取对应编码
codes = cat.codes
print(codes[:10])  # 输出全为2,符合预期

方案2:手动映射(简单直观)

用字典定义映射关系,通过map方法转换:

locale_map = {'Local': 0, 'Regional': 1, 'National': 2}
codes = X_train['locale'].map(locale_map)
print(codes[:10])  # 输出全为2

方案3:用sklearn工具处理

方法A:OrdinalEncoder(专门处理有序分类)
from sklearn.preprocessing import OrdinalEncoder

# 指定自定义分类顺序
encoder = OrdinalEncoder(categories=[['Local', 'Regional', 'National']])
# 注意输入需为二维数组,转换后展平为一维
codes = encoder.fit_transform(X_train[['locale']]).flatten()
print(codes[:10])  # 输出全为2
方法B:LabelEncoder配合有序分类

sklearn的LabelEncoder本身不支持自定义顺序,可先转成有序分类再编码:

from sklearn.preprocessing import LabelEncoder

# 先将列转为有序分类
X_train['locale_cat'] = pd.Categorical(X_train['locale'], categories=['Local', 'Regional', 'National'], ordered=True)
# 用LabelEncoder编码
le = LabelEncoder()
codes = le.fit_transform(X_train['locale_cat'])
print(codes[:10])  # 输出全为2

内容的提问来源于stack exchange,提问作者Katsu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 16:06:28