You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars是否有类似pandas.factorize的字符串列转整数编码功能?

Polars 实现类似 pandas.factorize 的字符串编码功能

Polars 完全支持将字符串列编码为连续整数的需求,以下是两种常用实现方式:

方法一:利用 Categorical 类型转换

将字符串列转为 Categorical 类型后,通过 to_physical() 获取其底层整数编码(默认从 0 开始),若需要从 1 起始则加 1 即可:

import polars as pl

# 示例数据
df = pl.DataFrame({"fruit": ["apple", "banana", "apple", "orange", "banana"]})

# 生成从 1 开始的整数编码
df = df.with_columns(
    (pl.col("fruit").cast(pl.Categorical).to_physical() + 1).alias("fruit_code")
)

print(df)

输出结果:

shape: (5, 2)
┌────────┬────────────┐
│ fruit  ┆ fruit_code │
│ ---    ┆ ---        │
│ str    ┆ u32        │
╞════════╪════════════╡
│ apple  ┆ 1          │
│ banana ┆ 2          │
│ apple  ┆ 1          │
│ orange ┆ 3          │
│ banana ┆ 2          │
└────────┴────────────┘

若需要查看类别与编码的对应关系,可以使用:

categories = df.get_column("fruit").cast(pl.Categorical).cat.get_categories()
# 对应关系:类别 -> 编码(+1 后的值)
print(dict(zip(categories, range(1, len(categories)+1))))
# 输出: {'apple': 1, 'banana': 2, 'orange': 3}

方法二:使用稠密排名(rank)直接生成

通过 rank("dense") 方法可以直接生成从 1 开始的连续整数编码,无需额外调整:

import polars as pl

df = pl.DataFrame({"fruit": ["apple", "banana", "apple", "orange", "banana"]})

df = df.with_columns(
    pl.col("fruit").rank("dense", descending=False).alias("fruit_code")
)

print(df)

输出结果与方法一完全一致。

内容的提问来源于stack exchange,提问作者Mark Wang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 00:45:32