You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python Polars数据框中获取字符串列字符的ASCII码?

在Polars数据框中获取字符的ASCII码

你可以通过两种方式实现需求,优先推荐Polars原生方法以获得更好的性能:

方法1:使用Polars原生str.code_points()方法

Polars的str.code_points()可以直接获取字符的Unicode码点(对于ASCII字符来说,码点就是对应的ASCII值)。结合你的切片操作,代码如下:

import polars as pl
df = pl.DataFrame({'a' : ['apple', 'banana']})

# 获取第一个字符的ASCII码(结果为列表形式)
result = df.with_columns(
    pl.col('a').str.slice(0, 1).str.code_points().alias('first_char_ascii')
)
print(result)

输出:

shape: (2, 2)
┌────────┬──────────────────┐
│ a      ┆ first_char_ascii │
│ ---    ┆ ---              │
│ str    ┆ list[u32]        │
╞════════╪══════════════════╡
│ apple  ┆ [97]             │
│ banana ┆ [98]             │
└────────┴──────────────────┘

如果需要将结果转为单个数值,可添加list.first()提取列表中的唯一元素:

result = df.with_columns(
    pl.col('a').str.slice(0, 1).str.code_points().list.first().alias('first_char_ascii')
)
print(result)

输出:

shape: (2, 2)
┌────────┬──────────────────┐
│ a      ┆ first_char_ascii │
│ ---    ┆ ---              │
│ str    ┆ u32              │
╞════════╪══════════════════╡
│ apple  ┆ 97               │
│ banana ┆ 98               │
└────────┴──────────────────┘

方法2:通过map_elements调用Python的ord()函数

如果想直接使用Python内置的ord(),可以用map_elements对每个切片后的字符进行处理:

import polars as pl
df = pl.DataFrame({'a' : ['apple', 'banana']})

result = df.with_columns(
    pl.col('a').str.slice(0, 1).map_elements(lambda x: ord(x)).alias('first_char_ascii')
)
print(result)

这种方式的输出和方法1的单数值结果一致,但性能不如原生方法,处理大数据集时更推荐方法1。


内容的提问来源于stack exchange,提问作者lmocsi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 08:47:18