You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars中如何去除字符串列的字符重音?

在Polars中去除文本列的字符重音

需求:去除文本列中的字符重音(例如将Piña转换为Pina),已知Pandas中的实现代码如下:

(names
 .str.normalize('NFKD')
 .str.encode('ascii', errors='ignore')
 .str.decode('utf-8'))

针对Polars的实现,有两种可行方案:

方案一:使用字符串编码解码(修正用法)

Polars的str.encode和str.decode可实现相同逻辑,只需按正确参数调用:

import polars as pl

df = pl.DataFrame({"names": ["Piña", "Café", "São Paulo"]})
df = df.with_columns(
    pl.col("names")
    .str.normalize("NFKD")
    .str.encode("ascii", errors="ignore")
    .str.decode("utf-8")
    .alias("names_no_accents")
)

方案二:正则替换非ASCII字符

通过str.normalize分解字符后,用正则替换掉所有非ASCII范围内的字符,同样能达到去重音效果:

df = df.with_columns(
    pl.col("names")
    .str.normalize("NFKD")
    .str.replace_all(r"[^\x00-\x7F]", "")
    .alias("names_no_accents")
)

两种方案最终都能得到无重音的文本结果,可根据个人习惯选择。

内容的提问来源于stack exchange,提问作者Bingbong Recto

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 19:12:41