You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Polars中实现忽略变音符号排序或自定义排序?

Polars 自定义变音符号排序解决方案

问题背景

需要对Polars DataFrame的Lexeme列排序时,让带重音的字符(如á、é)和对应基础字符(a、e)排在一起,而非默认Unicode排序中基础字符全部排在重音字符之前。尝试过locale.setlocale未生效,因为Polars默认排序不依赖系统locale,而是按Unicode码点排序。

方案1:忽略所有变音符号排序

通过Unicode规范化移除重音标记,生成无重音的排序键,以此实现a/á、e/é等同序排列:

代码实现

import os
import polars as pl
import unicodedata

# 读取数据并填充空值(修正原代码未赋值的问题)
LexiconPath = r'C:\Users\mawan\Documents\Code\WorldBuildingCode\KalagyonManNyal_Lexicon.csv'
UnsortedLexiconPath = r'C:\Users\mawan\Documents\Code\WorldBuildingCode\KalagyonManNyal_Lexicon_Sorted.csv'

df = pl.read_csv(LexiconPath).fill_null("-")

# 方法1:用unicodedata手动移除重音(适合小数据集)
def remove_accents(s):
    # NFKD规范化拆分重音与基础字符,过滤掉非间距标记
    return unicodedata.normalize('NFKD', s).encode('ascii', 'ignore').decode('utf-8')

df_sorted = df.sort(pl.col("Lexeme").map_elements(remove_accents))

# 方法2:用Polars内置字符串函数(更高效,适合大数据集)
df_sorted = df.sort(pl.col("Lexeme").str.normalize("NFKD").str.replace_all(r'\p{Mn}', ''))

# 输出排序结果并保存
print(df_sorted)
df_sorted.write_csv(UnsortedLexiconPath, include_bom=True)

原理

NFKD会将带重音的字符分解为「基础字符+重音标记」,\p{Mn}匹配所有Unicode非间距重音标记,替换后得到无重音的字符串,以此为排序键就能让基础字符和对应重音字符排在一起。

方案2:自定义部分变音符号规则

如果需要更灵活的规则(比如让部分重音字符独立排序,部分和基础字符合并),可以自定义字符映射生成排序键:

代码实现

import os
import polars as pl
from functools import reduce
import operator

LexiconPath = r'C:\Users\mawan\Documents\Code\WorldBuildingCode\KalagyonManNyal_Lexicon.csv'
UnsortedLexiconPath = r'C:\Users\mawan\Documents\Code\WorldBuildingCode\KalagyonManNyal_Lexicon_Sorted.csv'

df = pl.read_csv(LexiconPath).fill_null("-")

# 自定义映射:键是原字符,值是排序时的替代字符
# 示例:让á/a、é/e等同序,若需要某字符独立(如ç),可设为'c1'使其排在c之后
sort_mapping = {
    'á': 'a',
    'é': 'e',
    'í': 'i',
    'ó': 'o',
    'ú': 'u',
    # 'ç': 'c1'
}

# 构建Polars表达式批量替换字符
replace_expr = reduce(
    operator.add,
    [pl.col("Lexeme").str.replace(c, replacement) for c, replacement in sort_mapping.items()]
)

# 基于替换后的列排序
df_sorted = df.sort(replace_expr)

# 输出排序结果并保存
print(df_sorted)
df_sorted.write_csv(UnsortedLexiconPath, include_bom=True)

原理

通过将指定重音字符替换为对应基础字符,排序时这些字符会被视为相同键,从而排在一起;未在映射中的字符保持原Unicode顺序。

注意事项

  1. 原代码中df.fill_null("-")未赋值给df,需修正为df = df.fill_null("-")才会生效。
  2. map_elements适合小数据集,大数据集建议用Polars内置字符串函数(如str.normalize、str.replace_all),性能更优。

内容的提问来源于stack exchange,提问作者mavmav0

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 01:11:01