You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars中条件链式操作的规范写法咨询

优化Polars条件式文本处理代码的建议

你的代码功能逻辑是正确的,能完成小写转换、按语言移除停用词、统一文本清洗这三个核心需求。但从Polars的惯用写法、性能和简洁性来看,有不少可以优化的地方,具体如下:

主要优化点

  • 避免重复计算语言检测结果:原代码在两个when分支里重复调用map_elements(lang_identifier),会导致同一文本被检测两次,浪费计算资源。建议先把语言检测结果存储为单独列,后续直接复用。
  • 修复正则转义错误:原代码中使用r'\\b'会匹配字面意义的\b而非单词边界,应该改为r'\b'才能正确匹配停用词的单词边界。
  • 合并with_columns调用:Polars支持链式调用,多次with_columns可以合并为一次,让代码更紧凑。
  • 预编译正则表达式:提前编译停用词正则,避免每次str.replace_all时重复编译,提升处理大数据集的性能。
  • 封装重复逻辑:把Unicode归一化这类独立操作封装成函数,让代码更易读和维护。

优化后的完整代码

import polars as pl
from unicodedata import normalize 
from lingua import Language, LanguageDetectorBuilder
from nltk.corpus import stopwords
import re

# 初始化语言检测器
languages = [Language.ENGLISH, Language.FRENCH]
detector = LanguageDetectorBuilder.from_languages(*languages).build()

# 测试数据集
test = pl.DataFrame(
    {
        "id": [1, 2, 3, 4],
        "NaicsDescription": ['Full service restuarants', 'The Manufacturing of toys and trains', 'POWER GENERATING STATIONS', 'the short term rental of cottages']
    }
)

def lang_identifier(text: str):
    if isinstance(text, str):
        language = detector.detect_language_of(text)
        return language.iso_code_639_1.name if language else None
    return None

# 预编译停用词正则
en_stopwords = re.compile(r'\b(?:' + '|'.join(set(w.lower() for w in stopwords.words('english'))) + r')\b')
fr_stopwords = re.compile(r'\b(?:' + '|'.join(set(w.lower() for w in stopwords.words('french'))) + r')\b')
bilingual_stopwords = re.compile(r'\b(?:' + '|'.join(set(w.lower() for w in stopwords.words('english') + stopwords.words('french'))) + r')\b')

# 封装Unicode归一化函数
def normalize_text(text: str):
    return normalize('NFKD', text).encode('ascii', errors='ignore').decode('utf-8')

# 链式处理,合并所有操作
test = (
    test
    # 1. 小写转换 + 语言检测
    .with_columns(
        pl.col('NaicsDescription').str.to_lowercase().alias('NaicsDescription_'),
        pl.col('NaicsDescription').map_elements(lang_identifier, return_dtype=pl.String).alias('lang')
    )
    # 2. 按语言移除停用词
    .with_columns(
        pl.when(pl.col('lang') == 'FR')
          .then(pl.col('NaicsDescription_').str.replace_all(fr_stopwords, ' '))
          .when(pl.col('lang') == 'EN')
          .then(pl.col('NaicsDescription_').str.replace_all(en_stopwords, ' '))
          .otherwise(pl.col('NaicsDescription_').str.replace_all(bilingual_stopwords, ' '))
          .alias('NaicsDescription_')
    )
    # 3. 统一文本清洗
    .with_columns(
        pl.col('NaicsDescription_')
          .map_elements(normalize_text, return_dtype=pl.String)
          .str.replace_all(r'(?:[^\s\w]|_\d)+', ' ')
          .str.replace_all(r'\b(?:\d+|\w{1,2})\b', ' ')
          .str.replace_all(r'\s\s+', ' ')
          .str.strip_chars()
          .replace('', None)
          .alias('NaicsDescription_cleaned')
    )
    # 可选:移除中间列
    .drop('NaicsDescription_', 'lang')
)

print(test)

额外说明

  • 若处理超大规模数据集,可将map_elements替换为Polars的注册UDF(pl.register_udf),利用向量化处理能力进一步提升性能。
  • 停用词集合可提前缓存(比如本地存储下载好的nltk停用词文件),避免每次运行重复加载。

内容的提问来源于stack exchange,提问作者Jeff M

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 21:25:58