Polars中条件链式操作的规范写法咨询
优化Polars条件式文本处理代码的建议
你的代码功能逻辑是正确的,能完成小写转换、按语言移除停用词、统一文本清洗这三个核心需求。但从Polars的惯用写法、性能和简洁性来看,有不少可以优化的地方,具体如下:
主要优化点
- 避免重复计算语言检测结果:原代码在两个
when分支里重复调用map_elements(lang_identifier),会导致同一文本被检测两次,浪费计算资源。建议先把语言检测结果存储为单独列,后续直接复用。 - 修复正则转义错误:原代码中使用
r'\\b'会匹配字面意义的\b而非单词边界,应该改为r'\b'才能正确匹配停用词的单词边界。 - 合并
with_columns调用:Polars支持链式调用,多次with_columns可以合并为一次,让代码更紧凑。 - 预编译正则表达式:提前编译停用词正则,避免每次
str.replace_all时重复编译,提升处理大数据集的性能。 - 封装重复逻辑:把Unicode归一化这类独立操作封装成函数,让代码更易读和维护。
优化后的完整代码
import polars as pl from unicodedata import normalize from lingua import Language, LanguageDetectorBuilder from nltk.corpus import stopwords import re # 初始化语言检测器 languages = [Language.ENGLISH, Language.FRENCH] detector = LanguageDetectorBuilder.from_languages(*languages).build() # 测试数据集 test = pl.DataFrame( { "id": [1, 2, 3, 4], "NaicsDescription": ['Full service restuarants', 'The Manufacturing of toys and trains', 'POWER GENERATING STATIONS', 'the short term rental of cottages'] } ) def lang_identifier(text: str): if isinstance(text, str): language = detector.detect_language_of(text) return language.iso_code_639_1.name if language else None return None # 预编译停用词正则 en_stopwords = re.compile(r'\b(?:' + '|'.join(set(w.lower() for w in stopwords.words('english'))) + r')\b') fr_stopwords = re.compile(r'\b(?:' + '|'.join(set(w.lower() for w in stopwords.words('french'))) + r')\b') bilingual_stopwords = re.compile(r'\b(?:' + '|'.join(set(w.lower() for w in stopwords.words('english') + stopwords.words('french'))) + r')\b') # 封装Unicode归一化函数 def normalize_text(text: str): return normalize('NFKD', text).encode('ascii', errors='ignore').decode('utf-8') # 链式处理,合并所有操作 test = ( test # 1. 小写转换 + 语言检测 .with_columns( pl.col('NaicsDescription').str.to_lowercase().alias('NaicsDescription_'), pl.col('NaicsDescription').map_elements(lang_identifier, return_dtype=pl.String).alias('lang') ) # 2. 按语言移除停用词 .with_columns( pl.when(pl.col('lang') == 'FR') .then(pl.col('NaicsDescription_').str.replace_all(fr_stopwords, ' ')) .when(pl.col('lang') == 'EN') .then(pl.col('NaicsDescription_').str.replace_all(en_stopwords, ' ')) .otherwise(pl.col('NaicsDescription_').str.replace_all(bilingual_stopwords, ' ')) .alias('NaicsDescription_') ) # 3. 统一文本清洗 .with_columns( pl.col('NaicsDescription_') .map_elements(normalize_text, return_dtype=pl.String) .str.replace_all(r'(?:[^\s\w]|_\d)+', ' ') .str.replace_all(r'\b(?:\d+|\w{1,2})\b', ' ') .str.replace_all(r'\s\s+', ' ') .str.strip_chars() .replace('', None) .alias('NaicsDescription_cleaned') ) # 可选:移除中间列 .drop('NaicsDescription_', 'lang') ) print(test)
额外说明
- 若处理超大规模数据集,可将
map_elements替换为Polars的注册UDF(pl.register_udf),利用向量化处理能力进一步提升性能。 - 停用词集合可提前缓存(比如本地存储下载好的nltk停用词文件),避免每次运行重复加载。
内容的提问来源于stack exchange,提问作者Jeff M
相关产品推荐
相关产品推荐

