You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars中类似Pandas用列表推导式+Lambda的更优实现方式?

更优雅的Polars文本分割与正则替换实现

原实现回顾

Pandas 代码

df['sentences'] = df['content'].str.split(pattern2)
df['normal_text'] = df['sentences'].apply(lambda x: [re.sub(pattern3, ' ', sentence) for sentence in x])

原Polars 代码

df = df.with_columns(pl.col('content').map_elements(lambda x: re.split(pattern2, x)).alias('sentences'))
df = df.with_columns(pl.col('sentences').map_elements(lambda x: [re.sub(pattern3, ' ', sentence) for sentence in x]).alias('normal_text'))

优化后的Polars 实现

当然有更优雅且高效的写法——Polars提供了原生的字符串和列表处理API,无需依赖map_elements和Python的re模块做逐元素循环:

# 方式1:一次完成所有列的计算
df = df.with_columns(
    sentences=pl.col('content').str.split(pattern2),
    normal_text=pl.col('content').str.split(pattern2).list.eval(
        pl.element().str.replace_all(pattern3, ' ')
    )
)

# 方式2:复用sentences列,避免重复分割计算(更高效)
df = df.with_columns(
    sentences=pl.col('content').str.split(pattern2)
).with_columns(
    normal_text=pl.col('sentences').list.eval(
        pl.element().str.replace_all(pattern3, ' ')
    )
)

优化点说明

  • 替换map_elements为原生str.split:和Pandas的写法对齐,直接调用Polars原生的字符串分割方法,代码更简洁,避免Python lambda的开销。
  • 用list.eval+str.replace_all处理列表替换:通过list.eval遍历列表内的每个元素,结合原生的str.replace_all完成正则替换,全程是Polars的向量化操作,性能远高于Python层面的循环。
  • 无需引入re模块:所有字符串操作都通过Polars原生API完成,减少外部依赖。

内容的提问来源于stack exchange,提问作者dh Lin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 21:01:11