You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Polars原生方法将不同嵌套列表列转为字符串?

问题:Polars高效处理不同层级嵌套列表的拼接需求

我正在从Pandas迁移到Polars,现在有一个包含不同嵌套层级列表的Polars DataFrame,结构如下:

┌────────────────────────────────────┬────────────────────────────────────┬─────────────────┬──────┐
│ col1                               ┆ col2                               ┆ col3            ┆ col4 │
│ ---                                ┆ ---                                ┆ ---             ┆ ---  │
│ list[list[str]]                    ┆ list[list[str]]                    ┆ list[str]       ┆ str  │
╞════════════════════════════════════╪════════════════════════════════════╪═════════════════╪══════╡
│ [["a", "a"], ["b", "b"], ["c", "c"]┆ [["a", "a"], ["b", "b"], ["c", "c"]┆ ["A", "B", "C"] ┆ 1    │
│ [["a", "a"]]                       ┆ [["a", "a"]]                       ┆ ["A"]           ┆ 2    │
│ [["b", "b"], ["c", "c"]]           ┆ [["b", "b"], ["c", "c"]]           ┆ ["B", "C"]      ┆ 3    │
└────────────────────────────────────┴────────────────────────────────────┴─────────────────┴──────┘

我需要从内到外使用不同分隔符拼接列表,最终得到如下结果:

┌─────────────┬─────────────┬───────┬──────┐
│ col1        ┆ col2        ┆ col3  ┆ col4 │
│ ---         ┆ ---         ┆ ---   ┆ ---  │
│ str         ┆ str         ┆ str   ┆ str  │
╞═════════════╪═════════════╪═══════╪══════╡
│ a+a-b+b-c+c ┆ a+a-b+b-c+c ┆ A-B-C ┆ 1    │
│ a+a         ┆ a+a         ┆ A     ┆ 2    │
│ b+b-c+c     ┆ b+b-c+c     ┆ B-C   ┆ 3    │
└─────────────┴─────────────┴───────┴──────┘

目前我用map_elements加for循环实现,但效率很低,想知道有没有Polars原生方法可以实现。我的现有代码如下:

import polars as pl

df = pl.DataFrame({"col1": [[["a", "a"], ["b", "b"], ["c", "c"]], [["a", "a"]], [["b", "b"], ["c", "c"]]],
                   "col2": [[["a", "a"], ["b", "b"], ["c", "c"]], [["a", "a"]], [["b", "b"], ["c", "c"]]],
                   "col3": [["A", "B", "C"], ["A"], ["B", "C"]],
                   "col4": ["1", "2", "3"]})

nested_list_cols = ["col1", "col2"]
list_cols = ["col3"]

for col in nested_list_cols:
    df = df.with_columns(pl.lit(df[col].map_elements(lambda listed: ['+'.join(element) for element in listed], return_dtype=pl.List(pl.String))).alias(col)) # is the return_dtype always pl.List(pl.String)?
for col in list_cols + nested_list_cols:
    df = df.with_columns(pl.lit(df[col].list.join(separator='-')).alias(col))

优化方案:用Polars原生向量化操作替代低效循环

当然有!Polars提供了原生的列表操作API,完全可以替代map_elements这种Python层面的循环,效率提升非常明显,尤其是处理大数据量时。

具体实现代码

import polars as pl

df = pl.DataFrame({
    "col1": [[["a", "a"], ["b", "b"], ["c", "c"]], [["a", "a"]], [["b", "b"], ["c", "c"]]],
    "col2": [[["a", "a"], ["b", "b"], ["c", "c"]], [["a", "a"]], [["b", "b"], ["c", "c"]]],
    "col3": [["A", "B", "C"], ["A"], ["B", "C"]],
    "col4": ["1", "2", "3"]
})

result_df = df.with_columns(
    # 处理双层嵌套的col1、col2:先内层用+拼接,再外层用-拼接
    pl.col("col1", "col2")
        .list.eval(pl.element().list.join("+"))
        .list.join("-"),
    # 处理单层嵌套的col3:直接用-拼接
    pl.col("col3").list.join("-")
)

print(result_df)

代码详解

  1. list.eval(pl.element().list.join("+")):
    • 针对col1、col2这类双层嵌套列表,list.eval会遍历外层列表的每个元素(也就是内层的小列表)
    • 对每个内层小列表执行list.join("+"),把列表里的字符串用+拼接成单个字符串,最终得到一个单层的字符串列表
  2. .list.join("-"):把第一步生成的单层字符串列表,用-拼接成最终的完整字符串
  3. 单层列表处理:col3是单层列表,直接调用list.join("-")就能完成拼接
  4. 向量化优势:所有操作都是Polars内部的向量化处理,完全避免了Python循环,性能比原实现高几个量级

内容的提问来源于stack exchange,提问作者gernophil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 15:28:10