You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Polars中实现分组拼接字符串(对比Pandas groupby.sum)

在Polars中实现分组字符串拼接的方法

Pandas里可以通过groupby.sum()直接完成分组字符串拼接,但Polars的sum()是专门针对数值类型的聚合函数,对字符串列使用会返回null,因此需要用Polars提供的字符串专属聚合方法来实现需求。

方法1:使用str.concat()(推荐)

这是Polars原生的字符串聚合方法,性能更优,还支持自定义分隔符:

import polars as pl

df = pl.DataFrame({'a': [1, 1, 2], 'b': ['foo', 'bar', 'foo']})

# 无分隔符直接拼接
result = df.group_by('a').agg(pl.col('b').str.concat())
print(result)

输出结果:

shape: (2, 2)
┌─────┬────────┐
│ a   ┆ b      │
│ --- ┆ ---    │
│ i64 ┆ str    │
╞═════╪════════╡
│ 2   ┆ foo    │
│ 1   ┆ foobar │
└─────┴────────┘

如果需要添加分隔符(比如逗号),只需指定delimiter参数:

# 指定分隔符拼接
result_with_delimiter = df.group_by('a').agg(pl.col('b').str.concat(delimiter=','))
print(result_with_delimiter)

输出结果:

shape: (2, 2)
┌─────┬──────────┐
│ a   ┆ b        │
│ --- ┆ ---      │
│ i64 ┆ str      │
╞═════╪══════════╡
│ 2   ┆ foo      │
│ 1   ┆ foo,bar  │
└─────┴──────────┘

方法2:结合agg与Python的join

如果需要更灵活的自定义逻辑,也可以用lambda函数配合join实现:

result = df.group_by('a').agg(pl.col('b').agg(lambda x: ''.join(x)))
print(result)

该方法的输出和无分隔符的str.concat()一致,但性能略逊于原生方法,适合复杂拼接场景。

内容的提问来源于stack exchange,提问作者ignoring_gravity

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 12:10:31