如何用Python datatable实现pandas分组后字符串拼接功能
实现方案
核心代码
from datatable import dt, f, by df = dt.Frame(group1=[1, 1, 1, 2, 2, 2], group2=[1, 1, 2, 2, 2, 3], text=['a', 'b', 'c', 'd', 'e', 'f']) # 直接用datatable原生语法实现分组拼接,无需依赖pandas df2 = df[:, dt.str.join(' ', f.text), by(f.group1, f.group2)]
输出的df2结构和示例中pandas计算结果完全一致:
| group1 group2 text | int32 int32 str32 -- + ------ ------ ----- 0 | 1 1 a b 1 | 1 2 c 2 | 2 2 d e 3 | 2 3 f [4 rows x 3 columns]
注:
dt.str.join方法需要 datatable 1.0.0 及以上版本支持,若版本过低可执行pip install --upgrade datatable升级。
逻辑说明
by(f.group1, f.group2)对应pandas的groupby(['group1', 'group2']),指定分组字段dt.str.join(' ', f.text)是datatable原生的字符串拼接聚合函数,第一个参数为分隔符,第二个参数为要拼接的列,等价于pandas里的apply(' '.join)- 全程在datatable内存模型中计算,没有转换pandas的额外开销,也不需要依赖pandas库
内容的提问来源于stack exchange,提问作者peter
相关产品推荐
相关产品推荐

