You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效实现DataFrame中按id和hour分组拼接transaction_type?

更高效的DataFrame交易类型拼接实现方案

需求说明

现有包含id、hour、transaction_type三列的DataFrame,需将同一id、同一hour下的transaction_type值拼接为类似"A, B"的格式。

原实现代码

Hours=['24', '23', '22', '21', '20', '19', '18','17', '16', '15', '14', '13', '12', '11','10', '09', '08', '07', '06', '05', '04','03', '02', '01', '00']
result=[]
unique_id = df['id'].unique()

for i in unique_id:
    for j in Hours:
        filtered_data = df[(df['id'] == i) & (df['hour']==j)]
        if not filtered_data.empty:
            concatenated_names = ', '.join(filtered_data['transaction_type'])
            result.append(concatenated_names)

高效实现方式

利用Pandas内置的groupby+agg方法,通过向量化操作替代Python层面的嵌套循环,大幅提升效率:

基础版(不强制小时顺序)

# 按id和hour分组,拼接每组的transaction_type
grouped_result = df.groupby(['id', 'hour'])['transaction_type'].agg(', '.join)
# 转为与原代码格式一致的列表
result = grouped_result.tolist()

严格匹配Hours顺序版

如果需要输出结果严格遵循Hours列表中的小时顺序,可将hour列转为有序类别后再分组:

import pandas as pd

# 将hour列设置为有序类别,指定顺序为Hours列表
df['hour'] = pd.Categorical(df['hour'], categories=Hours, ordered=True)
# observed=True确保只保留数据中存在的类别组合
grouped_result = df.groupby(['id', 'hour'], observed=True)['transaction_type'].agg(', '.join)
result = grouped_result.tolist()

效率优势说明

原代码的嵌套循环需多次对DataFrame进行过滤切片,每次切片都会生成新的数据集,时间复杂度为O(n*m)(n为唯一id数量,m为小时数),大数据量下性能极差。

而groupby是Pandas底层优化的操作,基于C语言实现向量化处理,能一次性完成分组与聚合,时间复杂度远低于循环实现,处理十万级以上数据时速度提升可达数十倍。

内容的提问来源于stack exchange,提问作者Aditi Sahay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 06:33:29