You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何遍历pandas多列中的列表元素并完成zip配对与计数统计

实现方案

方案1:聚合阶段直接计算(性能最优,无冗余中间存储)

直接在groupby分组时完成配对统计,不需要先把两列转为列表存储,适合数据量较大的场景:

def count_pairs(group):
    # 对当前时间分组内的两行直接配对统计
    pairs = zip(group['DocumentTypeDE'], group['DocumentEventDE'])
    vc = pd.Series(pairs).value_counts().reset_index()
    vc.columns = ['Tuple', 'Count']
    vc['Mean'] = vc['Count'] / vc['Count'].sum()
    # 如需直接显示百分比可以替换为下一行
    # vc['Mean'] = vc['Count'].div(vc['Count'].sum()).apply(lambda x: f"{x:.2%}")
    return vc

# 直接得到你需要的长表结果
result = df.groupby('time_bin', as_index=False).apply(count_pairs).reset_index(drop=True)

方案2:基于已生成的timebin_df处理(符合你期望的zip+列表推导式写法)

如果你不想重新执行聚合逻辑,沿用你已经生成的存储了列表列的timebin_df,用一行列表推导式即可完成统计:

def get_stat_item(type_list, event_list):
    pairs = zip(type_list, event_list)
    vc = pd.Series(pairs).value_counts().rename('Count').reset_index()
    vc['Mean'] = vc['Count'].div(vc['Count'].sum()).apply(lambda x: f"{x:.2%}")
    return vc.to_dict('records')

# 完全符合你期望的zip遍历写法
timebin_df['stat_items'] = [get_stat_item(x, y) for x, y in zip(timebin_df['DocumentTypeDE'], timebin_df['DocumentEventDE'])]

# 展开得到最终结果
result = timebin_df.explode('stat_items').reset_index(drop=True)
result = pd.concat([result[['time_bin']], pd.json_normalize(result['stat_items'])], axis=1)

输出格式调整

如果需要和你给出的示例一样,重复时间箱名称合并显示,把对应列设为索引即可:

result = result.set_index(['time_bin', 'Tuple'])

内容的提问来源于stack exchange,提问作者nhoang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 07:45:03