You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python基于pandas DataFrame构建参与者-主办方共现频次矩阵

基于Pandas生成参与者-主办方共现频次矩阵的实现方法

我们有如下结构的pandas DataFrame:

import pandas as pd
df = pd.DataFrame({'participant_id' : [1608, 1608, 2089, 213, 1608, 1887, 2089, 4544, 6866, 2020, 2020],
               'organizer_id' : [1772, 1772, 1772, 1790, 1790, 1790, 1791, 1791, 1772, 1799, 1799]})

打印df的输出如下:

participant_id  organizer_id
0             1608          1772
1             1608          1772
2             2089          1772
3              213          1790
4             1608          1790
5             1887          1790
6             2089          1791
7             4544          1791
8             6866          1772
9             2020          1799
10            2020          1799

我们的目标是统计每位参与者参与各主办方任务的次数,生成指定格式的共现频次矩阵。


实现方法

方法1:使用crosstab(最简便)

pandas内置的crosstab函数专门用于统计分组频次,一行代码即可完成需求:

freq_matrix = pd.crosstab(index=df['participant_id'], columns=df['organizer_id'])

如果需要匹配示例中的浮点数输出格式,增加类型转换即可:

freq_matrix = freq_matrix.astype(float)

输出结果完全符合要求:

organizer_id    1772  1790  1791  1799
participant_id                        
1608             2.0   1.0   0.0   0.0
1887             0.0   1.0   0.0   0.0
2020             0.0   0.0   0.0   2.0
2089             1.0   0.0   1.0   0.0
213              0.0   1.0   0.0   0.0
4544             0.0   0.0   1.0   0.0
6866             1.0   0.0   0.0   0.0

如果需要去掉索引名和列名,调用freq_matrix.rename_axis(columns=None).rename_axis(index=None)即可。

方法2:使用groupby + unstack

如果习惯用分组聚合逻辑实现,也可以用如下写法:

freq_matrix = df.groupby(['participant_id', 'organizer_id']).size().unstack(fill_value=0).astype(float)

输出结果和crosstab方法完全一致。


内容的提问来源于stack exchange,提问作者Squid Game

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 01:45:03