You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何更简便地通过统计单元格值计算Pandas共现矩阵?

Clean Solution for Co-Occurrence Matrix in Pandas

Hey there! Your loop-based approach gets the job done, but we can leverage pandas' powerful vectorized operations to get the exact co-occurrence matrix you want in one concise line—no loops required. Let's walk through it step by step.

First, recap your data and goal

Your original 0-1 DataFrame:

import pandas as pd
df = pd.DataFrame({'a' : [1,1,0,0], 'b': [0,1,1,0], 'c': [0,0,1,1]})

You need a matrix where each entry [i,j] counts how many times column j has a 1 when column i has a 1.

The One-Line Vectorized Solution

This is exactly what a matrix dot product of the DataFrame with its transpose achieves:

co_occurrence_matrix = df.T @ df

Alternatively, using the explicit dot method (same result):

co_occurrence_matrix = df.transpose().dot(df)

Verify the Result

Running this code gives you the exact matrix you wanted:

a  b  c
a  2  1  0
b  1  2  1
c  0  1  2

Why This Works

Let's break down the logic:

  • When you compute df.T @ df, the value at position (i,j) is the dot product of column i and column j from the original DataFrame.
  • For your 0-1 data, the dot product counts the number of rows where both column i and column j are 1. This perfectly matches your rule: it sums up all the 1s in column j for every row where column i is 1.

This method isn't just shorter—it's also far faster for larger datasets, since pandas uses optimized vectorized operations instead of slow Python-level loops.

How It Compares to Your Original Code

Your loop approach is functional, but this vectorized version is more readable, easier to maintain, and efficient. No need to juggle iteritems, groupby, or lambda functions here!

内容的提问来源于stack exchange,提问作者Edward

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:05:35