如何更简便地通过统计单元格值计算Pandas共现矩阵?
Hey there! Your loop-based approach gets the job done, but we can leverage pandas' powerful vectorized operations to get the exact co-occurrence matrix you want in one concise line—no loops required. Let's walk through it step by step.
First, recap your data and goal
Your original 0-1 DataFrame:
import pandas as pd df = pd.DataFrame({'a' : [1,1,0,0], 'b': [0,1,1,0], 'c': [0,0,1,1]})
You need a matrix where each entry [i,j] counts how many times column j has a 1 when column i has a 1.
The One-Line Vectorized Solution
This is exactly what a matrix dot product of the DataFrame with its transpose achieves:
co_occurrence_matrix = df.T @ df
Alternatively, using the explicit dot method (same result):
co_occurrence_matrix = df.transpose().dot(df)
Verify the Result
Running this code gives you the exact matrix you wanted:
a b c a 2 1 0 b 1 2 1 c 0 1 2
Why This Works
Let's break down the logic:
- When you compute
df.T @ df, the value at position(i,j)is the dot product of columniand columnjfrom the original DataFrame. - For your 0-1 data, the dot product counts the number of rows where both column
iand columnjare 1. This perfectly matches your rule: it sums up all the 1s in columnjfor every row where columniis 1.
This method isn't just shorter—it's also far faster for larger datasets, since pandas uses optimized vectorized operations instead of slow Python-level loops.
How It Compares to Your Original Code
Your loop approach is functional, but this vectorized version is more readable, easier to maintain, and efficient. No need to juggle iteritems, groupby, or lambda functions here!
内容的提问来源于stack exchange,提问作者Edward

