大样本下基于dummy/binary变量的co-occurrence matrix及heat map制作咨询
Absolutely! Using dummy/binary variables is a perfect approach to calculate co-occurrence counts and build a heatmap—this method is both efficient and straightforward, even for your large dataset (20+ columns, 2000+ rows). Let’s break down the solution step by step, including how to transform your sample dataset.
When you convert your variables into binary (0/1) dummy columns, each row represents a sample, and each column indicates whether a specific variable is present (1) or absent (0) in that sample. The co-occurrence count between two variables is simply the number of rows where both columns are 1. We can compute this with matrix multiplication, then visualize the result as a heatmap.
Let’s use a sample dataset x to demonstrate. Suppose x has rows with multi-valued categories (common in co-occurrence scenarios):
import pandas as pd # Sample dataset x x = pd.DataFrame({ "sample_id": [1, 2, 3, 4, 5], "categories": ["A,B", "B,C", "A,C,D", "A", "B,D"] })
Convert the multi-valued column into binary dummy columns:
# Split comma-separated values and generate dummy variables dummy_df = x["categories"].str.get_dummies(sep=",") # Optional: Keep the sample ID for reference final_dummy_data = pd.concat([x["sample_id"], dummy_df], axis=1)
The resulting final_dummy_data will have columns like A, B, C, D with 0/1 values—exactly the format we need.
To get the count of how often each pair of variables appears together, multiply the transposed dummy matrix by the original dummy matrix:
# Compute co-occurrence count matrix co_occurrence_counts = dummy_df.T.dot(dummy_df) # Optional: Calculate normalized co-occurrence ratios (for relative frequency) total_occurrences = co_occurrence_counts.values.diagonal() co_occurrence_ratios = co_occurrence_counts / total_occurrences # Set diagonal to 0 if you want to ignore self-occurrences import numpy as np np.fill_diagonal(co_occurrence_ratios.values, 0)
co_occurrence_counts[i,j]: Number of times variableiandjappear togetherco_occurrence_ratios[i,j]: Ratio of timesiappears alongsidej(useful for comparing rare vs common variables)
Use seaborn to visualize the matrix—this works seamlessly for large datasets:
import seaborn as sns import matplotlib.pyplot as plt # Heatmap for raw counts plt.figure(figsize=(10, 8)) sns.heatmap(co_occurrence_counts, annot=True, fmt="d", cmap="Blues", cbar=True) plt.title("Variable Co-Occurrence Heatmap (Raw Counts)") plt.xlabel("Variables") plt.ylabel("Variables") plt.show() # Heatmap for normalized ratios plt.figure(figsize=(10, 8)) sns.heatmap(co_occurrence_ratios, annot=True, fmt=".2f", cmap="Greens", cbar=True) plt.title("Variable Co-Occurrence Heatmap (Normalized Ratios)") plt.xlabel("Variables") plt.ylabel("Variables") plt.show()
- Sparse Matrix Handling: If your dummy matrix is mostly 0s (sparse), use
scipy.sparseto save memory and speed up calculations:from scipy.sparse import csr_matrix sparse_dummy = csr_matrix(dummy_df.values) co_occurrence_sparse = sparse_dummy.T.dot(sparse_dummy) # Convert back to DataFrame for visualization co_occurrence_counts = pd.DataFrame(co_occurrence_sparse.todense(), index=dummy_df.columns, columns=dummy_df.columns) - Heatmap Readability: For 20+ columns, reduce annotation font size with
annot_kws={"size": 8}or disable annotations (annot=False) if they clutter the plot.
内容的提问来源于stack exchange,提问作者Tvdv

