You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

大样本下基于dummy/binary变量的co-occurrence matrix及heat map制作咨询

Absolutely! Using dummy/binary variables is a perfect approach to calculate co-occurrence counts and build a heatmap—this method is both efficient and straightforward, even for your large dataset (20+ columns, 2000+ rows). Let’s break down the solution step by step, including how to transform your sample dataset.

Step 1: Understand the Core Logic

When you convert your variables into binary (0/1) dummy columns, each row represents a sample, and each column indicates whether a specific variable is present (1) or absent (0) in that sample. The co-occurrence count between two variables is simply the number of rows where both columns are 1. We can compute this with matrix multiplication, then visualize the result as a heatmap.

Step 2: Transform Your Dataset to Dummy Variables

Let’s use a sample dataset x to demonstrate. Suppose x has rows with multi-valued categories (common in co-occurrence scenarios):

import pandas as pd

# Sample dataset x
x = pd.DataFrame({
    "sample_id": [1, 2, 3, 4, 5],
    "categories": ["A,B", "B,C", "A,C,D", "A", "B,D"]
})

Convert the multi-valued column into binary dummy columns:

# Split comma-separated values and generate dummy variables
dummy_df = x["categories"].str.get_dummies(sep=",")

# Optional: Keep the sample ID for reference
final_dummy_data = pd.concat([x["sample_id"], dummy_df], axis=1)

The resulting final_dummy_data will have columns like A, B, C, D with 0/1 values—exactly the format we need.

Step 3: Calculate the Co-Occurrence Matrix

To get the count of how often each pair of variables appears together, multiply the transposed dummy matrix by the original dummy matrix:

# Compute co-occurrence count matrix
co_occurrence_counts = dummy_df.T.dot(dummy_df)

# Optional: Calculate normalized co-occurrence ratios (for relative frequency)
total_occurrences = co_occurrence_counts.values.diagonal()
co_occurrence_ratios = co_occurrence_counts / total_occurrences
# Set diagonal to 0 if you want to ignore self-occurrences
import numpy as np
np.fill_diagonal(co_occurrence_ratios.values, 0)
  • co_occurrence_counts[i,j]: Number of times variable i and j appear together
  • co_occurrence_ratios[i,j]: Ratio of times i appears alongside j (useful for comparing rare vs common variables)
Step 4: Generate the Heatmap

Use seaborn to visualize the matrix—this works seamlessly for large datasets:

import seaborn as sns
import matplotlib.pyplot as plt

# Heatmap for raw counts
plt.figure(figsize=(10, 8))
sns.heatmap(co_occurrence_counts, annot=True, fmt="d", cmap="Blues", cbar=True)
plt.title("Variable Co-Occurrence Heatmap (Raw Counts)")
plt.xlabel("Variables")
plt.ylabel("Variables")
plt.show()

# Heatmap for normalized ratios
plt.figure(figsize=(10, 8))
sns.heatmap(co_occurrence_ratios, annot=True, fmt=".2f", cmap="Greens", cbar=True)
plt.title("Variable Co-Occurrence Heatmap (Normalized Ratios)")
plt.xlabel("Variables")
plt.ylabel("Variables")
plt.show()
Optimization Tips for Large Datasets
  • Sparse Matrix Handling: If your dummy matrix is mostly 0s (sparse), use scipy.sparse to save memory and speed up calculations:
    from scipy.sparse import csr_matrix
    sparse_dummy = csr_matrix(dummy_df.values)
    co_occurrence_sparse = sparse_dummy.T.dot(sparse_dummy)
    # Convert back to DataFrame for visualization
    co_occurrence_counts = pd.DataFrame(co_occurrence_sparse.todense(), 
                                        index=dummy_df.columns, 
                                        columns=dummy_df.columns)
    
  • Heatmap Readability: For 20+ columns, reduce annotation font size with annot_kws={"size": 8} or disable annotations (annot=False) if they clutter the plot.

内容的提问来源于stack exchange,提问作者Tvdv

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:01:11