如何基于文档词汇共现矩阵绘制热力图?
This is a perfect use case for a heatmap—it’ll make those co-occurrence patterns way easier to spot at a glance. Let’s walk through how to adjust your code to build the matrix and plot it properly.
Step 1: Finish Building the Co-Occurrence Matrix
First, let’s fix and complete your code to generate a square matrix where each cell (i,j) represents how many documents contain both names[i] and names[j]. Here’s the corrected version:
import matplotlib.pyplot as plt import numpy as np # Load your documents a = (open('A.txt').read()).lower().split() c = (open('C.txt').read()).lower().split() f = (open('F.txt').read()).lower().split() documents = [a, f, c] # Replace with your actual list of important words names = ["word1", "word2", "word3", "word4"] # Example list # Initialize an empty co-occurrence matrix num_words = len(names) co_occurrence_matrix = np.zeros((num_words, num_words), dtype=int) # Calculate co-occurrences for i, word_a in enumerate(names): for j, word_b in enumerate(names): count = 0 for doc in documents: # Check if both words are present in the document if word_a in doc and word_b in doc: count += 1 co_occurrence_matrix[i][j] = count
Note: If you don’t want to count a word co-occurring with itself (the diagonal cells), add a condition if word_a != word_b inside the document loop. Keeping the diagonal can still be useful, though—it shows how often each word appears across all documents.
Step 2: Plot the Heatmap with Seaborn
Seaborn makes heatmaps straightforward and customizable. First install it if you haven’t (pip install seaborn), then add this code:
import seaborn as sns # Set up a readable figure size plt.figure(figsize=(10, 8)) # Create the heatmap sns.heatmap(co_occurrence_matrix, annot=True, # Show exact count values in each cell fmt='d', # Format annotations as integers cmap='YlGnBu', # Choose a color palette (adjust as needed) xticklabels=names, yticklabels=names) # Add context with labels and title plt.title('Document Word Co-Occurrence Heatmap', fontsize=14) plt.xlabel('Words', fontsize=12) plt.ylabel('Words', fontsize=12) # Show the final plot plt.show()
Quick Tweaks to Improve the Visualization
- Normalize Counts: If you want to account for how often each word appears individually (instead of raw counts), divide each cell by the total occurrences of either word (or their average) to get a normalized co-occurrence score.
- Adjust Color Palette: Try
coolwarm,viridis, orRdBuif you prefer a different color scheme. - Rotate Labels: For long word lists, add
plt.xticks(rotation=45)to prevent label overlap.
This will give you a clear heatmap where more intense colors indicate higher co-occurrence between pairs of words.
内容的提问来源于stack exchange,提问作者Oumnia Amal Rizki

