如何在Python中绘制含共有基因数据的2D着色矩阵
Solution: Integrate CSV Comparison & Fix Matplotlib Coloring for Gene Matrix
Let's break down how to fix your code, integrate the CSV comparison with the visualization, and make the matrix dynamically adapt to your gene lists:
Step 1: Fix the CSV Gene Comparison & Generate a Proper 2D Matrix
First, we'll modify your CSV reading code to:
- Skip header rows (adjust this if your CSVs don't have a header row for gene names)
- Collect comparison results into a structured 2D numpy array instead of just printing values
import csv import numpy as np # Read genes from first CSV (column 1 holds gene names) with open('test1.csv','r') as input_file: reader = csv.reader(input_file) next(reader) # Remove this line if your CSV has no header row row_str1 = [row[1] for row in reader] # Read genes from second CSV with open('test2.csv','r') as input_file: reader = csv.reader(input_file) next(reader) # Remove this line if your CSV has no header row row_str2 = [row[1] for row in reader] # Generate 2D matrix: 1 = shared gene, 0 = not shared gene_matrix = np.zeros((len(row_str1), len(row_str2)), dtype=int) for i, gene1 in enumerate(row_str1): for j, gene2 in enumerate(row_str2): if gene1 == gene2: gene_matrix[i][j] = 1
Step 2: Fix Matplotlib Coloring & Dynamic Adaptation
The main issues with your original visualization code were incorrect color bounds (you used [0,10,20] for 0/1 data) and hardcoded elements that didn't adapt to your actual gene list lengths. Here's the fixed, integrated visualization code:
import matplotlib.pyplot as plt from matplotlib import colors # Set up color map: 0 = red (non-shared), 1 = green (shared) cmap = colors.ListedColormap(['red', 'green']) # Define bounds to correctly map 0 and 1 to their respective colors bounds = [-0.5, 0.5, 1.5] norm = colors.BoundaryNorm(bounds, cmap.N) # Create figure with dynamic size based on gene list lengths fig, ax = plt.subplots(figsize=(len(row_str2)*0.8, len(row_str1)*0.8)) ax.imshow(gene_matrix, cmap=cmap, norm=norm) # Add grid lines between matrix cells ax.grid(which='major', axis='both', linestyle='-', color='k', linewidth=2) # Set dynamic ticks matching the length of your gene lists ax.set_xticks(np.arange(len(row_str2))) ax.set_yticks(np.arange(len(row_str1))) # Optional: Uncomment to label ticks with gene names # ax.set_xticklabels(row_str2, rotation=90) # ax.set_yticklabels(row_str1) # Adjust layout to prevent label cutoff plt.tight_layout() plt.show()
Key Fixes Explained
- Structured Matrix: We now create a 2D array that matches the dimensions of your two gene lists, ensuring the visualization aligns with your actual data
- Correct Color Mapping: The
boundsare set to[-0.5, 0.5, 1.5]so values of 0 fall into the red range and values of 1 fall into the green range - Dynamic Adaptation: All elements (figure size, ticks, matrix dimensions) are tied to the length of your gene lists from the CSVs, so it works regardless of how many genes you're comparing
- Header Handling: Added
next(reader)to skip header rows—remove this line if your CSVs don't have a header at the top
内容的提问来源于stack exchange,提问作者Vinay Bharadhwaj
相关产品推荐
相关产品推荐

