You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中绘制含共有基因数据的2D着色矩阵

Solution: Integrate CSV Comparison & Fix Matplotlib Coloring for Gene Matrix

Let's break down how to fix your code, integrate the CSV comparison with the visualization, and make the matrix dynamically adapt to your gene lists:

Step 1: Fix the CSV Gene Comparison & Generate a Proper 2D Matrix

First, we'll modify your CSV reading code to:

  • Skip header rows (adjust this if your CSVs don't have a header row for gene names)
  • Collect comparison results into a structured 2D numpy array instead of just printing values
import csv
import numpy as np

# Read genes from first CSV (column 1 holds gene names)
with open('test1.csv','r') as input_file:
    reader = csv.reader(input_file)
    next(reader)  # Remove this line if your CSV has no header row
    row_str1 = [row[1] for row in reader]

# Read genes from second CSV
with open('test2.csv','r') as input_file:
    reader = csv.reader(input_file)
    next(reader)  # Remove this line if your CSV has no header row
    row_str2 = [row[1] for row in reader]

# Generate 2D matrix: 1 = shared gene, 0 = not shared
gene_matrix = np.zeros((len(row_str1), len(row_str2)), dtype=int)
for i, gene1 in enumerate(row_str1):
    for j, gene2 in enumerate(row_str2):
        if gene1 == gene2:
            gene_matrix[i][j] = 1

Step 2: Fix Matplotlib Coloring & Dynamic Adaptation

The main issues with your original visualization code were incorrect color bounds (you used [0,10,20] for 0/1 data) and hardcoded elements that didn't adapt to your actual gene list lengths. Here's the fixed, integrated visualization code:

import matplotlib.pyplot as plt
from matplotlib import colors

# Set up color map: 0 = red (non-shared), 1 = green (shared)
cmap = colors.ListedColormap(['red', 'green'])
# Define bounds to correctly map 0 and 1 to their respective colors
bounds = [-0.5, 0.5, 1.5]
norm = colors.BoundaryNorm(bounds, cmap.N)

# Create figure with dynamic size based on gene list lengths
fig, ax = plt.subplots(figsize=(len(row_str2)*0.8, len(row_str1)*0.8))
ax.imshow(gene_matrix, cmap=cmap, norm=norm)

# Add grid lines between matrix cells
ax.grid(which='major', axis='both', linestyle='-', color='k', linewidth=2)

# Set dynamic ticks matching the length of your gene lists
ax.set_xticks(np.arange(len(row_str2)))
ax.set_yticks(np.arange(len(row_str1)))

# Optional: Uncomment to label ticks with gene names
# ax.set_xticklabels(row_str2, rotation=90)
# ax.set_yticklabels(row_str1)

# Adjust layout to prevent label cutoff
plt.tight_layout()
plt.show()

Key Fixes Explained

  • Structured Matrix: We now create a 2D array that matches the dimensions of your two gene lists, ensuring the visualization aligns with your actual data
  • Correct Color Mapping: The bounds are set to [-0.5, 0.5, 1.5] so values of 0 fall into the red range and values of 1 fall into the green range
  • Dynamic Adaptation: All elements (figure size, ticks, matrix dimensions) are tied to the length of your gene lists from the CSVs, so it works regardless of how many genes you're comparing
  • Header Handling: Added next(reader) to skip header rows—remove this line if your CSVs don't have a header at the top

内容的提问来源于stack exchange,提问作者Vinay Bharadhwaj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:08:09