在R语言中基于两个DataFrame创建关联图谱(Contact Map)的方法
Alright, let's tackle how to create a contact map from your two data frames in R. First, I noticed your df2 is incomplete, so I'll start by filling in a reasonable structure with gene start/end positions (since we need genomic intervals to link motifs to genes). Here's a step-by-step approach:
First, let's finalize df2 with genomic intervals for each gene (this is necessary to map your motif positions to genes):
# Complete df2 with gene start/end positions df2 <- structure( list( geneID = c("E1", "E2", "E3", "E4", "E5"), chromosome = c("chr1", "chr1", "chr2", "chr2", "chr2"), start = c(500L, 1500L, 1000L, 2500L, 11000L), end = c(1500L, 2500L, 3500L, 4000L, 13000L) ), .Names = c("geneID", "chromosome", "start", "end"), class = "data.frame", row.names = c(NA,-5L)) # Your original df1 df1 <- structure( list( sample_id = c(1L, 1L, 1L, 1L, 2L, 2L), motif = c("CT-G.A", "TA-C.C", "TC-G.C", "TC-G.C", "CG-A.T", "CA-G.T"), chromosome = c("chr1", "chr1", "chr2", "chr2", "chr2", "chr2"), position = c(7300L, 1000L, 1200L, 3000L, 12000L, 2000L) ), .Names = c("sample_id", "motif", "chromosome", "position"), class = "data.frame", row.names = c(NA,-6L))
We'll use the GenomicRanges package to efficiently match motif positions to gene intervals (this is far more reliable than manual range checks):
library(dplyr) library(GenomicRanges) # Convert df1 to GenomicRanges (for motif positions) gr_motifs <- GRanges( seqnames = df1$chromosome, ranges = IRanges(start = df1$position, end = df1$position), sample_id = df1$sample_id, motif = df1$motif ) # Convert df2 to GenomicRanges (for gene intervals) gr_genes <- GRanges( seqnames = df2$chromosome, ranges = IRanges(start = df2$start, end = df2$end), geneID = df2$geneID ) # Find overlapping motifs and genes, then create a mapping dataframe motif_gene_links <- findOverlaps(gr_motifs, gr_genes) %>% {cbind(as.data.frame(gr_motifs[queryHits(.)]), as.data.frame(gr_genes[subjectHits(.)]))} %>% select(sample_id, motif, geneID, chromosome = seqnames, motif_position = start)
Now we'll create count matrices for the interactions you want to visualize. Two common options are:
Option A: Motif-to-Gene Contact Counts
# Count how many times each motif links to each gene per sample motif_gene_counts <- motif_gene_links %>% group_by(sample_id, geneID, motif) %>% summarise(contact_count = n(), .groups = "drop")
Option B: Gene-to-Gene Contact Counts
If you want to show interactions between genes (e.g., motifs that bridge different genes in the same sample/chromosome):
# Pair genes linked by motifs in the same sample/chromosome gene_gene_counts <- motif_gene_links %>% group_by(sample_id, chromosome) %>% mutate(partner_gene = lead(geneID)) %>% # Pair each gene with the next linked gene filter(!is.na(partner_gene)) %>% group_by(sample_id, geneID, partner_gene) %>% summarise(contact_count = n(), .groups = "drop")
We'll use ggplot2 to create heatmaps (the standard for contact maps).
Motif-to-Gene Contact Map
library(ggplot2) ggplot(motif_gene_counts, aes(x = geneID, y = motif, fill = contact_count)) + geom_tile(color = "white", size = 0.5) + facet_wrap(~sample_id, ncol = 1) + scale_fill_viridis_c(name = "Interaction\nCount") + labs(title = "Motif-Gene Contact Map", x = "Gene ID", y = "Motif Type") + theme_minimal() + theme(axis.text.x = element_text(angle = 45, hjust = 1))
Gene-to-Gene Contact Map (Hi-C Style)
ggplot(gene_gene_counts, aes(x = geneID, y = partner_gene, fill = contact_count)) + geom_tile(color = "white", size = 0.5) + facet_wrap(~sample_id, ncol = 1) + scale_fill_viridis_c(name = "Interaction\nCount") + labs(title = "Gene-Gene Contact Map", x = "Gene ID", y = "Partner Gene ID") + theme_minimal() + theme(axis.text.x = element_text(angle = 45, hjust = 1))
If you need a full Hi-C style contact map with genomic coordinates (not just gene IDs), use packages like HiCPlotR or plotgardener. These tools handle genome-wide interaction data and can plot maps aligned to chromosome coordinates.
内容的提问来源于stack exchange,提问作者H. dashti

