如何为ggplot的geom_tile图层添加含元数据的列?
Great question! To add metadata columns to your geom_tile plot, you’ll need to integrate your metadata into your dataset, reshape the data for ggplot compatibility, and then customize the plot to distinguish metadata from your genetic markers. Here’s a complete, reproducible example:
Step 1: Combine Your Genetic Data with Metadata
First, create your metadata variables (e.g., sample type, antibiotic resistance status) and merge them with your existing genetic data into a single data frame. I’ll use realistic metadata examples below:
# Your original genetic data set.seed(123) # For consistent random samples id <- 1:80 gyrA <- sample(c(1,0), 80, replace = TRUE) parC <- sample(c(1,0), 80, replace = TRUE) marR <- sample(c(1,0), 80, replace = TRUE) qnrS <- sample(c(1,0), 80, replace = TRUE) marA <- sample(c(1,0), 80, replace = TRUE) ydhE <- sample(c(1,0), 80, replace = TRUE) qnrA <- sample(c(1,0), 80, replace = TRUE) qnrB <- sample(c(1,0), 80, replace = TRUE) qnrD <- sample(c(1,0), 80, replace = TRUE) mcbE <- sample(c(1,0), 80, replace = TRUE) # Add example metadata sample_type <- sample(c("Clinical", "Environmental"), 80, replace = TRUE) abx_resistance <- sample(c("Resistant", "Susceptible"), 80, replace = TRUE) # Combine into one data frame df <- data.frame(id, sample_type, abx_resistance, gyrA, parC, marR, qnrS, marA, ydhE, qnrA, qnrB, qnrD, mcbE)
Step 2: Reshape Data to Long Format
geom_tile works best with long-format data (each row represents one sample-variable pair). We’ll use pivot_longer from the tidyr package to reshape the data, and add a column to categorize variables as either "Metadata" or "Gene":
library(tidyr) library(dplyr) df_long <- df %>% pivot_longer(cols = -id, # Keep sample ID as a separate column names_to = "Variable", values_to = "Value") %>% # Tag variables as metadata or genetic markers mutate(Var_Category = ifelse(Variable %in% c("sample_type", "abx_resistance"), "Metadata", "Gene"))
Step 3: Create the Tile Plot with Metadata Columns
Now we can build the plot. We’ll place samples on the y-axis, all variables (metadata + genes) on the x-axis, and style metadata differently to make it stand out:
library(ggplot2) ggplot(df_long, aes(x = Variable, y = factor(id), fill = Value)) + # Add tile layer with white borders for clarity geom_tile(color = "white") + # Add a thick vertical line to separate metadata from genes geom_vline(xintercept = which(unique(df_long$Variable) == "abx_resistance") + 0.5, color = "black", size = 1.2) + # Customize fill colors: distinct colors for metadata categories and gene presence/absence scale_fill_manual(values = c( "0" = "#ffffff", "1" = "#0072B2", "Clinical" = "#E69F00", "Environmental" = "#56B4E9", "Resistant" = "#D55E00", "Susceptible" = "#009E73" )) + # Clean up labels and theme labs(x = "", y = "Sample ID", fill = "Status") + theme_minimal() + theme( axis.text.x = element_text(angle = 45, hjust = 1, size = 10), axis.text.y = element_text(size = 8), panel.grid = element_blank(), legend.position = "bottom" )
Key Notes:
- Customization: Adjust the metadata variables, fill colors, and theme to match your actual data and preferences.
- Categorical Metadata: If your metadata is numerical, you can use a continuous fill scale instead of manual values.
- Variable Order: Use
factor(Variable, levels = c(...))to control the order of columns in the plot (e.g., place metadata first).
内容的提问来源于stack exchange,提问作者Haakonkas

