R语言新手求助:如何基于两个文件绘制指定基因行的分组箱线图
Hey there! I totally get that manually transposing files is a huge hassle—let's fix that with a streamlined, code-only workflow using the tidyverse suite of packages. This will let you go straight from your original test.txt and design.txt files to your desired boxplot without any manual data editing.
Step 1: Install & Load Required Packages
First, make sure you have the tidyverse package installed (it includes ggplot2 for plotting, dplyr for data wrangling, and tidyr for reshaping data):
# Install if you haven't already install.packages("tidyverse") # Load the package into your R session library(tidyverse)
Step 2: Read & Clean Your Input Files
Your design.txt is in a wide, paired format (sample name followed by its group), so we'll first reshape it into a clean two-column table:
# Read the gene expression data expr_data <- read.delim("test.txt", sep = "\t", header = TRUE, row.names = "gene") # Read the design file (note: it doesn't have a proper header, so we use header=FALSE) design_raw <- read.delim("design.txt", sep = "\t", header = FALSE) # Reshape the design data into a sample-to-group mapping design_clean <- tibble( sample = design_raw[1, seq(2, ncol(design_raw), by = 2)], # Grab every 2nd column starting at 2 (sample names) group = design_raw[1, seq(3, ncol(design_raw), by = 2)] # Grab every 2nd column starting at 3 (group labels) )
Step 3: Reshape Expression Data & Merge with Group Info
Next, we'll convert your wide expression data into long format, then combine it with the group information from the design file:
# Turn wide expression data into long format (gene, sample, expression value) expr_long <- expr_data %>% rownames_to_column("gene") %>% # Move row names (gene IDs) into a proper column pivot_longer(cols = -gene, names_to = "sample", values_to = "expression") # Merge with design data to add group labels to each sample's expression value combined_data <- expr_long %>% inner_join(design_clean, by = "sample")
Step 4: Create the Boxplot for Gene a1
Now we can filter for your target gene and build the boxplot:
# Filter the combined data to only include gene a1 a1_data <- combined_data %>% filter(gene == "a1") # Generate the boxplot ggplot(a1_data, aes(x = group, y = expression)) + stat_boxplot(geom = "errorbar", width = 0.15) + # Add error bars first (behind the boxplot) geom_boxplot(fill = "skyblue", alpha = 0.7) + # Boxplot with a light fill labs( title = "Expression of Gene a1", x = "Experimental Group", y = "Expression Level" ) + theme_minimal() # Optional: use a clean, modern theme
Step 5: Save the Plot
To save the plot with your desired filename:
ggsave("boxplot-a1.png", width = 6, height = 4, dpi = 300)
This workflow is fully reproducible—if you want to plot a different gene later, just change filter(gene == "a1") to the gene ID you need (like "a2" or "a3"). No more manual file editing!
内容的提问来源于stack exchange,提问作者Peter Trinh

