在使用ChemoSpec包前,用R计算重复样品数据的均值
Hey there! No need to apologize—we all start somewhere 😊. Absolutely, you can automate this workflow in R, which will save you tons of time as your sample count grows. Below is a straightforward, scalable solution tailored to your setup:
Step 1: Install & Load Required Packages
We’ll use the tidyverse collection of packages (it makes reading files and manipulating data much more intuitive). If you haven’t installed it yet, run this first:
install.packages("tidyverse") library(tidyverse)
Step 2: Read All CSV Files
First, make sure all your 30 CSV files are in a single folder (you can set this folder as your working directory with setwd("path/to/your/folder")). Then we’ll read all files and add a column to track which original file each row came from:
# Get list of all CSV files in the working directory file_list <- list.files(pattern = "\\.csv$") # Read all files into a single data frame, adding a "file_id" column all_data <- map_dfr(file_list, read_csv, .id = "file_id") %>% # Convert file_id from character (like "1") to numeric mutate(file_id = as.numeric(file_id))
Step 3: Group Replicates & Calculate Mean Absorbance
Since your files are paired (1&2 = Sample 1, 3&4 = Sample 2, etc.), we’ll create a sample_id column to group each pair. Then we’ll calculate the mean absorbance for each wavelength within each sample:
mean_absorbance <- all_data %>% # Create sample_id: pair files into groups of 2 mutate(sample_id = ceiling(file_id / 2)) %>% # Group by sample and wavelength, then calculate mean absorbance group_by(sample_id, wavelength) %>% summarize(mean_absorbance = mean(absorbance), .groups = "drop")
Step 4 (Optional): Export the Results to CSV
If you want to save the averaged data for later use (or for ChemoSpec), you can export it to a CSV file:
write_csv(mean_absorbance, "averaged_ftir_data.csv")
How This Works with Your Example
Using your simplified csv1 and csv2 data:
wavelength <- c(500, 550, 600) absorbance <- c(2, 4, 3) csv1 <- data.frame(wavelength, absorbance) csv2 <- data.frame(wavelength, absorbance) # Simulate the file reading step all_data_example <- bind_rows( csv1 %>% mutate(file_id = 1), csv2 %>% mutate(file_id = 2) ) # Calculate mean mean_example <- all_data_example %>% mutate(sample_id = ceiling(file_id / 2)) %>% group_by(sample_id, wavelength) %>% summarize(mean_absorbance = mean(absorbance), .groups = "drop") print(mean_example)
This will output exactly the mean values you calculated manually:
# A tibble: 3 × 3 sample_id wavelength mean_absorbance <dbl> <dbl> <dbl> 1 1 500 2 2 1 550 4 3 1 600 3
For ChemoSpec Users
If you plan to use this data directly in ChemoSpec, you can convert the mean_absorbance data frame into the format ChemoSpec expects. For example, pivot the data to have wavelengths as rows and samples as columns:
chemo_spec_ready <- mean_absorbance %>% pivot_wider(names_from = sample_id, values_from = mean_absorbance) %>% column_to_rownames("wavelength")
This approach scales seamlessly—even if you end up with 100 samples (200 CSV files), you won’t need to change any code!
内容的提问来源于stack exchange,提问作者G Jones

