如何在R中导出stat_peaks识别的峰位置至数据框?
Absolutely! You can pull the peak data (sample ID, x position, and intensity) into a clean dataframe exactly like you want—ggspectra has built-in tools to make this straightforward, no need to hack around plot objects (though that's an option too). Let's break this down:
Option 1: Directly Extract Peaks with find_peaks() (Recommended)
The stat_peaks() function you're using relies on ggspectra::find_peaks() under the hood, which can return peak data directly without first creating a plot. This is the most reliable method, as you can match the exact peak-detection parameters you use for visualization.
Step 1: Prep your data
First, make sure your data is in long format with columns for:
ID: Sample identifierx: Spectral axis (wavelength, wavenumber, etc.)y: Intensity value
Here's a dummy example matching your structure (replace with your real data):
library(tibble) set.seed(123) # For reproducible dummy peaks spectra_data <- tibble( ID = rep(c("sample1", "sample2"), each = 500), x = rep(seq(200, 500, length.out = 500), 2), y = c( # Peaks for sample1 at ~209 and ~467 dnorm(seq(200,500, length.out=500), mean=209, sd=10)*10000 + dnorm(seq(200,500, length.out=500), mean=467, sd=15)*20000 + rnorm(500, 0, 50), # Peaks for sample2 at ~213 and ~468 dnorm(seq(200,500, length.out=500), mean=213, sd=10)*8000 + dnorm(seq(200,500, length.out=500), mean=468, sd=15)*12000 + rnorm(500, 0, 50) ) )
Step 2: Extract peaks per sample
Use dplyr to group your data by sample ID, then apply find_peaks() to each group. Adjust parameters like threshold (relative peak height) and span (smoothing window) to match what you used in stat_peaks:
library(ggspectra) library(dplyr) library(tidyr) peaks_df <- spectra_data %>% group_by(ID) %>% # Apply find_peaks to each sample's spectral data summarize( peaks = list(find_peaks( x = x, y = y, threshold = 0.01, # Peaks must be at least 1% of max intensity span = 5 # Window size for peak detection (adjust as needed) )), .groups = "drop" ) %>% # Unnest the list column into individual rows unnest(peaks) %>% # Rename columns to match your desired format rename(`y (intensity)` = y_peak) # View the result print(peaks_df)
This will output exactly the dataframe structure you want:
# A tibble: 4 × 3 ID x `y (intensity)` <chr> <dbl> <dbl> 1 sample1 209 1048. 2 sample1 467 2247. 3 sample2 213 849. 4 sample2 468 1298.
Option 2: Extract Peaks from an Existing ggplot Object
If you already have a plot created with stat_peaks(), you can pull the peak data directly from the plot layer:
# Example plot with stat_peaks p <- ggplot(spectra_data, aes(x = x, y = y)) + geom_line(aes(color = ID)) + stat_peaks(aes(label = x), span = 5, threshold = 0.01) # Extract peak data (adjust `i = 2` to match the layer index of stat_peaks) peaks_from_plot <- layer_data(p, i = 2) %>% select(ID = colour, x = x, `y (intensity)` = y)
Key Notes
- Match Parameters: Make sure the
threshold,span, and other peak-detection arguments infind_peaks()are identical to those in yourstat_peaks()call—this ensures the peaks you extract match what's plotted. - Baseline Correction: If your spectra have a noisy baseline, consider applying baseline correction first (e.g., with
ggspectra::baseline()or thebaselinepackage) to avoid detecting spurious peaks. - Customization:
find_peaks()has additional arguments likeignore_threshold(absolute intensity threshold) andstrict(strict peak definition) to fine-tune peak detection to your data.
内容的提问来源于stack exchange,提问作者Agnese

