在R语言中按固定分辨率对时间序列数据求平均
色谱CSV文件批量降分辨率处理方案
核心思路
无需手动创建时间区间虚拟表,通过数值计算自动生成分组标签,再按分组对RID求平均,实现快速降分辨率。核心逻辑是用floor(time / 分辨率) * 分辨率将每个时间点映射到对应区间的起始值,直接完成分组。
单文件快速处理
方法1:Base R实现(无需额外包)
# 设置目标分辨率 resolution <- 0.05 # 读取CSV文件(替换为你的文件路径) df <- read.csv("your_chromatogram.csv") # 生成区间起始标签 df$time_group <- floor(df$time / resolution) * resolution # 按分组聚合求RID平均值 downsampled_df <- aggregate(RID ~ time_group, data = df, FUN = mean) # 重命名列名匹配需求 colnames(downsampled_df) <- c("time", "RID") # 查看结果 print(downsampled_df)
方法2:dplyr实现(代码更简洁易读)
library(dplyr) resolution <- 0.05 downsampled_df <- read.csv("your_chromatogram.csv") %>% # 生成区间起始作为新的time列 mutate(time = floor(time / resolution) * resolution) %>% # 按time分组 group_by(time) %>% # 计算RID平均值(自动跳过NA值) summarise(RID = mean(RID, na.rm = TRUE)) %>% # 取消分组 ungroup() print(downsampled_df)
批量处理整个文件夹
以下代码可自动遍历指定文件夹内所有CSV文件,完成降分辨率后保存到新文件夹:
library(dplyr) # 配置参数 resolution <- 0.05 input_dir <- "./your_input_folder" # 替换为你的输入文件夹路径 output_dir <- "./downsampled_results" # 输出结果的文件夹 # 创建输出文件夹(不存在则自动创建) if (!dir.exists(output_dir)) { dir.create(output_dir) } # 定义单个文件的处理函数 process_single_csv <- function(file_path) { # 读取文件 raw_data <- read.csv(file_path, stringsAsFactors = FALSE) # 降分辨率处理 processed_data <- raw_data %>% mutate(time = floor(time / resolution) * resolution) %>% group_by(time) %>% summarise(RID = mean(RID, na.rm = TRUE)) %>% ungroup() # 生成输出文件路径 file_name <- basename(file_path) output_path <- file.path(output_dir, file_name) # 保存结果 write.csv(processed_data, output_path, row.names = FALSE) return(processed_data) } # 获取文件夹内所有CSV文件的完整路径 all_csv_files <- list.files(input_dir, pattern = "\\.csv$", full.names = TRUE) # 批量处理所有文件 lapply(all_csv_files, process_single_csv)
注意事项
- 确保所有输入CSV文件包含
time和RID列,若列名不同需修改代码中对应的列名。 - 如果数据存在缺失值,
mean()中的na.rm = TRUE会自动跳过,避免计算错误。 - 可根据需求调整
resolution参数(如0.1、0.2等),适配不同的降分辨率需求。
内容的提问来源于stack exchange,提问作者Raffaello
相关产品推荐
相关产品推荐

