如何为DataFrame的path列不同行子集应用不同时间缩放因子且不拆分数据?
合并CSV后生成实验时间戳的优化方案
问题背景
将多个CSV文件合并为一个DataFrame后,生成的path列对应每个CSV的ID,需要通过数学转换将其转为实验时间戳。合并后的combined_files tibble包含14列,path列为第一列。
数据规则:path1-7(对应前336行)与path8-100之间存在0.5小时的间隔,path1-7转换规则为(path-1)*0.5,path8-100则在此基础上加0.5小时偏移。
原代码错误分析
尝试直接对行子集做转换时出现语法错误,原代码片段:
files <- list.files(pattern = "*.csv") files <- stringr::str_sort(files, numeric = TRUE) combined_files <- readr::read_csv(files, id = "path") %>% mutate(path = str_remove(path, "cap_data_"))%>% mutate(path = str_remove(path, ".csv")) combined_files$path <- as.numeric(combined_files$path) # 错误的子集操作 (combined_files[combined_files[c(1:336), 1]] - 1)*0.5
报错内容:
Error in `vectbl_as_col_location()`: ! Can't subset columns with `combined_files[c(1:336), 1]`. x `combined_files[c(1:336), 1]` must be logical, numeric, or character, not a <tbl_df/tbl/data.frame> object.
错误原因:combined_files[c(1:336), 1]返回的是一个tibble对象,而数据框子集索引需要的是逻辑向量、数值向量或字符向量,因此语法不合法。
现有临时解决方案
通过拆分数据框再合并的方式实现转换,代码如下:
combined_files_1st <- subset(combined_files, combined_files$path < 7.5 ) combined_files_2nd <- subset(combined_files, combined_files$path > 7) combined_files_1st$path <- (combined_files_1st$path - 1)*0.5 combined_files_2nd$path <- (combined_files_2nd$path-1)*0.5 +0.5 combined_join <- rbind(combined_files_1st, combined_files_2nd)
无需拆分的优化方案
使用dplyr的case_when或ifelse函数,直接在原数据框中完成条件转换,无需拆分合并:
方案1:用case_when实现(逻辑清晰,易扩展)
library(dplyr) library(stringr) library(readr) # 合并CSV并一步完成path清洗与时间戳转换 combined_files <- read_csv(files, id = "path") %>% mutate( # 先清洗path列,转为数值型 path = str_remove(path, "cap_data_") %>% str_remove(".csv") %>% as.numeric(), # 按规则转换为时间戳,直接替换原path列 path = case_when( path <= 7 ~ (path - 1)*0.5, path >= 8 ~ (path - 1)*0.5 + 0.5, TRUE ~ NA_real_ # 处理异常值 ) )
方案2:用ifelse简化(逻辑简单时更简洁)
combined_files <- combined_files %>% mutate(path = ifelse(path <=7, (path-1)*0.5, (path-1)*0.5 +0.5))
示例验证
针对提供的示例数据:
| path | p1 | p2 |
|---|---|---|
| 1 | 10 | 20 |
| 1 | 20 | 30 |
| 1 | 30 | 54 |
| 2 | 10 | 20 |
| 2 | 20 | 30 |
| 2 | 30 | 54 |
| 3 | 10 | 20 |
| 3 | 20 | 30 |
| 3 | 30 | 54 |
按照示例规则(path1-2转换为(path-1)*0.5,path3转换为(path-1)*0.5+1),用case_when实现:
sample_df <- tibble( path = c(1,1,1,2,2,2,3,3,3), p1 = c(10,20,30,10,20,30,10,20,30), p2 = c(20,30,54,20,30,54,20,30,54) ) sample_df <- sample_df %>% mutate(path = case_when( path <=2 ~ (path-1)*0.5, path ==3 ~ (path-1)*0.5 +1, TRUE ~ NA_real_ ))
转换后结果符合预期:
| path | p1 | p2 |
|---|---|---|
| 0 | 10 | 20 |
| 0 | 20 | 30 |
| 0 | 30 | 54 |
| 0.5 | 10 | 20 |
| 0.5 | 20 | 30 |
| 0.5 | 30 | 54 |
| 2 | 10 | 20 |
| 2 | 20 | 30 |
| 2 | 30 | 54 |
内容的提问来源于stack exchange,提问作者Kieran
相关产品推荐
相关产品推荐

