R语言filter函数递归索引失败求助:动物追踪数据时间过滤
问题
处理动物追踪数据集时,需要筛选特定时间段,简化测试代码如下:
k=1 i=1 Intrusion5_animaldata_mv %>% filter(t %in% time_windows_utc2[[i]][[k]])
执行后触发错误:
Error in
filter():
ℹ In argument:t %in% time_windows_utc2[[i]][[k]].
Caused by error intime_windows_utc2[[i]]:
! recursive indexing failed at level 3
数据集结构
- Intrusion5_animaldata_mv 片段:
structure(list(t = structure(c(1505965920, 1505965980, 1505966040, 1505966100, 1505966160), tzone = "Etc/GMT-2", class = c("POSIXct", "POSIXt")), i = c(1L, 1L, 1L, 1L, 1L), x = c(26570L, 26569L, 26566L, 26565L, 26563L), y = c(3747L, 3745L, 3742L, 3740L, 3737L ), temp = c(1850L, 1850L, 1860L, 1860L, 1860L), acs.mean.mean = c(0L, 0L, 0L, 0L, 0L), acs.mean.peak = c(0L, 0L, 0L, 0L, 0L), acs.mean.sd = c(0L, 0L, 0L, 0L, 0L), acs.sd.mean = c(0L, 0L, 0L, 0L, 0L), acs.sd.peak = c(0L, 0L, 0L, 0L, 0L), acs.sd.sd = c(0L, 0L, 0L, 0L, 0L), acs.fracmove = c(0L, 0L, 0L, 0L, 0L), x_fix = c(590457, 590456.9, 590456.6, 590456.5, 590456.3), y_fix = c(7319174.7, 7319174.5, 7319174.2, 7319174, 7319173.7)), row.names = c(NA, -5L), class = c("tbl_df", "tbl", "data.frame"))
- time_windows_utc2 是嵌套列表:外层含8个列表,每个子列表包含6个列表,最内层是62个值的POSIXct向量,结构示例:
List of 8 $ :List of 6 ..$ : POSIXct[1:62], format: "2017-09-21 05:52:00" "2017-09-21 05:51:00" "2017-09-21 05:50:00" "2017-09-21 05:49:00" ... ..$ : POSIXct[1:62], format: "2017-09-21 06:12:00" "2017-09-21 06:11:00" "2017-09-21 06:10:00" "2017-09-21 06:09:00" ... ..$ : POSIXct[1:62], format: "2017-09-21 06:32:00" "2017-09-21 06:31:00" "2017-09-21 06:30:00" "2017-09-21 06:29:00" ... ..$ : POSIXct[1:62], format: "2017-09-21 06:52:00" "2017-09-21 06:51:00" "2017-09-21 06:50:00" "2017-09-21 06:49:00" ... ..$ : POSIXct[1:62], format: "2017-09-21 07:12:00" "2017-09-21 07:11:00" "2017-09-21 07:10:00" "2017-09-21 07:09:00" ... ..$ : POSIXct[1:62], format: "2017-09-21 07:32:00" "2017-09-21 07:31:00" "2017-09-21 07:30:00" "2017-09-21 07:29:00" ...
直接使用time_windows_utc2[[1]][[1]]能正常筛选,但用变量i、k索引(无论是否在for循环中)都会触发递归索引失败错误。
解决方案
原因分析
错误核心是dplyr的filter函数会优先将代码中的变量视为数据框内的列名,而非全局环境中的变量。你的数据框Intrusion5_animaldata_mv恰好有名为i的列,导致filter尝试用数据框的i列去索引time_windows_utc2,而非你定义的全局变量i,从而触发层级错误。
解决方法
方法1:用.env明确指定变量来源
通过.env告诉filter使用全局环境中的变量:
k=1 i=1 Intrusion5_animaldata_mv %>% filter(t %in% time_windows_utc2[[.env$i]][[.env$k]])
方法2:提前提取时间窗口
先把目标时间向量提取到全局变量中,再传入filter:
k=1 i=1 target_times <- time_windows_utc2[[i]][[k]] Intrusion5_animaldata_mv %>% filter(t %in% target_times)
这种方式更直观,彻底避免变量混淆问题。
方法3:重命名冲突列
如果数据框中的i列不是必须保留的名称,可以先重命名列:
k=1 i=1 Intrusion5_animaldata_mv %>% rename(animal_id = i) %>% # 将原i列重命名为animal_id filter(t %in% time_windows_utc2[[i]][[k]])
内容的提问来源于stack exchange,提问作者Thomas Osinga
相关产品推荐
相关产品推荐

