如何用R语言提取累计距离每增0.01km对应的GPS时间?
提取累计距离每增加10米对应的GPS时间点
问题描述
我有一个包含time(对应样本中的taken_at)、longitude(lng)、latitude(lat)的CSV文件,需要提取**累计距离每增加10米(0.01km)**对应的时间。目前已通过以下代码计算出每行相对于起点的累计距离:
gps <- read.csv("SP1ST1.csv") gps_sp <- SpatialPoints(cbind(gps$lng,gps$lat)) test <- spDistsN1(gps_sp, gps_sp[1,], longlat=TRUE)
得到的累计距离结果如下:
[1] 0.000000000 0.001586483 0.004574098 0.004493954 0.004887035 0.005405389 0.005930999 0.006443206 0.006991742 0.007595466 0.009693191 [12] 0.010654023 0.010231435 0.010082614 0.012005496 0.012905777 0.013896484 0.014873557 0.015857558 0.016905208 0.013991941 0.017441699 [23] 0.017797154 0.018539821 0.019254225 0.019914940 0.020634398 0.021411878 0.022246358 0.023037314 0.023832587 0.024608449 0.023977990
观察发现第一次累计距离突破0.01km在第1-11行之间,第二次突破0.02km在第11-26行之间。需要自动识别所有这类阈值点(距离增量并非恰好0.01km,且数据存在回退波动),并关联原始GPS数据提取对应时间。
补充的样本数据:
# 样本数据dput结果 structure(list(filename = c("20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV", "20230718_GSL_SP1ST1_4k_01.MOV" ), taken_at = c("14:11:05", "14:11:08", "14:11:12", "14:11:13", "14:11:14", "14:11:15", "14:11:16", "14:11:17", "14:11:18", "14:11:19", "14:11:20", "14:11:21", "14:11:22", "14:11:23", "14:11:24", "14:11:25", "14:11:26", "14:11:27", "14:11:28", "14:11:29", "14:11:30", "14:11:31", "14:11:32", "14:11:33", "14:11:34", "14:11:35", "14:11:36", "14:11:37", "14:11:38", "14:11:39"), lng = c(-65.36897, -65.36898, -65.36899, -65.36898, -65.36897, -65.36897, -65.36896, -65.36896, -65.36895, -65.36895, -65.369, -65.36899, -65.36901, -65.36903, -65.36899, -65.36898, -65.36897, -65.36896, -65.36895, -65.36894, -65.36899, -65.36901, -65.36901, -65.369, -65.36899, -65.36898, -65.36897, -65.36896, -65.36901, -65.36901), lat = c(49.95216, 49.95218, 49.9522, 49.9522, 49.95221, 49.95221, 49.95222, 49.95222, 49.95222, 49.95223, 49.95225, 49.95226, 49.95225, 49.95224, 49.95227, 49.95228, 49.95229, 49.9523, 49.9523, 49.95231, 49.95229, 49.95232, 49.95232, 49.95233, 49.95234, 49.95234, 49.95235, 49.95236, 49.95236, 49.95237 ), gps_altitude = c(-31.625, -31.373, -31.254, -31.604, -31.419, -31.432, -31.445, -31.459, -31.472, -31.485, -31.328, -31.322, -31.462, -31.614, -31.272, -31.189, -31.102, -31.015, -30.927, -30.838, -32.265, -31.533, -31.781, -31.921, -32.056, -32.188, -32.32, -32.452, -31.729, -31.705)), class = "data.frame", row.names = c(NA, -30L))
实现步骤与代码
1. 整合累计距离到原始数据
先把计算好的累计距离test添加到gps数据框,方便后续关联:
gps$cumulative_dist <- test
2. 定义目标距离阈值
根据需求生成0.01km、0.02km...的阈值序列,最大阈值不超过累计距离的最大值:
max_dist <- max(gps$cumulative_dist) thresholds <- seq(0.01, max_dist, by = 0.01)
3. 匹配每个阈值对应的第一个有效点
针对数据存在距离回退的情况,取第一次超过阈值的行,后续回退不重复计数:
result_list <- list() for (thresh in thresholds) { candidate_rows <- which(gps$cumulative_dist >= thresh) if (length(candidate_rows) > 0) { target_row <- candidate_rows[1] result_list[[as.character(thresh)]] <- data.frame( threshold_km = thresh, taken_at = gps$taken_at[target_row], actual_dist_km = gps$cumulative_dist[target_row], row_index = target_row ) } } # 转换为结构化数据框 result_df <- do.call(rbind, result_list)
4. 忽略回退波动的优化方案(可选)
如果只需要追踪距离的持续前进,可先计算累计距离的峰值序列,再匹配阈值:
# 生成只保留历史最大值的峰值序列 gps$peak_dist <- cummax(gps$cumulative_dist) result_peak_list <- list() for (thresh in thresholds) { candidate_rows <- which(gps$peak_dist >= thresh) if (length(candidate_rows) > 0) { target_row <- candidate_rows[1] result_peak_list[[as.character(thresh)]] <- data.frame( threshold_km = thresh, taken_at = gps$taken_at[target_row], peak_dist_km = gps$peak_dist[target_row], row_index = target_row ) } } result_peak_df <- do.call(rbind, result_peak_list)
5. 查看结果
运行后result_df或result_peak_df即为所需的时间点数据,针对样本数据的前两行结果示例:
threshold_km taken_at actual_dist_km row_index 0.01 0.01 14:11:20 0.01065402 12 0.02 0.02 14:11:36 0.02063440 27
内容的提问来源于stack exchange,提问作者Emily
相关产品推荐
相关产品推荐

