使用R spatstat识别GPS点最近邻并关联原数据集ID
问题描述
我有一个包含488个GPS点(经度、纬度)的dataframe,想为每个点找到它的2个最近邻。目前已经用R的spatstat包创建了ppp点模式对象,也算出了每个点到前2个最近邻的距离,但还需要把这些最近邻和原数据里的hhid关联起来,得到包含原点hhid、近邻距离及对应近邻hhid的结果。
当前实现代码
# 1. store x and y coords in two vectors lon <- data$longitude lat <- data$latitude # 2. create two vectors xrange and yrange with dimensions of triangle that contain all points xrange <- range(lon, na.rm=T) yrange <- range(lat, na.rm=T) # 3. create ppp lf <- ppp(lon, lat, xrange, yrange) plot(lf) nndist(lf, k = 1:2)
当前输出(前5行示例)
dist.1 dist.2 [1,] 1.426925e-03 0.0017007414 [2,] 1.017287e-03 0.0015574895 [3,] 6.502012e-04 0.0010172867 [4,] 6.502012e-04 0.0007202307 [5,] 7.202307e-04 0.0010472445
期望输出格式
hhid dist.1 dist.1.hhid dist.2 dist.2.hhid 1 1.426925e-03 7 0.0017007414 3 2 1.017287e-03 6 0.0015574895 4 3 6.502012e-04 10 0.0010172867 5 4 6.502012e-04 2 0.0007202307 8 5 7.202307e-04 1 0.0010472445 13
原数据集前20行
structure(list(hhid = c(2004L, 2006L, 2009L, 2012L, 2013L, 2020L, 2022L, 2023L, 2028L, 2029L, 2035L, 2036L, 2043L, 2046L, 2047L, 2059L, 2062L, 2063L, 2065L, 2066L), longitude = c(-1.478302479, -1.477469802, -1.476488709, -1.476146936, -1.47547996, -1.475799441, -1.475903392, -1.476232767, -1.476053953, -1.477196693, -1.476906657, -1.478778243, -1.480723381, -1.433436394, -1.433033824, -1.428791046, -1.431989908, -1.432058454, -1.43134892, -1.430848002), latitude = c(12.10552216, 12.10700512, 12.10673618, 12.10618305, 12.10645485, 12.10846806, 12.1080761, 12.10830975, 12.11114883, 12.11076546, 12.11197853, 12.11345387, 12.10725021, 12.1183548, 12.11699867, 12.11466122, 12.1154108, 12.11545277, 12.11554337, 12.11567497)), row.names = c(NA, 20L), class = "data.frame")
解决方案
要关联近邻的hhid,核心是用spatstat包的nnwhich()函数获取每个点最近邻的索引,再通过索引从原数据中提取对应hhid。
完整代码
library(spatstat) # 加载原数据(测试时可直接运行上述原数据集结构代码) # data <- structure(...) # 创建ppp对象 lf <- ppp(data$longitude, data$latitude, xrange = range(data$longitude, na.rm = TRUE), yrange = range(data$latitude, na.rm = TRUE)) # 计算前2个最近邻的距离 nn_dist <- nndist(lf, k = 1:2) # 计算前2个最近邻在原数据中的索引 nn_index <- nnwhich(lf, k = 1:2) # 构建最终结果数据框 result <- data.frame( hhid = data$hhid, dist.1 = nn_dist[, 1], dist.1.hhid = data$hhid[nn_index[, 1]], dist.2 = nn_dist[, 2], dist.2.hhid = data$hhid[nn_index[, 2]] ) # 查看前5行结果 head(result, 5)
代码说明
nnwhich(lf, k = 1:2):返回每个点的第1、2近邻在ppp对象中的位置索引,该索引与原dataframe的行顺序完全一致- 通过
data$hhid[nn_index[,1]]可提取每个点第1近邻的hhid,同理可得第2近邻的hhid - 最后将原点
hhid、近邻距离、近邻hhid合并为符合期望的结果格式
内容的提问来源于stack exchange,提问作者bellbyrne
相关产品推荐
相关产品推荐

