You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R spatstat识别GPS点最近邻并关联原数据集ID

问题描述

我有一个包含488个GPS点(经度、纬度)的dataframe,想为每个点找到它的2个最近邻。目前已经用R的spatstat包创建了ppp点模式对象,也算出了每个点到前2个最近邻的距离,但还需要把这些最近邻和原数据里的hhid关联起来,得到包含原点hhid、近邻距离及对应近邻hhid的结果。

当前实现代码

# 1. store x and y coords in two vectors
lon <- data$longitude
lat <- data$latitude

# 2. create two vectors xrange and yrange with dimensions of triangle that contain all points
xrange <- range(lon, na.rm=T)
yrange <- range(lat, na.rm=T)

# 3. create ppp
lf <- ppp(lon, lat, xrange, yrange)

plot(lf)

nndist(lf, k = 1:2)

当前输出(前5行示例)

dist.1       dist.2
  [1,] 1.426925e-03 0.0017007414
  [2,] 1.017287e-03 0.0015574895
  [3,] 6.502012e-04 0.0010172867
  [4,] 6.502012e-04 0.0007202307
  [5,] 7.202307e-04 0.0010472445

期望输出格式

hhid         dist.1  dist.1.hhid         dist.2     dist.2.hhid
  1    1.426925e-03             7  0.0017007414                 3
  2    1.017287e-03             6  0.0015574895                 4
  3    6.502012e-04            10  0.0010172867                 5
  4    6.502012e-04             2  0.0007202307                 8
  5    7.202307e-04             1  0.0010472445                13

原数据集前20行

structure(list(hhid = c(2004L, 2006L, 2009L, 2012L, 2013L, 2020L, 
2022L, 2023L, 2028L, 2029L, 2035L, 2036L, 2043L, 2046L, 2047L, 
2059L, 2062L, 2063L, 2065L, 2066L), longitude = c(-1.478302479, 
-1.477469802, -1.476488709, -1.476146936, -1.47547996, -1.475799441, 
-1.475903392, -1.476232767, -1.476053953, -1.477196693, -1.476906657, 
-1.478778243, -1.480723381, -1.433436394, -1.433033824, -1.428791046, 
-1.431989908, -1.432058454, -1.43134892, -1.430848002), latitude = c(12.10552216, 
12.10700512, 12.10673618, 12.10618305, 12.10645485, 12.10846806, 
12.1080761, 12.10830975, 12.11114883, 12.11076546, 12.11197853, 
12.11345387, 12.10725021, 12.1183548, 12.11699867, 12.11466122, 
12.1154108, 12.11545277, 12.11554337, 12.11567497)), row.names = c(NA, 
20L), class = "data.frame")
解决方案

要关联近邻的hhid,核心是用spatstat包的nnwhich()函数获取每个点最近邻的索引,再通过索引从原数据中提取对应hhid。

完整代码

library(spatstat)

# 加载原数据(测试时可直接运行上述原数据集结构代码)
# data <- structure(...)

# 创建ppp对象
lf <- ppp(data$longitude, data$latitude, 
          xrange = range(data$longitude, na.rm = TRUE),
          yrange = range(data$latitude, na.rm = TRUE))

# 计算前2个最近邻的距离
nn_dist <- nndist(lf, k = 1:2)
# 计算前2个最近邻在原数据中的索引
nn_index <- nnwhich(lf, k = 1:2)

# 构建最终结果数据框
result <- data.frame(
  hhid = data$hhid,
  dist.1 = nn_dist[, 1],
  dist.1.hhid = data$hhid[nn_index[, 1]],
  dist.2 = nn_dist[, 2],
  dist.2.hhid = data$hhid[nn_index[, 2]]
)

# 查看前5行结果
head(result, 5)

代码说明

  • nnwhich(lf, k = 1:2):返回每个点的第1、2近邻在ppp对象中的位置索引,该索引与原dataframe的行顺序完全一致
  • 通过data$hhid[nn_index[,1]]可提取每个点第1近邻的hhid,同理可得第2近邻的hhid
  • 最后将原点hhid、近邻距离、近邻hhid合并为符合期望的结果格式

内容的提问来源于stack exchange,提问作者bellbyrne

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 10:32:37