You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何spThin包thin函数每次运行输出不同?如何选择结果?

问题描述

我尝试对存在数据进行空间去重,以减少用于MaxEnt模型的仅存在数据集的采样偏差,使用的R代码及第一次运行输出如下:

library(spThin)
thinned_dataset_full.100.1 <- thin(loc.data = AP.bioclim.combined.3, lat.col = "Latitude", long.col = "Longitude", spec.col = "Species", thin.par = 5, reps = 100, locs.thinned.list.return = TRUE, write.files = FALSE, write.log.file = FALSE)

运行输出:

Beginning Spatial Thinning.
 Script Started at: Wed Jun 17 17:02:32 2026
lat.long.thin.count
1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1206 
   2    6    4   11    6   20   17   14    7    6    6    1 
[1] "Maximum number of records after thinning: 1206"
[1] "Number of data.frames with max records: 1"
[1] "No files written for this run."

我发现多次运行同一代码,每次输出结果都不同,例如第二次和第三次运行:

第二次运行

#Attempt 2
library(spThin)
thinned_dataset_full.100.2 <- thin(loc.data = AP.bioclim.combined.3, lat.col = "Latitude", long.col = "Longitude", spec.col = "Species", thin.par = 5, reps = locs.thinned.list.return = TRUE, write.files = FALSE, write.log.file = FALSE)

输出:

Beginning Spatial Thinning.
 Script Started at: Wed Jun 17 17:04:03 2026
lat.long.thin.count
1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1205 1206 
   1    4    4    3   12   12   13   13    9   12    5    9    2    1 
[1] "Maximum number of records after thinning: 1206"
[1] "Number of data.frames with max records: 1"
[1] "No files written for this run."

第三次运行

#Attempt 3
library(spThin)
thinned_dataset_full.100.3 <- thin(loc.data = AP.bioclim.combined.3, lat.col = "Latitude", long.col = "Longitude", spec.col = "Species", thin.par = 5, reps = 100, locs.thinned.list.return = TRUE, write.files = FALSE, write.log.file = FALSE)

输出:

Beginning Spatial Thinning.
 Script Started at: Wed Jun 17 17:06:25 2026
lat.long.thin.count
1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1206 
   1    1    5    5   11   11   15   17   12   11    4    4    3 
[1] "Maximum number of records after thinning: 1206"
[1] "Number of data.frames with max records: 3"
[1] "No files written for this run."

请问为何会出现这种情况?我该如何选择要解读的输出结果?


解答

一、多次运行结果不同的原因

spThin包的thin()函数在空间去重时,当某一区域内存在多个满足距离阈值(你设置的thin.par=5,即5个单位距离内仅保留一条记录)的采样点时,会随机选择保留其中一个点。每次运行时随机数生成器的初始状态不同,导致选择结果存在差异,最终输出的去重数据集数量分布和具体保留的采样点都会不一样。

另外注意到你第二次运行的代码里reps参数写错了,写成了reps = locs.thinned.list.return = TRUE,正确写法应为reps=100,不过这没影响最终最大记录数,只是可能导致实际重复次数不符合预期,但核心差异还是来自函数内置的随机采样逻辑。

二、如何选择输出结果

针对MaxEnt模型的需求,推荐以下几种选择方式:

  • 选记录数最多的数据集:从多次重复结果里,挑出保留记录数最多的那些数据集(比如你第三次运行里有3个记录数为1206的数据集),这类数据集既满足空间去重要求,又保留了最多的物种分布信息,能在减少采样偏差的同时避免丢失过多有效数据。
  • 多数据集集成分析:如果算力允许,用多个不同的去重数据集分别训练MaxEnt模型,然后对模型结果进行集成(比如取平均预测值、合并变量重要性),这样能消除单次随机选择带来的不确定性,让模型结果更稳健。
  • 固定随机种子:如果需要结果可重复,可以在调用thin()函数前设置随机种子,比如set.seed(123),这样每次运行thin()函数时,随机选择的逻辑会完全一致,输出结果也会相同。

内容的提问来源于stack exchange,提问作者Insect_biologist

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 10:40:37