为何spThin包thin函数每次运行输出不同?如何选择结果?
问题描述
我尝试对存在数据进行空间去重,以减少用于MaxEnt模型的仅存在数据集的采样偏差,使用的R代码及第一次运行输出如下:
library(spThin) thinned_dataset_full.100.1 <- thin(loc.data = AP.bioclim.combined.3, lat.col = "Latitude", long.col = "Longitude", spec.col = "Species", thin.par = 5, reps = 100, locs.thinned.list.return = TRUE, write.files = FALSE, write.log.file = FALSE)
运行输出:
Beginning Spatial Thinning. Script Started at: Wed Jun 17 17:02:32 2026 lat.long.thin.count 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1206 2 6 4 11 6 20 17 14 7 6 6 1 [1] "Maximum number of records after thinning: 1206" [1] "Number of data.frames with max records: 1" [1] "No files written for this run."
我发现多次运行同一代码,每次输出结果都不同,例如第二次和第三次运行:
第二次运行
#Attempt 2 library(spThin) thinned_dataset_full.100.2 <- thin(loc.data = AP.bioclim.combined.3, lat.col = "Latitude", long.col = "Longitude", spec.col = "Species", thin.par = 5, reps = locs.thinned.list.return = TRUE, write.files = FALSE, write.log.file = FALSE)
输出:
Beginning Spatial Thinning. Script Started at: Wed Jun 17 17:04:03 2026 lat.long.thin.count 1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1205 1206 1 4 4 3 12 12 13 13 9 12 5 9 2 1 [1] "Maximum number of records after thinning: 1206" [1] "Number of data.frames with max records: 1" [1] "No files written for this run."
第三次运行
#Attempt 3 library(spThin) thinned_dataset_full.100.3 <- thin(loc.data = AP.bioclim.combined.3, lat.col = "Latitude", long.col = "Longitude", spec.col = "Species", thin.par = 5, reps = 100, locs.thinned.list.return = TRUE, write.files = FALSE, write.log.file = FALSE)
输出:
Beginning Spatial Thinning. Script Started at: Wed Jun 17 17:06:25 2026 lat.long.thin.count 1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1206 1 1 5 5 11 11 15 17 12 11 4 4 3 [1] "Maximum number of records after thinning: 1206" [1] "Number of data.frames with max records: 3" [1] "No files written for this run."
请问为何会出现这种情况?我该如何选择要解读的输出结果?
解答
一、多次运行结果不同的原因
spThin包的thin()函数在空间去重时,当某一区域内存在多个满足距离阈值(你设置的thin.par=5,即5个单位距离内仅保留一条记录)的采样点时,会随机选择保留其中一个点。每次运行时随机数生成器的初始状态不同,导致选择结果存在差异,最终输出的去重数据集数量分布和具体保留的采样点都会不一样。
另外注意到你第二次运行的代码里reps参数写错了,写成了reps = locs.thinned.list.return = TRUE,正确写法应为reps=100,不过这没影响最终最大记录数,只是可能导致实际重复次数不符合预期,但核心差异还是来自函数内置的随机采样逻辑。
二、如何选择输出结果
针对MaxEnt模型的需求,推荐以下几种选择方式:
- 选记录数最多的数据集:从多次重复结果里,挑出保留记录数最多的那些数据集(比如你第三次运行里有3个记录数为1206的数据集),这类数据集既满足空间去重要求,又保留了最多的物种分布信息,能在减少采样偏差的同时避免丢失过多有效数据。
- 多数据集集成分析:如果算力允许,用多个不同的去重数据集分别训练MaxEnt模型,然后对模型结果进行集成(比如取平均预测值、合并变量重要性),这样能消除单次随机选择带来的不确定性,让模型结果更稳健。
- 固定随机种子:如果需要结果可重复,可以在调用
thin()函数前设置随机种子,比如set.seed(123),这样每次运行thin()函数时,随机选择的逻辑会完全一致,输出结果也会相同。
内容的提问来源于stack exchange,提问作者Insect_biologist
相关产品推荐
相关产品推荐

