group_by后summarize中自定义函数异常:所有分组结果一致
聚类数据分组计算凸包面积异常问题
我有一份带聚类标签的点数据,想要计算每个聚类的点数和凸包面积。我写了一个计算单聚类凸包面积的函数,单独处理每个聚类时结果正确,但在dataframe %>% group_by() %>% summarize()链式调用中,函数没有按聚类分别计算面积,而是把整个数据集当成单个聚类计算,再把这个值返回给所有聚类。
示例数据集
x <- c(0, 0, 5, 5, 8, 8, 10, 10 ) y <- c(0, 5, 5, 0, 0, 5, 5, 0) cluster <- c(1, 1, 1, 1, 2, 2, 2, 2) dat.clustered <- data.frame(x, y, cluster) ## 每个聚类有4个点
聚类1是5×5的正方形,凸包面积为25;聚类2是2×5的矩形,面积为10。如果把所有点视为单个聚类,形成的是5×10的矩形,面积为50(这一点很关键)。
可以给每个聚类添加内部点,让场景更贴近实际聚类,同时验证n()在summarize()中能正常工作:
dat.clustered[nrow(dat.clustered) + 1,] <- c(2, 3, 1) # 聚类1内部 dat.clustered[nrow(dat.clustered) + 1,] <- c(4, 3, 1) # 聚类1内部 dat.clustered[nrow(dat.clustered) + 1,] <- c(9, 1, 2) # 聚类2内部 ## 现在聚类1共有6个点,聚类2共有5个点
自定义凸包面积计算函数
Polygon()来自sp包,函数实现如下:
area.calc <- function(data){ area = chull(data) %>% c(., .[1]) %>% data[.,] %>% .[, 1:2] %>% sp::Polygon(., hole = F) %>% .@area return(area) }
函数单独调用验证
单独处理每个聚类或整个数据集时,函数结果正确:
area.calc(filter(dat.clustered, cluster == 1)) # 返回25 area.calc(filter(dat.clustered, cluster == 2)) # 返回10 area.calc(dat.clustered) # 返回50(整个数据集作为单个聚类)
分组调用时的异常结果
使用group_by()+summarize()调用函数时,得到错误结果:
clust.smy <- dat.clustered %>% group_by(cluster) %>% summarize(count = n(), area = area.calc(data = .))
返回结果:
# A tibble: 2 × 3 cluster count area <dbl> <int> <dbl> 1 1 6 50 2 2 5 50
可以看到,count返回的各聚类点数正确,但area都返回了整个数据集的凸包面积50,而不是对应聚类的面积。
预期结果
| cluster | count | area |
|---|---|---|
| 1 | 6 | 25 |
| 2 | 5 | 10 |
尝试的解决方法及问题
- 使用
cur_data()替代.
尝试用cur_data()传递分组数据,出现错误:
clust.smy <- dat.clustered %>% group_by(cluster) %>% summarize(count = n(), area = area.calc(data = cur_data()))
错误信息:
Error in `summarize()`: ! Problem while computing `area = area.calc(data = cur_data())`. ℹ The error occurred in group 1: cluster = 1. Caused by error: ! error in evaluating the argument 'obj' in selecting a method for function 'coordinates': Can't subset elements past the end. ℹ Locations 4, 2, 3, and 4 don't exist. ℹ There is only 1 element. Run `rlang::last_error()` to see where the error occurred.
- 直接在
summarize()中写管道逻辑
不调用自定义函数,直接把计算逻辑写在summarize()里,结果依然错误:
clust.smy2 <- dat.clustered %>% group_by(cluster) %>% summarize(count = n(), area = (.) %>% chull() %>% c(., .[1]) %>% dat.clustered[.,] %>% .[, 1:2] %>% sp::Polygon(., hole = F) %>% .@area)
返回结果:
# A tibble: 2 × 3 cluster count area <dbl> <int> <dbl> 1 1 6 50 2 2 5 50
- 用
cur_data()替代.直接写逻辑
把上述逻辑中的.换成cur_data(),结果变成两个聚类都返回聚类1的面积:
clust.smy2 <- dat.clustered %>% group_by(cluster) %>% summarize(count = n(), area = cur_data() %>% chull() %>% c(., .[1]) %>% dat.clustered[.,] %>% .[, 1:2] %>% sp::Polygon(., hole = F) %>% .@area)
返回结果:
> clust.smy2 # A tibble: 2 × 3 cluster count area <dbl> <int> <dbl> 1 1 6 25 2 2 5 25
会话信息
> sessionInfo() R version 4.2.0 (2022-04-22 ucrt) Platform: x86_64-w64-mingw32/x64 (64-bit) Running under: Windows 10 x64 (build 22000) Matrix products: default locale: [1] LC_COLLATE=English_United States.utf8 LC_CTYPE=English_United States.utf8 LC_MONETARY=English_United States.utf8 [4] LC_NUMERIC=C LC_TIME=English_United States.utf8 attached base packages: [1] stats graphics grDevices utils datasets methods base other attached packages: [1] dplyr_1.0.9 plyr_1.8.7 loaded via a namespace (and not attached): [1] Rcpp_1.0.8.3 rstudioapi_0.13 magrittr_2.0.3 tidyselect_1.1.2 munsell_0.5.0 lattice_0.20-45 colorspace_2.0-3 R6_2.5.1 [9] rlang_1.0.6 factoextra_1.0.7 fansi_1.0.3 tools_4.2.0 grid_4.2.0 gtable_0.3.0 utf8_1.2.2 cli_3.4.1 [17] DBI_1.1.3 ellipsis_0.3.2 assertthat_0.2.1 tibble_3.1.7 lifecycle_1.0.3 crayon_1.5.1 purrr_0.3.4 ggplot2_3.4.0 [25] vctrs_0.5.1 ggrepel_0.9.2 glue_1.6.2 sp_1.5-1 compiler_4.2.0 pillar_1.7.0 generics_0.1.2 scales_1.2.0 [33] pkgconfig_2.0.3
内容的提问来源于stack exchange,提问作者njero
相关产品推荐
相关产品推荐

