基于min/max列重叠性分配字母分组的R语言实现及小数问题
问题描述
我有一个包含min和max列的数据框:
df <- data.frame(min=c(2, 4, 3, 3, 2, 6), max=c(2.9, 5.9, 3.9, 4.9, 7.9, 7.9))
希望基于这两列的重叠/非重叠关系创建并分配分组:若某行与之前所有行均不重叠,则分配新字母;所有与该行重叠的行都使用该字母,完全包含于其他行的行也使用相同字母标识。
例如前两行无重叠,分别分配"a"和"b";第1行与第5行重叠,故第5行也为"a"。目标输出如下:
min max group 1 2 2.9 a 2 4 5.9 b 3 3 3.9 c 4 3 4.9 bc 5 2 7.9 abcd 6 6 7.9 d
这类似multcomp包中compact letter display(紧凑字母显示)的统计显著性标识逻辑,但针对的是min/max值。最终需将字母标注在ggplot2的geom_errorbar上方。
使用@Mark的方案时发现,当min/max列小数位数超过1位时,结果出现差异,示例数据如下:
testdf2 <- structure(list(site = structure(1:9, levels = c("TR", "SF", "AR", "MR", "PD", "PL", "GC", "WM", "EM"), class = c("ordered", "factor")), mean = c(18.160173892, 17.4521769872, 12.078365989, 20.4040729028, 25.272373546, 29.216489076, 7.1171799882, 7.3105614156, 9.781365245), min = c(14.1, 12.6, 9.2, 15.1, 19.8, 22.3, 3.9, 4.8, 6.2), max = c(22.3, 22.9, 15.2, 31.2, 31.3, 36.5, 11.1, 9.8, 14.7), min2 = c(14.12, 12.62, 9.21, 15.06, 19.76, 22.26, 3.87, 4.84, 6.24), max2 = c(22.35, 22.95, 15.22, 31.16, 31.32, 36.49, 11.11, 9.84, 14.7), overlap = c("a", "ab", "ac", "ab", "ab", "b", "c", "c", "ac"), overlap2 = c("a", "a", "ab", "a", "a", "a", "b", "b", "ab")), row.names = c(1L, 2L, 3L, 4L, 5L, 6L, 7L, 8L, 9L), class = "data.frame")
通过绘图可更直观看到差异:
ggplot(data=testdf2, aes(x = site, y = mean)) + geom_errorbar(aes(ymin=min, ymax=max), size=1, width=0.3) + geom_col(alpha=0.5) + ylab("One-decimal min max") + geom_text(aes(label = overlap, y = max), vjust = -.5) + theme_bw() ggplot(data=testdf2, aes(x = site, y = mean)) + geom_errorbar(aes(ymin=min2, ymax=max2), size=1, width=0.3) + geom_col(alpha=0.5) + ylab("Two-decimal min max") + geom_text(aes(label = overlap2, y = max2), vjust = -.5) + theme_bw()
请问该如何解决小数位数导致分组结果不同的问题?
解决方案
问题根源是浮点数精度误差:当小数位数增加时,浮点数的二进制存储特性会导致微小的数值偏差,原本应该判定为重叠的区间可能被误判,反之亦然。以下两种方案可解决该问题:
方案1:统一舍入小数位数
选择与原始数据有效位数匹配的小数位数,对所有min/max列做舍入处理,消除精度偏差后再执行分组逻辑:
# 定义舍入函数,指定目标小数位数 round_tolerance <- function(x, digits = 2) { round(x, digits = digits) } # 对testdf2的高精度列进行舍入 testdf2$min2_rounded <- round_tolerance(testdf2$min2, digits = 2) testdf2$max2_rounded <- round_tolerance(testdf2$max2, digits = 2) # 核心分组逻辑(兼容舍入后的数据) assign_groups <- function(df, min_col, max_col) { n <- nrow(df) groups <- vector("list", n) current_letter <- 1 for (i in 1:n) { current_min <- df[[min_col]][i] current_max <- df[[max_col]][i] # 检查与已有组的重叠关系 overlap_groups <- which(sapply(1:(current_letter-1), function(g) { any(df[[min_col]][groups[[g]]] <= current_max & df[[max_col]][groups[[g]]] >= current_min) })) if (length(overlap_groups) == 0) { # 无重叠则新建组 groups[[current_letter]] <- i current_letter <- current_letter + 1 } else { # 加入所有重叠的组 for (g in overlap_groups) { groups[[g]] <- c(groups[[g]], i) } } } # 生成有序的字母标签 group_labels <- rep("", n) for (g in 1:(current_letter-1)) { label <- letters[g] group_labels[groups[[g]]] <- paste0(group_labels[groups[[g]]], label) } group_labels <- sapply(group_labels, function(x) paste(sort(strsplit(x, "")[[1]]), collapse = "")) return(group_labels) } # 用舍入后的数据计算修正后的分组 testdf2$overlap2_fixed <- assign_groups(testdf2, "min2_rounded", "max2_rounded")
方案2:引入容差的重叠判断
在区间重叠判断中加入一个极小的容差值,允许数值在微小范围内的偏差仍被视为重叠:
# 带容差的分组逻辑 assign_groups_with_tol <- function(df, min_col, max_col, tol = 1e-4) { n <- nrow(df) groups <- vector("list", n) current_letter <- 1 for (i in 1:n) { current_min <- df[[min_col]][i] current_max <- df[[max_col]][i] overlap_groups <- which(sapply(1:(current_letter-1), function(g) { # 带容差的重叠判断,抵消浮点数精度误差 any(df[[min_col]][groups[[g]]] <= current_max + tol & df[[max_col]][groups[[g]]] >= current_min - tol) })) if (length(overlap_groups) == 0) { groups[[current_letter]] <- i current_letter <- current_letter + 1 } else { for (g in overlap_groups) { groups[[g]] <- c(groups[[g]], i) } } } group_labels <- rep("", n) for (g in 1:(current_letter-1)) { label <- letters[g] group_labels[groups[[g]]] <- paste0(group_labels[groups[[g]]], label) } group_labels <- sapply(group_labels, function(x) paste(sort(strsplit(x, "")[[1]]), collapse = "")) return(group_labels) } # 用带容差的函数计算修正后的分组 testdf2$overlap2_fixed <- assign_groups_with_tol(testdf2, "min2", "max2", tol = 1e-4)
验证修正效果
使用修正后的分组标签绘制图形,结果会和单小数位的分组逻辑一致:
ggplot(data=testdf2, aes(x = site, y = mean)) + geom_errorbar(aes(ymin=min2, ymax=max2), size=1, width=0.3) + geom_col(alpha=0.5) + ylab("Two-decimal min max (fixed)") + geom_text(aes(label = overlap2_fixed, y = max2), vjust = -.5) + theme_bw()
内容的提问来源于stack exchange,提问作者Bradley S
相关产品推荐
相关产品推荐

