You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于min/max列重叠性分配字母分组的R语言实现及小数问题

问题描述

我有一个包含min和max列的数据框:

df <- data.frame(min=c(2, 4, 3, 3, 2, 6),
                 max=c(2.9, 5.9, 3.9, 4.9, 7.9, 7.9))

希望基于这两列的重叠/非重叠关系创建并分配分组:若某行与之前所有行均不重叠,则分配新字母;所有与该行重叠的行都使用该字母,完全包含于其他行的行也使用相同字母标识。

例如前两行无重叠,分别分配"a"和"b";第1行与第5行重叠,故第5行也为"a"。目标输出如下:

min max group
1   2 2.9     a
2   4 5.9     b
3   3 3.9     c
4   3 4.9    bc
5   2 7.9  abcd
6   6 7.9     d

这类似multcomp包中compact letter display(紧凑字母显示)的统计显著性标识逻辑,但针对的是min/max值。最终需将字母标注在ggplot2的geom_errorbar上方。

使用@Mark的方案时发现,当min/max列小数位数超过1位时,结果出现差异,示例数据如下:

testdf2 <- structure(list(site = structure(1:9, levels = c("TR", "SF", 
"AR", "MR", "PD", "PL", "GC", "WM", "EM"), class = c("ordered", 
"factor")), mean = c(18.160173892, 17.4521769872, 12.078365989, 
20.4040729028, 25.272373546, 29.216489076, 7.1171799882, 7.3105614156, 
9.781365245), min = c(14.1, 12.6, 9.2, 15.1, 19.8, 22.3, 3.9, 
4.8, 6.2), max = c(22.3, 22.9, 15.2, 31.2, 31.3, 36.5, 11.1, 
9.8, 14.7), min2 = c(14.12, 12.62, 9.21, 15.06, 19.76, 22.26, 
3.87, 4.84, 6.24), max2 = c(22.35, 22.95, 15.22, 31.16, 31.32, 
36.49, 11.11, 9.84, 14.7), overlap = c("a", "ab", "ac", "ab", 
"ab", "b", "c", "c", "ac"), overlap2 = c("a", "a", "ab", "a", 
"a", "a", "b", "b", "ab")), row.names = c(1L, 2L, 3L, 4L, 5L, 
6L, 7L, 8L, 9L), class = "data.frame")

通过绘图可更直观看到差异:

ggplot(data=testdf2, aes(x = site, y = mean)) +
  geom_errorbar(aes(ymin=min, ymax=max), size=1, width=0.3) +
  geom_col(alpha=0.5) +
  ylab("One-decimal min max") +
  geom_text(aes(label = overlap, y = max), vjust = -.5) +
  theme_bw()

ggplot(data=testdf2, aes(x = site, y = mean)) +
  geom_errorbar(aes(ymin=min2, ymax=max2), size=1, width=0.3) +
  geom_col(alpha=0.5) +
  ylab("Two-decimal min max") +
  geom_text(aes(label = overlap2, y = max2), vjust = -.5) +
  theme_bw()

请问该如何解决小数位数导致分组结果不同的问题?


解决方案

问题根源是浮点数精度误差:当小数位数增加时,浮点数的二进制存储特性会导致微小的数值偏差,原本应该判定为重叠的区间可能被误判,反之亦然。以下两种方案可解决该问题:

方案1:统一舍入小数位数

选择与原始数据有效位数匹配的小数位数,对所有min/max列做舍入处理,消除精度偏差后再执行分组逻辑:

# 定义舍入函数,指定目标小数位数
round_tolerance <- function(x, digits = 2) {
  round(x, digits = digits)
}

# 对testdf2的高精度列进行舍入
testdf2$min2_rounded <- round_tolerance(testdf2$min2, digits = 2)
testdf2$max2_rounded <- round_tolerance(testdf2$max2, digits = 2)

# 核心分组逻辑(兼容舍入后的数据)
assign_groups <- function(df, min_col, max_col) {
  n <- nrow(df)
  groups <- vector("list", n)
  current_letter <- 1
  
  for (i in 1:n) {
    current_min <- df[[min_col]][i]
    current_max <- df[[max_col]][i]
    
    # 检查与已有组的重叠关系
    overlap_groups <- which(sapply(1:(current_letter-1), function(g) {
      any(df[[min_col]][groups[[g]]] <= current_max & df[[max_col]][groups[[g]]] >= current_min)
    }))
    
    if (length(overlap_groups) == 0) {
      # 无重叠则新建组
      groups[[current_letter]] <- i
      current_letter <- current_letter + 1
    } else {
      # 加入所有重叠的组
      for (g in overlap_groups) {
        groups[[g]] <- c(groups[[g]], i)
      }
    }
  }
  
  # 生成有序的字母标签
  group_labels <- rep("", n)
  for (g in 1:(current_letter-1)) {
    label <- letters[g]
    group_labels[groups[[g]]] <- paste0(group_labels[groups[[g]]], label)
  }
  group_labels <- sapply(group_labels, function(x) paste(sort(strsplit(x, "")[[1]]), collapse = ""))
  return(group_labels)
}

# 用舍入后的数据计算修正后的分组
testdf2$overlap2_fixed <- assign_groups(testdf2, "min2_rounded", "max2_rounded")

方案2:引入容差的重叠判断

在区间重叠判断中加入一个极小的容差值,允许数值在微小范围内的偏差仍被视为重叠:

# 带容差的分组逻辑
assign_groups_with_tol <- function(df, min_col, max_col, tol = 1e-4) {
  n <- nrow(df)
  groups <- vector("list", n)
  current_letter <- 1
  
  for (i in 1:n) {
    current_min <- df[[min_col]][i]
    current_max <- df[[max_col]][i]
    
    overlap_groups <- which(sapply(1:(current_letter-1), function(g) {
      # 带容差的重叠判断,抵消浮点数精度误差
      any(df[[min_col]][groups[[g]]] <= current_max + tol & df[[max_col]][groups[[g]]] >= current_min - tol)
    }))
    
    if (length(overlap_groups) == 0) {
      groups[[current_letter]] <- i
      current_letter <- current_letter + 1
    } else {
      for (g in overlap_groups) {
        groups[[g]] <- c(groups[[g]], i)
      }
    }
  }
  
  group_labels <- rep("", n)
  for (g in 1:(current_letter-1)) {
    label <- letters[g]
    group_labels[groups[[g]]] <- paste0(group_labels[groups[[g]]], label)
  }
  group_labels <- sapply(group_labels, function(x) paste(sort(strsplit(x, "")[[1]]), collapse = ""))
  return(group_labels)
}

# 用带容差的函数计算修正后的分组
testdf2$overlap2_fixed <- assign_groups_with_tol(testdf2, "min2", "max2", tol = 1e-4)

验证修正效果

使用修正后的分组标签绘制图形,结果会和单小数位的分组逻辑一致:

ggplot(data=testdf2, aes(x = site, y = mean)) +
  geom_errorbar(aes(ymin=min2, ymax=max2), size=1, width=0.3) +
  geom_col(alpha=0.5) +
  ylab("Two-decimal min max (fixed)") +
  geom_text(aes(label = overlap2_fixed, y = max2), vjust = -.5) +
  theme_bw()

内容的提问来源于stack exchange,提问作者Bradley S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 14:35:54