You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对含数值的规则字符串进行聚类?技术问询

问题描述

我拥有以规则形式存在的含数字字符串,生成代码如下:

make_rule <- function(n=5){
x1 <- sample(1:20,n)
x2 <- sample(c(" <= "," >= "),n,replace = T)
x3 <- round(rnorm(n),4)
res <- paste0("X[,",x1,"]", x2 , x3,collapse = " & ")
return( res )
}
rules <- lapply(1:10000, function(x)  make_rule(n=sample(2:5,1))  )

生成的规则示例:

head(rules)
[[1]]
[1] "X[,16] <= -0.664 & X[,11] <= -2.1891 & X[,17] >= -0.4138"

[[2]]
[1] "X[,1] <= 1.2584 & X[,11] >= -0.7082 & X[,20] <= -0.0806"

[[3]]
[1] "X[,3] <= -0.6363 & X[,7] >= -0.853 & X[,10] >= -3.4069"

[[4]]
[1] "X[,7] <= 0.5907 & X[,1] <= 0.5892 & X[,8] >= -2.1081 & X[,9] >= -1.4655 & X[,18] >= 0.0914"

这些字符串长度不同,且内部包含数值。因此我认为可能需要应用两种距离度量:

  • 针对数值部分:提取规则中的索引数字(如16、11)和阈值数字(如-0.664、-2.1891)
  • 针对字符部分:提取规则中的结构片段(如X[,] <= 、& X[,] >=)

请问如何对这类数据进行聚类?


解决方案

步骤1:规则结构化解析

首先要把每个规则字符串拆解为标准化的结构化数据,这是聚类的基础。可以用正则表达式提取每个规则的核心元素:

# 定义解析函数
parse_rule <- function(rule_str) {
  # 拆分每个子规则
  sub_rules <- strsplit(rule_str, " & ")[[1]]
  # 提取每个子规则的特征:列索引、运算符、阈值
  features <- lapply(sub_rules, function(sr) {
    matches <- regmatches(sr, regexec("X\\[(,)(\\d+)\\]\\s*(<=|>=)\\s*(-?\\d+\\.?\\d*)", sr))[[1]]
    list(col = as.integer(matches[3]), op = matches[4], threshold = as.numeric(matches[5]))
  })
  # 按列索引排序,消除顺序影响
  features <- features[order(sapply(features, function(x) x$col))]
  return(features)
}

# 解析所有规则
parsed_rules <- lapply(rules, parse_rule)

步骤2:定义自定义距离度量

因为规则的结构和数值都需要考虑,我们可以构建一个加权的混合距离:

  1. 结构距离:计算两个规则的公共子规则结构(列索引+运算符)的匹配度,用Jaccard相似性的补值作为距离。
  2. 数值距离:对匹配的子规则,计算阈值的绝对差,再取平均值。

示例距离函数:

rule_distance <- function(r1, r2) {
  # 提取两个规则的结构标识(列+运算符)
  s1 <- sapply(r1, function(x) paste(x$col, x$op, sep = "_"))
  s2 <- sapply(r2, function(x) paste(x$col, x$op, sep = "_"))
  
  # 结构相似度:Jaccard系数
  common <- length(intersect(s1, s2))
  union <- length(union(s1, s2))
  struct_sim <- ifelse(union == 0, 0, common / union)
  struct_dist <- 1 - struct_sim
  
  # 数值距离:仅对匹配的结构计算阈值差
  matched_idx <- match(s1, s2)
  valid_matches <- !is.na(matched_idx)
  if (any(valid_matches)) {
    threshold_diffs <- abs(sapply(which(valid_matches), function(i) {
      r1[[i]]$threshold - r2[[matched_idx[i]]]$threshold
    }))
    numeric_dist <- mean(threshold_diffs)
  } else {
    numeric_dist <- 1  # 无匹配结构时数值距离设为最大值
  }
  
  # 加权混合距离,可根据需求调整权重
  weight_struct <- 0.6
  weight_numeric <- 0.4
  total_dist <- weight_struct * struct_dist + weight_numeric * numeric_dist
  
  return(total_dist)
}

# 生成距离矩阵(10000个样本计算量极大,建议先抽样)
sample_rules <- parsed_rules[1:1000]
dist_matrix <- as.matrix(proxy::dist(sample_rules, method = rule_distance))

步骤3:执行聚类

有了距离矩阵后,可以用层次聚类或DBSCAN等方法:

# 层次聚类
hc <- hclust(as.dist(dist_matrix), method = "ward.D2")
plot(hc)
# 切割聚类
clusters <- cutree(hc, k = 5)

# 或者DBSCAN(适合大样本)
library(dbscan)
db <- dbscan(dist_matrix, eps = 0.3, minPts = 5)

关键注意点

  • 规则顺序无关性:解析时一定要按列索引排序,避免因为子规则顺序不同被误判为不同规则。
  • 权重调整:根据业务需求调整结构和数值部分的权重,比如如果结构一致性更重要,就提高结构权重。
  • 大样本优化:10000个样本直接计算距离矩阵会非常耗内存,建议先抽样验证方法,或者用近似距离算法。

内容的提问来源于stack exchange,提问作者mr.T

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 04:55:21