如何对含数值的规则字符串进行聚类?技术问询
问题描述
我拥有以规则形式存在的含数字字符串,生成代码如下:
make_rule <- function(n=5){ x1 <- sample(1:20,n) x2 <- sample(c(" <= "," >= "),n,replace = T) x3 <- round(rnorm(n),4) res <- paste0("X[,",x1,"]", x2 , x3,collapse = " & ") return( res ) } rules <- lapply(1:10000, function(x) make_rule(n=sample(2:5,1)) )
生成的规则示例:
head(rules) [[1]] [1] "X[,16] <= -0.664 & X[,11] <= -2.1891 & X[,17] >= -0.4138" [[2]] [1] "X[,1] <= 1.2584 & X[,11] >= -0.7082 & X[,20] <= -0.0806" [[3]] [1] "X[,3] <= -0.6363 & X[,7] >= -0.853 & X[,10] >= -3.4069" [[4]] [1] "X[,7] <= 0.5907 & X[,1] <= 0.5892 & X[,8] >= -2.1081 & X[,9] >= -1.4655 & X[,18] >= 0.0914"
这些字符串长度不同,且内部包含数值。因此我认为可能需要应用两种距离度量:
- 针对数值部分:提取规则中的索引数字(如
16、11)和阈值数字(如-0.664、-2.1891) - 针对字符部分:提取规则中的结构片段(如
X[,] <=、& X[,] >=)
请问如何对这类数据进行聚类?
解决方案
步骤1:规则结构化解析
首先要把每个规则字符串拆解为标准化的结构化数据,这是聚类的基础。可以用正则表达式提取每个规则的核心元素:
# 定义解析函数 parse_rule <- function(rule_str) { # 拆分每个子规则 sub_rules <- strsplit(rule_str, " & ")[[1]] # 提取每个子规则的特征:列索引、运算符、阈值 features <- lapply(sub_rules, function(sr) { matches <- regmatches(sr, regexec("X\\[(,)(\\d+)\\]\\s*(<=|>=)\\s*(-?\\d+\\.?\\d*)", sr))[[1]] list(col = as.integer(matches[3]), op = matches[4], threshold = as.numeric(matches[5])) }) # 按列索引排序,消除顺序影响 features <- features[order(sapply(features, function(x) x$col))] return(features) } # 解析所有规则 parsed_rules <- lapply(rules, parse_rule)
步骤2:定义自定义距离度量
因为规则的结构和数值都需要考虑,我们可以构建一个加权的混合距离:
- 结构距离:计算两个规则的公共子规则结构(列索引+运算符)的匹配度,用Jaccard相似性的补值作为距离。
- 数值距离:对匹配的子规则,计算阈值的绝对差,再取平均值。
示例距离函数:
rule_distance <- function(r1, r2) { # 提取两个规则的结构标识(列+运算符) s1 <- sapply(r1, function(x) paste(x$col, x$op, sep = "_")) s2 <- sapply(r2, function(x) paste(x$col, x$op, sep = "_")) # 结构相似度:Jaccard系数 common <- length(intersect(s1, s2)) union <- length(union(s1, s2)) struct_sim <- ifelse(union == 0, 0, common / union) struct_dist <- 1 - struct_sim # 数值距离:仅对匹配的结构计算阈值差 matched_idx <- match(s1, s2) valid_matches <- !is.na(matched_idx) if (any(valid_matches)) { threshold_diffs <- abs(sapply(which(valid_matches), function(i) { r1[[i]]$threshold - r2[[matched_idx[i]]]$threshold })) numeric_dist <- mean(threshold_diffs) } else { numeric_dist <- 1 # 无匹配结构时数值距离设为最大值 } # 加权混合距离,可根据需求调整权重 weight_struct <- 0.6 weight_numeric <- 0.4 total_dist <- weight_struct * struct_dist + weight_numeric * numeric_dist return(total_dist) } # 生成距离矩阵(10000个样本计算量极大,建议先抽样) sample_rules <- parsed_rules[1:1000] dist_matrix <- as.matrix(proxy::dist(sample_rules, method = rule_distance))
步骤3:执行聚类
有了距离矩阵后,可以用层次聚类或DBSCAN等方法:
# 层次聚类 hc <- hclust(as.dist(dist_matrix), method = "ward.D2") plot(hc) # 切割聚类 clusters <- cutree(hc, k = 5) # 或者DBSCAN(适合大样本) library(dbscan) db <- dbscan(dist_matrix, eps = 0.3, minPts = 5)
关键注意点
- 规则顺序无关性:解析时一定要按列索引排序,避免因为子规则顺序不同被误判为不同规则。
- 权重调整:根据业务需求调整结构和数值部分的权重,比如如果结构一致性更重要,就提高结构权重。
- 大样本优化:10000个样本直接计算距离矩阵会非常耗内存,建议先抽样验证方法,或者用近似距离算法。
内容的提问来源于stack exchange,提问作者mr.T
相关产品推荐
相关产品推荐

