You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求Azure Data Explorer中关联原行的Kusto聚类查询方案

在Azure Data Explorer中获取聚类对应的原始行

要解决autocluster/basket仅返回计数、无法关联原始行的问题,可以通过将聚类结果与原始表动态关联的方式实现,核心思路是利用聚类输出的Pattern特征列匹配原始数据。以下是两种可行方案:

方案1:获取每个聚类对应的原始行集合

先通过autocluster生成聚类特征,再用apply为每个聚类筛选出匹配的原始行:

// 1. 先获取聚类结果(包含SegmentId、计数、特征Pattern)
let cluster_info = YourTableName
| evaluate autocluster(0.01) // 0.01为聚类阈值,可根据需求调整,值越小聚类越细分
| project SegmentId, ClusterRowCount = Count, Pattern;

// 2. 为每个聚类匹配原始行
cluster_info
| apply matched_rows = (
    YourTableName
    | where 
        // 动态匹配Pattern中的所有特征列,无需手动指定每一列
        array_all(bag_keys(Pattern), column_name => 
            todynamic(YourTableName[column_name]) == Pattern[column_name]
        )
)
| project SegmentId, ClusterRowCount, MatchedRawRows = matched_rows

说明:

  • array_all(bag_keys(Pattern), ...)会自动遍历Pattern中的所有特征列,匹配原始表对应列的值,无需手动编写每列的过滤条件,适配任意列数的表。
  • 最终结果中,MatchedRawRows列是一个动态数组,包含当前聚类对应的所有原始行数据;也可将其展开,用mv-expand MatchedRawRows将每个原始行单独成一条记录。

方案2:为每个原始行标记所属聚类ID

如果需要给每条原始行直接打上所属的聚类标签(支持一行属于多个聚类的场景),可以用mv-apply实现:

// 1. 先获取聚类结果
let cluster_info = YourTableName
| evaluate autocluster(0.01)
| project SegmentId, Pattern;

// 2. 为原始行匹配所属聚类
YourTableName
| mv-apply cluster = cluster_info on (
    where array_all(bag_keys(cluster.Pattern), column_name => 
        todynamic(YourTableName[column_name]) == cluster.Pattern[column_name]
    )
    | project ClusterSegmentId = cluster.SegmentId
)
| project-away cluster

说明:

  • 每条原始行如果匹配多个聚类,会生成多条记录,每条记录对应一个所属的ClusterSegmentId。
  • 若需仅保留最匹配的聚类(比如最大的聚类),可在cluster_info中按Count降序排序后,添加limit 1到mv-apply的子查询中。

注意事项

  • autocluster的threshold参数:取值范围0~1,值越小聚类数量越多,特征越细分;值越大聚类越聚合,数量越少。
  • 若原始表数据量极大,建议先通过where过滤出目标数据集再做聚类,提升性能。

内容的提问来源于stack exchange,提问作者user8561039

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 08:43:38