You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Kusto执行summarize时如何过滤每组step4结束后的冗余日志记录

实现方案

这个逻辑完全可以实现,核心思路是先预先计算每个group的有效时间边界,再用边界过滤数据后执行聚合,完全可以避开step4之后的无效记录。

具体实现步骤

  • 第一步:先统计每个group的有效时间范围,即该组step1的发生时间、step4的发生时间,生成每个group的边界表
  • 第二步:将边界表与原日志表按group字段关联,过滤出时间落在有效区间内的记录
  • 第三步:在过滤后的有效数据集上,直接写你原本需要的summarize聚合逻辑即可

下面以Kusto查询语法为例给出示例代码,其他SQL类查询引擎逻辑完全通用:

// 第一步:生成每个group的有效时间边界
let group_time_boundary = 你的原日志表
| where step in ("step1", "step4")
| summarize 
    step1_ts = minif(timestamp, step == "step1"),
    step4_ts = minif(timestamp, step == "step4")
  by group
// 过滤掉没有走完step1到step4的异常分组,可根据你的业务需求调整
| where isnotnull(step1_ts) and isnotnull(step4_ts);

// 第二步:关联边界过滤+执行聚合
你的原日志表
| lookup group_time_boundary on group
| where timestamp between (step1_ts .. step4_ts)
// 第三步:这里写你原本需要的聚合逻辑即可,所有统计都只会用到有效区间内的数据
| summarize 
    avg_step3_percent = avgif(percent_value, step == "step3"),
    total_cost = step4_ts - step1_ts
  by group

注意事项

如果你的业务场景中单个group存在多轮step1到step4的完整流程,需要先对同group的多轮流程做会话拆分,再按单轮流程计算边界即可,整体逻辑不变。

内容的提问来源于stack exchange,提问作者Gene Parmesan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 19:18:02