如何对指定时间段内所有事件的error.count数值求和并配置告警
调整后的查询语句
index=prod_service service.error | bin span=1h _time | stats sum(eval(tonumber('error.count'))) as total_error by _time | where total_error > 10
修改说明
- 修正数据源筛选规则:把原语句中错误的
index=prod-service、service.count替换为日志实际对应的index=prod_service、service.error,确保匹配到所有带error.count的事件(含error.count为0的事件) - 新增按小时分组逻辑:
bin span=1h _time会将所有事件按1小时的时间窗口切割分组 - 优化聚合逻辑:由于日志中
error.count是带双引号的字符串格式,先通过tonumber()转成数值类型再求和,最终将每小时的求和结果命名为total_error - 修正触发条件:将过滤规则调整为判断每小时的
total_error大于10,只要查询返回结果即可触发告警
如果你的日志平台已经自动把error.count识别为数值类型,可简化聚合部分:
index=prod_service service.error | bin span=1h _time | stats sum('error.count') as total_error by _time | where total_error > 10
内容的提问来源于stack exchange,提问作者Yoey
相关产品推荐
相关产品推荐

