You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Fluentd正则多行解析器匹配失败问题求助

问题:Fluentd正则解析多行日志失败(在线正则工具可正常解析)

尝试用Fluentd的regex解析器处理多行日志时,在线正则编辑器能正常解析,但配置到td-agent.conf后提示模式不匹配,无有效日志被解析。

示例日志

2023-03-09T12:40:35.135000+00:00
IOD_SESS.monitor begin:2023-03-09T12:40:32.856 end:2023-03-09T12:40:33.134 interval:+000000000 00:00:00.278381000 monitored_sessions:0 killed_sessions:0
2023-03-09T12:40:35 (IOD_SESS.snap_blocked_sessions) begin
2023-03-09T12:40:35 (IOD_SESS.snap_blocked_sessions) end (rows:2)
2023-03-09T12:41:40.135000+00:00
IOD_SESS.monitor begin:2023-03-09T12:40:32.856 end:2023-03-09T12:40:33.134 interval:+000000000 00:00:00.278381000 monitored_sessions:0 killed_sessions:0
2023-03-09T12:41:40 (IOD_SESS.snap_blocked_sessions) begin
2023-03-09T12:41:40 (IOD_SESS.snap_blocked_sessions) end (rows:2)
2023-03-09T12:42:35.135000+00:00
IOD_SESS.monitor begin:2023-03-09T12:40:32.856 end:2023-03-09T12:40:33.134 interval:+000000000 00:00:00.278381000 monitored_sessions:0 killed_sessions:0
2023-03-09T12:42:35 (IOD_SESS.snap_blocked_sessions) begin
2023-03-09T12:42:35 (IOD_SESS.snap_blocked_sessions) end (rows:2)

当前td-agent.conf配置

<source>
  @type tail
  path /tmp/gg.txt
  read_from_head true
  tag db
  <parse>
  @type regexp
  expression /^(?<timestamp>)\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}.\d{6}\+\d{2}:\d{2}(?<message>.*(?:\n(?!\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}.\d{6}\+\d{2}:\d{2}).*)*)/gm
  </parse>
</source>

<match db>
  @type stdout
</match>

错误输出

2023-03-12 23:12:31 +0530 [info]: #0 starting fluentd worker pid=42838 ppid=42836 worker=0
2023-03-12 23:12:31 +0530 [info]: #0 following tail of /tmp/gg.txt
2023-03-12 23:12:31 +0530 [warn]: #0 pattern not matched: "IOD_SESS.monitor begin:2023-03-09T12:40:32.856 end:2023-03-09T12:40:33.134 interval:+000000000 00:00:00.278381000 monitored_sessions:0 killed_sessions:0"
2023-03-12 23:12:31 +0530 [warn]: #0 pattern not matched: "2023-03-09T12:40:35 (IOD_SESS.snap_blocked_sessions) begin"
2023-03-12 23:12:31 +0530 [warn]: #0 pattern not matched: "2023-03-09T12:40:35 (IOD_SESS.snap_blocked_sessions) end (rows:2)"
2023-03-12 23:12:31 +0530 [warn]: #0 pattern not matched: "IOD_SESS.monitor begin:2023-03-09T12:40:32.856 end:2023-03-09T12:40:33.134 interval:+000000000 00:00:00.278381000 monitored_sessions:0 killed_sessions:0"
2023-03-12 23:12:31 +0530 [warn]: #0 pattern not matched: "2023-03-09T12:41:40 (IOD_SESS.snap_blocked_sessions) begin"
2023-03-12 23:12:31 +0530 [warn]: #0 pattern not matched: "2023-03-09T12:41:40 (IOD_SESS.snap_blocked_sessions) end (rows:2)"
2023-03-12 23:12:31 +0530 [warn]: #0 pattern not matched: "IOD_SESS.monitor begin:2023-03-09T12:40:32.856 end:2023-03-09T12:40:33.134 interval:+000000000 00:00:00.278381000 monitored_sessions:0 killed_sessions:0"
2023-03-12 23:12:31 +0530 [warn]: #0 pattern not matched: "2023-03-09T12:42:35 (IOD_SESS.snap_blocked_sessions) begin"
2023-03-12 23:12:31 +0530 [warn]: #0 pattern not matched: "2023-03-09T12:42:35 (IOD_SESS.snap_blocked_sessions) end (rows:2)"
2023-03-12 23:12:31.893575000 +0530 db: {"timestamp":"","message":""}
2023-03-12 23:12:31.893900000 +0530 db: {"timestamp":"","message":""}
2023-03-12 23:12:31.894040000 +0530 db: {"timestamp":"","message":""}

问题原因及解决方案

核心问题

  1. 正则分组错误:(?<timestamp>)分组内未包含匹配逻辑,导致时间字段捕获为空。
  2. 多行模式不兼容:Fluentd的regex解析器不支持gm修饰符,默认按行读取日志,无法直接匹配跨多行的日志块。
  3. 特殊字符未转义:正则中的.未转义,会匹配任意字符,导致匹配逻辑混乱。

方案一:使用multiline解析器(推荐)

Fluentd的multiline解析器专门用于处理多行日志,逻辑更清晰:

<source>
  @type tail
  path /tmp/gg.txt
  read_from_head true
  tag db
  <parse>
    @type multiline
    format_firstline /^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}\.\d{6}\+\d{2}:\d{2}$/
    format1 /^(?<timestamp>\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}\.\d{6}\+\d{2}:\d{2})$/
    format2 /^(?<message>.*)$/
    keep_time_key true
    time_key timestamp
    time_format %Y-%m-%dT%H:%M:%S.%N%:z
  </parse>
</source>

<match db>
  @type stdout
</match>

方案二:修正regex解析器配置

如果坚持用regex解析器,需调整正则并启用多行参数:

<source>
  @type tail
  path /tmp/gg.txt
  read_from_head true
  tag db
  <parse>
    @type regexp
    expression /^(?<timestamp>\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}\.\d{6}\+\d{2}:\d{2})(?<message>(?:\n(?!\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}\.\d{6}\+\d{2}:\d{2}).*)*)/
    multiline true
    keep_time_key true
    time_key timestamp
    time_format %Y-%m-%dT%H:%M:%S.%N%:z
  </parse>
</source>

<match db>
  @type stdout
</match>

关键修正说明

  • 将时间匹配逻辑移入(?<timestamp>)分组,确保能捕获到时间值
  • 转义正则中的.为\.,精准匹配时间格式中的小数点
  • 移除gm修饰符,改用multiline true启用Fluentd的多行匹配逻辑
  • 配置time_key和time_format,让Fluentd正确解析日志自带的时间,避免被替换为接收时间

内容的提问来源于stack exchange,提问作者Hariprasad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 00:24:58