Fluentd正则多行解析器匹配失败问题求助
问题:Fluentd正则解析多行日志失败(在线正则工具可正常解析)
尝试用Fluentd的regex解析器处理多行日志时,在线正则编辑器能正常解析,但配置到td-agent.conf后提示模式不匹配,无有效日志被解析。
示例日志
2023-03-09T12:40:35.135000+00:00 IOD_SESS.monitor begin:2023-03-09T12:40:32.856 end:2023-03-09T12:40:33.134 interval:+000000000 00:00:00.278381000 monitored_sessions:0 killed_sessions:0 2023-03-09T12:40:35 (IOD_SESS.snap_blocked_sessions) begin 2023-03-09T12:40:35 (IOD_SESS.snap_blocked_sessions) end (rows:2) 2023-03-09T12:41:40.135000+00:00 IOD_SESS.monitor begin:2023-03-09T12:40:32.856 end:2023-03-09T12:40:33.134 interval:+000000000 00:00:00.278381000 monitored_sessions:0 killed_sessions:0 2023-03-09T12:41:40 (IOD_SESS.snap_blocked_sessions) begin 2023-03-09T12:41:40 (IOD_SESS.snap_blocked_sessions) end (rows:2) 2023-03-09T12:42:35.135000+00:00 IOD_SESS.monitor begin:2023-03-09T12:40:32.856 end:2023-03-09T12:40:33.134 interval:+000000000 00:00:00.278381000 monitored_sessions:0 killed_sessions:0 2023-03-09T12:42:35 (IOD_SESS.snap_blocked_sessions) begin 2023-03-09T12:42:35 (IOD_SESS.snap_blocked_sessions) end (rows:2)
当前td-agent.conf配置
<source> @type tail path /tmp/gg.txt read_from_head true tag db <parse> @type regexp expression /^(?<timestamp>)\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}.\d{6}\+\d{2}:\d{2}(?<message>.*(?:\n(?!\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}.\d{6}\+\d{2}:\d{2}).*)*)/gm </parse> </source> <match db> @type stdout </match>
错误输出
2023-03-12 23:12:31 +0530 [info]: #0 starting fluentd worker pid=42838 ppid=42836 worker=0 2023-03-12 23:12:31 +0530 [info]: #0 following tail of /tmp/gg.txt 2023-03-12 23:12:31 +0530 [warn]: #0 pattern not matched: "IOD_SESS.monitor begin:2023-03-09T12:40:32.856 end:2023-03-09T12:40:33.134 interval:+000000000 00:00:00.278381000 monitored_sessions:0 killed_sessions:0" 2023-03-12 23:12:31 +0530 [warn]: #0 pattern not matched: "2023-03-09T12:40:35 (IOD_SESS.snap_blocked_sessions) begin" 2023-03-12 23:12:31 +0530 [warn]: #0 pattern not matched: "2023-03-09T12:40:35 (IOD_SESS.snap_blocked_sessions) end (rows:2)" 2023-03-12 23:12:31 +0530 [warn]: #0 pattern not matched: "IOD_SESS.monitor begin:2023-03-09T12:40:32.856 end:2023-03-09T12:40:33.134 interval:+000000000 00:00:00.278381000 monitored_sessions:0 killed_sessions:0" 2023-03-12 23:12:31 +0530 [warn]: #0 pattern not matched: "2023-03-09T12:41:40 (IOD_SESS.snap_blocked_sessions) begin" 2023-03-12 23:12:31 +0530 [warn]: #0 pattern not matched: "2023-03-09T12:41:40 (IOD_SESS.snap_blocked_sessions) end (rows:2)" 2023-03-12 23:12:31 +0530 [warn]: #0 pattern not matched: "IOD_SESS.monitor begin:2023-03-09T12:40:32.856 end:2023-03-09T12:40:33.134 interval:+000000000 00:00:00.278381000 monitored_sessions:0 killed_sessions:0" 2023-03-12 23:12:31 +0530 [warn]: #0 pattern not matched: "2023-03-09T12:42:35 (IOD_SESS.snap_blocked_sessions) begin" 2023-03-12 23:12:31 +0530 [warn]: #0 pattern not matched: "2023-03-09T12:42:35 (IOD_SESS.snap_blocked_sessions) end (rows:2)" 2023-03-12 23:12:31.893575000 +0530 db: {"timestamp":"","message":""} 2023-03-12 23:12:31.893900000 +0530 db: {"timestamp":"","message":""} 2023-03-12 23:12:31.894040000 +0530 db: {"timestamp":"","message":""}
问题原因及解决方案
核心问题
- 正则分组错误:
(?<timestamp>)分组内未包含匹配逻辑,导致时间字段捕获为空。 - 多行模式不兼容:Fluentd的regex解析器不支持
gm修饰符,默认按行读取日志,无法直接匹配跨多行的日志块。 - 特殊字符未转义:正则中的
.未转义,会匹配任意字符,导致匹配逻辑混乱。
方案一:使用multiline解析器(推荐)
Fluentd的multiline解析器专门用于处理多行日志,逻辑更清晰:
<source> @type tail path /tmp/gg.txt read_from_head true tag db <parse> @type multiline format_firstline /^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}\.\d{6}\+\d{2}:\d{2}$/ format1 /^(?<timestamp>\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}\.\d{6}\+\d{2}:\d{2})$/ format2 /^(?<message>.*)$/ keep_time_key true time_key timestamp time_format %Y-%m-%dT%H:%M:%S.%N%:z </parse> </source> <match db> @type stdout </match>
方案二:修正regex解析器配置
如果坚持用regex解析器,需调整正则并启用多行参数:
<source> @type tail path /tmp/gg.txt read_from_head true tag db <parse> @type regexp expression /^(?<timestamp>\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}\.\d{6}\+\d{2}:\d{2})(?<message>(?:\n(?!\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}\.\d{6}\+\d{2}:\d{2}).*)*)/ multiline true keep_time_key true time_key timestamp time_format %Y-%m-%dT%H:%M:%S.%N%:z </parse> </source> <match db> @type stdout </match>
关键修正说明
- 将时间匹配逻辑移入
(?<timestamp>)分组,确保能捕获到时间值 - 转义正则中的
.为\.,精准匹配时间格式中的小数点 - 移除
gm修饰符,改用multiline true启用Fluentd的多行匹配逻辑 - 配置
time_key和time_format,让Fluentd正确解析日志自带的时间,避免被替换为接收时间
内容的提问来源于stack exchange,提问作者Hariprasad
相关产品推荐
相关产品推荐

