Logstash过滤器中多grok模式的正确配置方法咨询
Hey there! Let's figure out how to correctly set up multiple grok patterns in your Logstash filters, and address why your current setup isn't working as expected.
First: Let's Compare Your Two Configurations
Configuration 1: Multiple match entries in one grok block
else if [pipeline] == "tomcat_all" { grok { match => [ "message", "%{MONTH}%{SPACE}%{MONTHDAY},%{SPACE}%{YEAR}%{SPACE}%{HOUR}:?%{MINUTE}(?::?%{SECOND})%{SPACE}(?:AM|PM)%{SPACE}%{NOTSPACE:class}%{SPACE}%{NOTSPACE:type_log}%{SPACE}%{WORD:loglevel}:%{SPACE}%{GREEDYDATA:log_text}" ] match => [ "message", "%{TIME:timestamp}%{SPACE}\|-%{WORD:loglevel}%{SPACE}in%{SPACE}%{NOTSPACE:class}%{SPACE}%{GREEDYDATA:log_text}" ] } }
Configuration 2: Separate grok blocks
else if [pipeline] == "123" { grok { match => [ "message", "%{MONTH}%{SPACE}%{MONTHDAY},%{SPACE}%{YEAR}%{SPACE}%{HOUR}:?%{MINUTE}(?::?%{SECOND})%{SPACE}(?:AM|PM)%{SPACE}%{NOTSPACE:class}%{SPACE}%{NOTSPACE:type_log}%{SPACE}%{WORD:loglevel}:%{SPACE}%{GREEDYDATA:log_text}" ] } grok { match => [ "message", "%{TIME:timestamp}%{SPACE}\|-%{WORD:loglevel}%{SPACE}in%{SPACE}%{NOTSPACE:class}%{SPACE}%{GREEDYDATA:log_text}" ] } }
Which One Is Valid? (And Why Yours Isn't Working)
Neither of these is the correct way to set up multiple grok patterns for the same field, and here's why:
Configuration 1 Problem:
When you define multiplematchentries in a singlegrokblock, Logstash will overwrite the firstmatchwith the second one. That means only your second pattern will ever be tried—your first pattern is completely ignored. That's why your parsing isn't working for logs that should match the first pattern.Configuration 2 Problem:
Using separategrokblocks means Logstash will run both grok filters, regardless of whether the first one matched. Even if the first grok successfully parses the log, the second one will still try to parse the samemessagefield. If the second pattern doesn't match, it might add_grokparsefailuretags unnecessarily. If it does match (unlikely for the same log line), it could overwrite fields you already parsed.
The Correct Way to Configure Multiple Grok Patterns
To have Logstash try multiple patterns for the same field (and stop when it finds a match), you need to put all your patterns into a single match array for the field. Here's how to fix Configuration 1:
else if [pipeline] == "tomcat_all" { grok { match => [ "message", "%{MONTH}%{SPACE}%{MONTHDAY},%{SPACE}%{YEAR}%{SPACE}%{HOUR}:?%{MINUTE}(?::?%{SECOND})%{SPACE}(?:AM|PM)%{SPACE}%{NOTSPACE:class}%{SPACE}%{NOTSPACE:type_log}%{SPACE}%{WORD:loglevel}:%{SPACE}%{GREEDYDATA:log_text}", "%{TIME:timestamp}%{SPACE}\|-%{WORD:loglevel}%{SPACE}in%{SPACE}%{NOTSPACE:class}%{SPACE}%{GREEDYDATA:log_text}" ] # Optional: Keep this as true (default) to stop trying patterns once one matches break_on_match => true } }
Alternatively, you can use the hash syntax which is equally valid:
grok { match => { "message" => [ "pattern1", "pattern2" ] } }
Key Notes:
break_on_match => true(default behavior) tells Logstash to stop trying further patterns once one matches. If you want to try all patterns even after a match (rare case), set this tofalse.- Always test your grok patterns first! Use tools like the built-in Logstash grok debugger to verify each pattern works with your log lines before adding them to your config.
- If you still see
_grokparsefailuretags, check your log lines against each pattern—there might be subtle differences (like extra spaces, special characters) that your patterns aren't accounting for.
Why Your Logstash Started Without Errors
Logstash's configuration parser doesn't throw errors for duplicate match entries or multiple grok blocks—it just processes them as per its rules (overwriting or running sequentially). That's why your instance started fine, but the parsing didn't work as expected.
内容的提问来源于stack exchange,提问作者Dennis

