如何用Grok提取日志message字段指定内容?修复Dissect输出 字符问题
Hey there, let's tackle your two Logstash questions one by one!
Absolutely! Grok is built specifically for parsing unstructured log data into structured fields, so it's perfect for pulling out timestamps and specific terms like OrderCreated. The key is to define a Grok pattern that matches your actual log format (the COMBINEDAPACHELOG pattern you used earlier is only for Apache web logs, which doesn't fit your custom log structure).
Looking at your sample log line:
General 2018-05-17 15:47:33.149 : StatusInformationSomeData.Unsubscribe()
You can create a custom Grok pattern tailored to this structure. For example:
filter { grok { # Match log level, timestamp, and message content match => ["message", "%{WORD:log_level} %{TIMESTAMP_ISO8601:timestamp} : %{DATA:message_content}"] } # If you want to flag or extract logs containing "OrderCreated" specifically if "OrderCreated" in [message_content] { # Add a tag to easily identify these events mutate { add_tag => ["order_created_event"] } # Optional: Extract the term as a dedicated field if needed grok { match => ["message_content", "%{DATA}OrderCreated%{DATA}"] add_field => ["event_type", "OrderCreated"] } } # Add your existing mutate rules to remove unwanted fields here mutate { remove_field => ["@timestamp", "@version", "host", "path"] } }
To refine this pattern, use Logstash's built-in Grok Debugger (accessible at http://localhost:9600/_node/plugins/pipeline/filters/grok/grokdebug if Logstash is running locally) — paste your sample log line and tweak the pattern until it matches all fields correctly.
Those \r characters are Windows-style carriage returns (since Windows uses \r\n for newlines, while Unix uses just \n). There are a few straightforward fixes:
Option 1: Clean up the message field before Dissect
Add a mutate filter to strip out all \r characters before passing the message to Dissect. This is the most reliable approach:
filter { # Remove all carriage returns from the message field first mutate { gsub => ["message", "\r", ""] } dissect { mapping => { "message" => "%{ts} %{+ts} %{+ts} %{src} %{} : %{msg}" } } # Filter out empty lines (like the standalone "\r" entries in your output) if [message] =~ /^\s*$/ { drop {} } # Your existing mutate rules to remove unwanted fields mutate { remove_field => ["@timestamp", "@version", "host", "path", "src", "pid", "prog"] } }
Option 2: Adjust the Dissect pattern to ignore trailing \r
If \r only appears at the end of log lines, you can add a placeholder to the Dissect pattern to consume it:
dissect { mapping => { "message" => "%{ts} %{+ts} %{+ts} %{src} %{} : %{msg}%{}" } }
This works for trailing \r, but Option 1 is better if \r might appear elsewhere in the log.
Bonus: Filter out empty lines
Your output shows a standalone "\r" entry — adding the if [message] =~ /^\s*$/ { drop {} } rule will discard these empty lines entirely, so they don't end up in Elasticsearch or your stdout.
内容的提问来源于stack exchange,提问作者Daniel

