Grok新手求助:Graylog2环境下拆分日志剩余内容
Hey there! It looks like you're off to a great start with your Syslog Grok pattern—nice job getting the timestamp, logsource, and prog parts sorted out. Let's tackle that remaining hyphen-separated chunk step by step.
First, let's clarify the structure of the remaining log content: each field is clearly separated by - (space + hyphen + space). We can build targeted patterns for each segment and append them to your existing SYSLOGBASE.
Step 1: Map out the remaining fields
Let's break down the sample log's unsplit portion to identify each field:
2017-12-20 18:46:00 068611 +0100 - server-04.location-2 - 14086/0x00007f093b7fe700 - processname/SIMServer - 00000000173d9b6b - info - work: You have 2 connections running
Here's what each segment represents:
- Script-specific timestamp (with millis and timezone):
2017-12-20 18:46:00 068611 +0100 - Server location identifier:
server-04.location-2 - PID/hex thread ID pair:
14086/0x00007f093b7fe700 - Process category/service name:
processname/SIMServer - Alphanumeric trace ID:
00000000173d9b6b - Log severity level:
info - Human-readable log message:
work: You have 2 connections running
Step 2: Build the extended Grok pattern
We'll add patterns for each of these fields to your existing SYSLOGBASE, using the - delimiter to connect them:
SYSLOGBASE %{SYSLOGTIMESTAMP:timestamp} (?:%{SYSLOGFACILITY} )?%{SYSLOGHOST:logsource} %{SYSLOGPROG}: %{DATA:script_timestamp} - %{DATA:location} - %{NUMBER:pid}/%{DATA:thread_id} - %{DATA:process_category}/%{DATA:service_name} - %{DATA:trace_id} - %{LOGLEVEL:loglevel} - %{GREEDYDATA:message}
Let's explain the new pattern components:
%{DATA:script_timestamp}: Captures the full custom timestamp string (since it doesn't perfectly match standardTIMESTAMP_ISO8601; if you want to split this into sub-fields, use%{DATE:script_date} %{TIME:script_time} %{NUMBER:script_millis} %{INT:script_timezone}instead)%{DATA:location}: Matches the server location string (supports dots and hyphens)%{NUMBER:pid}/%{DATA:thread_id}: Splits the numeric PID and hex thread ID into separate, usable fields%{DATA:process_category}/%{DATA:service_name}: Separates the parent process category and specific service name%{DATA:trace_id}: Captures the alphanumeric trace ID (swap with%{HEX:trace_id}if you want to enforce hex-only formatting)%{LOGLEVEL:loglevel}: Uses Graylog's built-in pattern to match standard severity levels likeinfo,warn, orerror%{GREEDYDATA:message}: Safely captures everything after the last delimiter, even if the message contains spaces or hyphens (better thanDATAfor unstructured message text)
Step 3: Test and refine
- Use Graylog's built-in Grok Debugger (under System > Inputs > Manage Inputs > Select your input > Grok Debugger) to test this pattern against your sample log line. It'll show you exactly which fields are being captured.
- If you notice fields are being misaligned, double-check that the delimiter in your pattern matches exactly (
-, with spaces on both sides of the hyphen). - For stricter validation, replace broad
DATApatterns with regex where needed—for example, use%{WORD:process_category}if that segment never contains special characters.
内容的提问来源于stack exchange,提问作者user2835733

