Logstash聚合拆分的Java堆栈跟踪Syslog日志问题求助
Hey there, let's get those split Java stack traces merged into single, usable events. The issues with your current setup are twofold:
- Your Aggregate filter uses
%{timestamp}as thetask_id—this is too vague, since multiple errors can fire at the exact same time, leading to incorrect grouping. - If you tried the Multiline plugin before, you likely used it in the Filter stage instead of the Input stage (Multiline is designed to handle raw input lines before any parsing happens).
Option 1: Use Multiline at the Input Stage (Recommended)
Multiline is the simplest, most efficient way to handle cross-line logs like Java stack traces. Since your Syslog messages are sent line-by-line, merging them at the input level avoids messy filter logic later.
Example Input Configuration
input { syslog { port => 514 type => "syslog" # Configure Multiline to merge stack trace lines codec => multiline { pattern => "^%{SYSLOGTIMESTAMP}" # Match lines starting with a Syslog timestamp (new event trigger) negate => true # Reverse the match: lines NOT starting with timestamp belong to the previous event what => "previous" # Append matching lines to the prior event max_lines => 1000 # Prevent memory bloat by capping merged lines timeout => 3 # Force output after 3s if no more lines arrive } } }
Updated Filter Configuration
Now that Multiline has merged the full stack into a single message field, your filter logic can stay mostly the same—just remove the broken Aggregate block:
filter { if [type] == "syslog" { grok { match => { "message" => "%{SYSLOGTIMESTAMP:timestamp} %{SYSLOGHOST:hostname} %{DATA:application}(?:\[%{POSINT:pid}\])?: %{GREEDYDATA:raw_message}" } } date { match => [ "timestamp", "MMM d HH:mm:ss", "MMM dd HH:mm:ss"] timezone => "Europe/Zurich" } if [hostname] == "app-name.example.com" { json { source => "raw_message" } mutate { add_field => {"[@metadata][es_index]" => "app-name-devl-logs-%{+YYYY.MM.dd}"} add_field => {"[@metadata][document_id]" => "%{[req_id]}" } remove_field => ["raw_message", "host", "timestamp"] } } else { fingerprint { source => "message" target => "[@metadata][document_id]" method => "MURMUR3" } mutate { add_field => {"[@metadata][es_index]" => "app-name-devl-logs-%{+YYYY.MM.dd}"} remove_field => ["message", "host", "timestamp"] } } } }
Key Notes:
- The
patterntargets Syslog timestamps because every new error starts with one—stack trace lines don't, so they get appended to the prior event. - The
timeoutensures you don't hang waiting for missing stack lines; after 3 seconds, Logstash outputs whatever it has.
Option 2: Use Aggregate Filter in the Filter Stage (For Complex Scenarios)
If you absolutely need to handle aggregation in the Filter stage (e.g., you're parsing first before merging), you need a unique task_id to group stack lines correctly. Combine timestamps, hostname, application, and the error's starting line to avoid grouping unrelated errors.
Corrected Aggregate Filter Configuration
filter { if [type] == "syslog" { grok { match => { "message" => "%{SYSLOGTIMESTAMP:timestamp} %{SYSLOGHOST:hostname} %{DATA:application}(?:\[%{POSINT:pid}\])?: %{GREEDYDATA:raw_message}" } } date { match => [ "timestamp", "MMM d HH:mm:ss", "MMM dd HH:mm:ss"] timezone => "Europe/Zurich" } if [application] == "app-name" { # Identify the start of a stack trace (lines with "Exception:") if "Exception:" in [raw_message] { # Generate a unique task ID using timestamp, hostname, app, and the error line fingerprint { source => ["timestamp", "hostname", "application", "raw_message"] target => "[@metadata][stack_task_id]" method => "MURMUR3" } # Initialize the aggregate map with the first line of the error aggregate { task_id => "%{[@metadata][stack_task_id]}" code => " map['full_stack_trace'] = event.get('raw_message') map['@timestamp'] = event.get('@timestamp') map['hostname'] = event.get('hostname') map['application'] = event.get('application') " push_previous_map_as_event => false timeout => 3 } # Drop the initial line—we'll output the full stack later drop {} } else { # Append stack trace lines to the existing aggregate aggregate { # Use the same task ID pattern as the starting line task_id => "%{timestamp}_%{hostname}_%{application}" code => " map['full_stack_trace'] += '\n' + event.get('raw_message') " push_previous_map_as_event => false timeout => 3 } # Drop intermediate stack lines drop {} } # Output the full aggregated event when the timeout hits aggregate { task_id => "%{[@metadata][stack_task_id]}" code => " event.set('full_stack_trace', map['full_stack_trace']) event.set('@timestamp', map['@timestamp']) event.set('hostname', map['hostname']) event.set('application', map['application']) " push_previous_map_as_event => true timeout => 3 } } # Clean up unused fields mutate { remove_field => ["host", "timestamp"] } # Add your existing JSON parsing/index logic here, adjusted to use `full_stack_trace` instead of `raw_message` } }
Key Notes:
- The
task_idcombines multiple unique identifiers to ensure only related stack lines are grouped. - We drop intermediate lines so only the full merged stack trace is sent to Elasticsearch.
Testing Tips
- Start with the Multiline option—it's far simpler and less error-prone for this use case.
- Validate your config with
logstash -f your-config.conf --config.test_and_exitto catch syntax errors. - Run Logstash in verbose mode (
logstash -f your-config.conf --verbose) to watch how lines are merged in real time.
内容的提问来源于stack exchange,提问作者V.Widmer

