You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

同一第三方设备多格式日志的Logstash配置及Grok解析方案咨询

Handling Variable Log Formats from a Third-Party Device in Logstash

I’ve dealt with this exact headache before—third-party devices spitting out totally different log structures based on the commands users run, like your examples:

日志A:Jun 2 16:45:49 host-A; rule='a', type='a', pattern='a', actions_taken='a', event_data='a'
日志B:Jun 2 16:52:19 host-A; event='bbb', user='sss', com='111'

Trying to parse these with a single Grok pattern is impossible, but we can fix this with smart Logstash configuration and targeted Grok rules. Let’s break down both your questions.

1. Configuring Logstash .conf Files for Mixed Log Types

The core solution here is using conditional logic in the filter section to route each log type to its own parsing rules. First, extract any shared fields all logs have (like timestamp and host), then use if/else if blocks to match unique identifiers for each log variant.

Here’s a practical, tested configuration example:

input {
  # Replace with your actual input (file, Filebeat, etc.)
  file {
    path => "/var/log/third-party-device/*.log"
    start_position => "beginning"
  }
}

filter {
  # Step 1: Parse the shared prefix all logs have
  grok {
    match => { "message" => "%{MONTH:month} %{MONTHDAY:day} %{TIME:time} %{HOSTNAME:host}; " }
    add_tag => ["parsed_common_fields"]
  }

  # Step 2: Handle Log Type A (identifiable by the 'rule' field)
  if "parsed_common_fields" in [tags] and [message] =~ /rule='/ {
    grok {
      match => { "message" => "%{MONTH} %{MONTHDAY} %{TIME} %{HOSTNAME}; rule='%{DATA:rule}', type='%{DATA:type}', pattern='%{DATA:pattern}', actions_taken='%{DATA:actions_taken}', event_data='%{DATA:event_data}'" }
      add_tag => ["log_type_a"]
    }
  }

  # Step 3: Handle Log Type B (identifiable by the 'event' field)
  else if "parsed_common_fields" in [tags] and [message] =~ /event='/ {
    grok {
      match => { "message" => "%{MONTH} %{MONTHDAY} %{TIME} %{HOSTNAME}; event='%{DATA:event}', user='%{DATA:user}', com='%{DATA:com}'" }
      add_tag => ["log_type_b"]
    }
  }

  # Optional: Catch unrecognized logs to avoid data loss
  else {
    add_tag => ["unparsed_log"]
    # You can add a debug field here if needed: add_field => { "debug_message" => "Unrecognized log format: %{message}" }
  }
}

output {
  # Replace with your actual output (Elasticsearch, stdout, etc.)
  elasticsearch {
    hosts => ["http://localhost:9200"]
    index => "device-logs-%{+YYYY.MM.dd}"
  }
  stdout { codec => rubydebug }
}

Key notes for this setup:

  • Parsing shared fields first avoids repeating redundant Grok logic across each log type.
  • The regex checks (=~ /rule='/) are simple, reliable ways to identify log variants—adjust them if your logs have more unique markers.
  • Tags make it easy to track which parsing path each log took, simplifying troubleshooting later.

2. Configuring the Grok Plugin for Inconsistent Fields

There are three effective approaches to Grok parsing for this scenario, depending on your preferences and log volume:

Approach 1: Targeted Patterns with Conditionals (as above)

This is the most reliable method. Create a unique Grok pattern for each log type, then use Logstash conditionals to apply the right pattern to the right log. It’s easy to maintain and avoids accidental mis-parsing.

Approach 2: Multi-Pattern Matching in a Single Grok Block

If you want a more condensed setup, you can pass an array of patterns to the match parameter. Grok will try each pattern in order until one matches:

filter {
  grok {
    match => {
      "message" => [
        "%{MONTH} %{MONTHDAY} %{TIME} %{HOSTNAME}; rule='%{DATA:rule}', type='%{DATA:type}', pattern='%{DATA:pattern}', actions_taken='%{DATA:actions_taken}', event_data='%{DATA:event_data}'",
        "%{MONTH} %{MONTHDAY} %{TIME} %{HOSTNAME}; event='%{DATA:event}', user='%{DATA:user}', com='%{DATA:com}'"
      ]
    }
    add_tag => ["grok_parsed"]
  }
}

Caveat: Only use this if your log types are distinct enough that Grok won’t accidentally match the wrong pattern. If there’s any overlap in structure, stick with conditional logic.

Approach 3: Combine Dissect + Grok for Performance

For logs with consistent separators (like ; and , in your examples), use the dissect filter first to split logs into chunks—it’s faster than Grok for simple splitting. Then parse each chunk with targeted Grok rules:

filter {
  dissect {
    mapping => { "message" => "%{timestamp} %{host}; %{log_details}" }
  }

  if [log_details] =~ /rule='/ {
    grok {
      match => { "log_details" => "rule='%{DATA:rule}', type='%{DATA:type}', pattern='%{DATA:pattern}', actions_taken='%{DATA:actions_taken}', event_data='%{DATA:event_data}'" }
    }
  } else if [log_details] =~ /event='/ {
    grok {
      match => { "log_details" => "event='%{DATA:event}', user='%{DATA:user}', com='%{DATA:com}'" }
    }
  }
}

This is ideal if you’re dealing with high log volumes and want to optimize parsing speed.

内容的提问来源于stack exchange,提问作者disney82231

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 18:52:48