You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否用Logstash aggregate插件合并多容器的多行Docker日志?

Solution: Merging Docker Multi-Line Logs with Logstash Aggregate Filter

Absolutely, you can use Logstash's aggregate filter to solve this multi-line log merging problem—and it handles interleaved logs from different ContainerIDs perfectly, as long as you configure it correctly. Let's break down how to set this up and why it works even when logs are mixed.

How It Works

The aggregate filter groups events by a unique identifier (in your case, ContainerID), so logs from separate containers never interfere with each other. We'll:

  1. Identify which lines start a new log entry (those with your PREFIX | marker)
  2. Group lines by ContainerID
  3. Append continuation lines to the initial log entry
  4. Output the full merged log once no more continuation lines arrive (via a timeout)

Step-by-Step Configuration

Here's a complete Logstash filter configuration tailored to your use case:

filter {
  # First, identify start lines and split content
  grok {
    # Match lines that start with your prefix
    match => { "message" => "^PREFIX \| %{GREEDYDATA:message_part}" }
    # Tag start-of-log lines
    add_tag => ["start_of_log"]
    # Tag lines that don't match (these are continuations)
    tag_on_failure => ["continuation"]
  }

  # Core aggregation logic
  aggregate {
    # Group events by ContainerID—this ensures isolation between containers
    task_id => "%{ContainerID}"
    code => "
      if 'start_of_log' in event.get('tags')
        # Initialize a new aggregated message for this container
        map['full_message'] = event.get('message_part')
        # Preserve key fields from the original event
        map['ContainerID'] = event.get('ContainerID')
        # Cancel the original event so we don't output it raw
        event.cancel()
      else
        # Append continuation lines to the aggregated message
        if map['full_message']
          map['full_message'] += '\n' + event.get('message')
        end
        # Cancel the continuation line event
        event.cancel()
      end
    "
    # Output the aggregated log when no new lines arrive for this container
    push_map_as_event_on_timeout => true
    # Adjust timeout based on your log frequency (30s is a safe default)
    timeout => 30
    # Tag aggregated events for easy filtering later
    timeout_tags => ["aggregated"]
  }

  # Clean up fields and tags for the final output
  mutate {
    # Replace the original message with our merged full message
    rename => { "full_message" => "message" }
    # Remove temporary fields
    remove_field => ["message_part"]
    # Clean up tags we used for processing
    remove_tag => ["start_of_log", "continuation"]
  }
}

Handling Interleaved Logs from Different Containers

This setup won't have issues with interleaved logs because the task_id => "%{ContainerID}" parameter creates a separate "task" (a memory map) for each unique container. When a log line from Container A comes in, it only updates the map for Container A—lines from Container B will update their own independent map. The timeout ensures each container's merged log is output once it stops receiving new lines, regardless of other containers' activity.

Key Notes for Success

  • Verify ContainerID Field: Make sure the field name matches what Docker's GELF driver sends (it might be container_id instead of ContainerID—check your raw events to confirm).
  • Adjust Timeout: If your containers generate long-running multi-line logs (e.g., stack traces), increase the timeout value to avoid splitting logs prematurely. If logs are frequent, you can lower it to reduce memory usage.
  • Tweak the Grok Pattern: If your PREFIX isn't static, update the grok match to match your actual log start pattern (e.g., ^%{WORD:prefix} \| %{GREEDYDATA:message_part} for dynamic prefixes).
  • Monitor Memory: If you have hundreds/thousands of containers, the aggregate filter will hold a map for each active container until timeout. Ensure your Logstash instance has enough memory, and set a reasonable timeout to prevent memory leaks.

内容的提问来源于stack exchange,提问作者herm

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:12:51