You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Logstash中Grok模式优化求助:高CPU占用问题

Optimizing Logstash Grok Patterns for High-Volume Log Processing

Let's tackle your CPU spike issue head-on. The problem here stems from two main factors: your dual Grok match rules forcing every log line to be tested twice, and overuse of broad, resource-heavy patterns like DATA and GREEDYDATA. With 550k events to process in 5 minutes, these inefficiencies quickly push CPU to 100%. Here's how to fix it:

1. Merge Your Two Grok Patterns Into One

Your two log formats only differ by the optional trailing ,'{jobSpecificMetaData}' field. Instead of making Logstash test each line against two separate patterns, we can make that final segment optional with a non-capturing group. This cuts the number of pattern tests per log line in half.

Revised Single Grok Pattern:

grok {
  # Optional: Define custom patterns first (more on this below)
  patterns_dir => ["./patterns"]
  match => { "message" => "%{TIMESTAMP_ISO8601:created_timestamp},%{UUID:request_id},%{DATA:tenant},%{DATA:username},%{DATA:job_code},%{DATA:stepname},%{MY_TIMESTAMP:quartz_trigger_timestamp},%{DATA:execution_level},%{DATA:facility_name},%{DATA:channel_code},%{WORD:status},%{NUMBER:current_step_time_ms:int},%{NUMBER:total_time_ms:int},'%{DATA:error_message}',%{DATA:tenant_mode},%{DATA:channel_src_code}(?:,'%{GREEDYDATA:jobSpecificMetaData}')?" }
  break_on_match => true # Ensure we stop matching once we find a hit
}

2. Replace Broad Patterns With Precise Custom Regex

DATA (.*?) and GREEDYDATA (.*) are flexible but computationally expensive because they require extensive regex backtracking. Swap them for targeted patterns where possible:

  • created_timestamp: Use %{TIMESTAMP_ISO8601} instead of DATA (matches your ISO 8601 timestamp format perfectly).
  • request_id: Use %{UUID} instead of DATA (matches your UUID format exactly).
  • status: Use %{WORD} instead of DATA (values like SUCCESSFUL are single words).
  • current_step_time_ms/total_time_ms: Use %{NUMBER:int} to directly parse as integers, avoiding post-processing with mutate.
  • quartz_trigger_timestamp: Create a custom pattern for that specific date format (Wed Jan 11 15:16:13 IST 2023). Add this to a custom patterns file (e.g., ./patterns/custom):
    MY_TIMESTAMP %{DAY} %{MONTH} %{MONTHDAY} %{TIME} %{WORD} %{YEAR}
    

3. Use Dissect + Grok for Faster Processing

For comma-separated logs, Dissect is far more efficient than Grok because it works on fixed delimiters instead of regex. Use it to split the log into fields first, then only apply Grok to the fields that need complex pattern matching. This reduces regex processing to just the necessary fields.

Example Dissect + Grok Setup:

# First split the log line with Dissect (super fast!)
dissect {
  mapping => {
    "message" => "%{created_timestamp},%{request_id},%{tenant},%{username},%{job_code},%{stepname},%{quartz_trigger_timestamp},%{execution_level},%{facility_name},%{channel_code},%{status},%{current_step_time_ms},%{total_time_ms},'%{error_message}',%{tenant_mode},%{channel_src_code}%{?jobSpecificMetaData}"
  }
}

# Clean up the optional jobSpecificMetaData field if it exists
if [jobSpecificMetaData] {
  mutate {
    gsub => ["jobSpecificMetaData", "^,'(.*)'$", "\1"]
  }
}

# Convert numeric fields to integers
mutate {
  convert => {
    "current_step_time_ms" => "integer"
    "total_time_ms" => "integer"
  }
}

# Only use Grok for the non-standard timestamp format
grok {
  patterns_dir => ["./patterns"]
  match => { "quartz_trigger_timestamp" => "%{MY_TIMESTAMP:quartz_trigger_timestamp}" }
}

4. Tune Logstash JVM Settings

High CPU usage can also be exacerbated by insufficient heap memory, leading to frequent garbage collection. Adjust your jvm.options file to allocate enough memory (based on your server's resources):

-Xms4g
-Xmx4g

Start with 4GB and adjust up if needed (aim for 50-70% of available RAM).

Why This Works

  • Merging patterns eliminates redundant regex checks per log line.
  • Precise patterns reduce regex backtracking, which is the #1 cause of Grok CPU bloat.
  • Dissect handles the simple delimiter-based splitting, leaving Grok to only do the complex work it's good at.

With these changes, you should see a dramatic drop in CPU usage while maintaining accurate log parsing.

内容的提问来源于stack exchange,提问作者Gajendar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 18:35:29