Logstash中Grok模式优化求助:高CPU占用问题
Let's tackle your CPU spike issue head-on. The problem here stems from two main factors: your dual Grok match rules forcing every log line to be tested twice, and overuse of broad, resource-heavy patterns like DATA and GREEDYDATA. With 550k events to process in 5 minutes, these inefficiencies quickly push CPU to 100%. Here's how to fix it:
1. Merge Your Two Grok Patterns Into One
Your two log formats only differ by the optional trailing ,'{jobSpecificMetaData}' field. Instead of making Logstash test each line against two separate patterns, we can make that final segment optional with a non-capturing group. This cuts the number of pattern tests per log line in half.
Revised Single Grok Pattern:
grok { # Optional: Define custom patterns first (more on this below) patterns_dir => ["./patterns"] match => { "message" => "%{TIMESTAMP_ISO8601:created_timestamp},%{UUID:request_id},%{DATA:tenant},%{DATA:username},%{DATA:job_code},%{DATA:stepname},%{MY_TIMESTAMP:quartz_trigger_timestamp},%{DATA:execution_level},%{DATA:facility_name},%{DATA:channel_code},%{WORD:status},%{NUMBER:current_step_time_ms:int},%{NUMBER:total_time_ms:int},'%{DATA:error_message}',%{DATA:tenant_mode},%{DATA:channel_src_code}(?:,'%{GREEDYDATA:jobSpecificMetaData}')?" } break_on_match => true # Ensure we stop matching once we find a hit }
2. Replace Broad Patterns With Precise Custom Regex
DATA (.*?) and GREEDYDATA (.*) are flexible but computationally expensive because they require extensive regex backtracking. Swap them for targeted patterns where possible:
created_timestamp: Use%{TIMESTAMP_ISO8601}instead ofDATA(matches your ISO 8601 timestamp format perfectly).request_id: Use%{UUID}instead ofDATA(matches your UUID format exactly).status: Use%{WORD}instead ofDATA(values likeSUCCESSFULare single words).current_step_time_ms/total_time_ms: Use%{NUMBER:int}to directly parse as integers, avoiding post-processing withmutate.quartz_trigger_timestamp: Create a custom pattern for that specific date format (Wed Jan 11 15:16:13 IST 2023). Add this to a custom patterns file (e.g.,./patterns/custom):MY_TIMESTAMP %{DAY} %{MONTH} %{MONTHDAY} %{TIME} %{WORD} %{YEAR}
3. Use Dissect + Grok for Faster Processing
For comma-separated logs, Dissect is far more efficient than Grok because it works on fixed delimiters instead of regex. Use it to split the log into fields first, then only apply Grok to the fields that need complex pattern matching. This reduces regex processing to just the necessary fields.
Example Dissect + Grok Setup:
# First split the log line with Dissect (super fast!) dissect { mapping => { "message" => "%{created_timestamp},%{request_id},%{tenant},%{username},%{job_code},%{stepname},%{quartz_trigger_timestamp},%{execution_level},%{facility_name},%{channel_code},%{status},%{current_step_time_ms},%{total_time_ms},'%{error_message}',%{tenant_mode},%{channel_src_code}%{?jobSpecificMetaData}" } } # Clean up the optional jobSpecificMetaData field if it exists if [jobSpecificMetaData] { mutate { gsub => ["jobSpecificMetaData", "^,'(.*)'$", "\1"] } } # Convert numeric fields to integers mutate { convert => { "current_step_time_ms" => "integer" "total_time_ms" => "integer" } } # Only use Grok for the non-standard timestamp format grok { patterns_dir => ["./patterns"] match => { "quartz_trigger_timestamp" => "%{MY_TIMESTAMP:quartz_trigger_timestamp}" } }
4. Tune Logstash JVM Settings
High CPU usage can also be exacerbated by insufficient heap memory, leading to frequent garbage collection. Adjust your jvm.options file to allocate enough memory (based on your server's resources):
-Xms4g -Xmx4g
Start with 4GB and adjust up if needed (aim for 50-70% of available RAM).
Why This Works
- Merging patterns eliminates redundant regex checks per log line.
- Precise patterns reduce regex backtracking, which is the #1 cause of Grok CPU bloat.
- Dissect handles the simple delimiter-based splitting, leaving Grok to only do the complex work it's good at.
With these changes, you should see a dramatic drop in CPU usage while maintaining accurate log parsing.
内容的提问来源于stack exchange,提问作者Gajendar

