如何在Lua中使用正则表达式对多份大日志文件做字符串替换
Alright, let's tackle this string replacement task for your large log files (1000+ lines each) using Lua's pattern matching. First, let's break down the structure of each log entry to understand what we're working with:
Each entry follows this clunky pattern:
{@BLOCK|{variable_name}|00 {@A-JUM|0|{numeric_value}{@LIM2|{upper_limit}|{lower_limit}}}} }
Our goal is to extract those key pieces of data (variable name, value, limits) and replace the messy original format with something more readable, or adjust it to whatever your specific needs are.
Step 1: Define the Matching Pattern
Lua uses its own pattern syntax (similar but not identical to standard regex). We'll use capture groups to pull out the parts we care about:
([^|]+): Captures everything until the next|(grabs the variable name like1%r1331)([^%{]+): Captures everything until the next{(grabs the numeric value like-9.352000E+06)([^|]+)and([^%}]+): Capture the upper and lower limits from theLIM2section
Step 2: Example Replacement Code
Here's a working example that converts each BLOCK entry into a clean, human-readable line:
-- Sample log line (replace this with lines read from your file) local sample_log = "{@BLOCK|1%r1331|00 {@A-JUM|0|-9.352000E+06{@LIM2|+9.999999E+99|+1.000000E+04}} } {@BLOCK|1%x1001_swp|00 {@A-JUM|0|+3.362121E+00{@LIM2|+2.000000E+01|+0.000000E+00}} }" -- Pattern to match each BLOCK entry and capture key data local match_pattern = "{@BLOCK|([^|]+)|00 {@A-JUM|0|([^%{]+){@LIM2|([^|]+)|([^%}]+)}} }" -- Customize this replacement string to fit your needs! local replacement_format = "BLOCK: %1 | Value: %2 | Limits: %4 to %3" -- Run the global substitution local cleaned_log = string.gsub(sample_log, match_pattern, replacement_format) print(cleaned_log)
Sample Output:
BLOCK: 1%r1331 | Value: -9.352000E+06 | Limits: +1.000000E+04 to +9.999999E+99 BLOCK: 1%x1001_swp | Value: +3.362121E+00 | Limits: +0.000000E+00 to +2.000000E+01
Step 3: Handle Large Log Files Efficiently
Since your files have 1000+ lines, it's better to process them line-by-line instead of loading the entire file into memory. Here's a script that reads from an input log and writes cleaned content to an output file:
-- Replace these paths with your actual file paths local input_path = "raw_log.log" local output_path = "cleaned_log.log" local input_file = io.open(input_path, "r") local output_file = io.open(output_path, "w") if not input_file or not output_file then print("Error: Could not open input or output file!") return end -- Reuse our pattern and replacement format local match_pattern = "{@BLOCK|([^|]+)|00 {@A-JUM|0|([^%{]+){@LIM2|([^|]+)|([^%}]+)}} }" local replacement_format = "BLOCK: %1 | Value: %2 | Limits: %4 to %3" -- Process each line one at a time for line in input_file:lines() do local cleaned_line = string.gsub(line, match_pattern, replacement_format) output_file:write(cleaned_line .. "\n") end -- Clean up input_file:close() output_file:close() print("Success! Cleaned log saved to " .. output_path)
Customization Tips
- Adjust the output format: Modify
replacement_formatto fit your needs. For example, if you only need the variable name and value, use"[%1] = %2". - Remove unwanted sections: If you want to strip out the
LIM2limits entirely, tweak the pattern to ignore that part and simplify the replacement. - Handle edge cases: If some log entries have slight variations (like different numbers in the
A-JUMsection), adjust the pattern to match those differences (e.g., replace0with%d+if the number can change).
内容的提问来源于stack exchange,提问作者costica_p

