Shell脚本中带分隔符的grep用法及日志grep操作咨询
Hey there! Let's break down how to parse your Hive logs with grep (plus some complementary tools) in a shell script, focusing on structured field separation using delimiters.
First, let's recap your log structure—each line follows this pattern (simplified):
2018-03-20T15:26:34,397 INFO [2da4e66f-6092-46a7-9542-60afc0611205 HiveServer2-Handler-Pool: Thread-32([])]: ql.Driver (Driver.java:compile(429)) - Compiling command(queryId=hive_20180320152634_a6ef02a1-e018-4085-8ceb-8a8d2733b427): select * from reportingperiod limit 5
Here are practical, script-friendly methods to handle this:
Grep alone doesn't handle field splitting, but pairing it with awk lets you split matching lines into clean, delimited fields. For example, splitting into timestamp, log level, and SQL statement:
# Filter lines with "select", then split using "- " and ": " as delimiters grep -i "select" logfile | awk -F ' - |: ' '{print "Timestamp: "$1"\nLog Level: "$2"\nSQL Query: "$NF}'
grep -i "select": Matches all lines with "select" (case-insensitive; remove-iif you need exact case)awk -F ' - |: ': Uses two delimiters (-and:) to split the line into segments$1grabs the timestamp,$2the log level (INFO), and$NFthe last field (your full SQL query)
If you just need to extract the SQL part without extra log noise, use grep's -o flag to output only the matched portion:
# Extract everything from "select" to the end of the line grep -o 'select.*$' logfile
-o: Tells grep to print only the matched text, not the entire lineselect.*$: Matches from "select" all the way to the end of the line
For more precision (to avoid accidental matches elsewhere), narrow it down to queries after the Compiling command marker:
grep -o 'Compiling command.*: select.*$' logfile | sed 's/Compiling command.*: //'
The sed command strips off the prefix, leaving just your SQL query.
If you need the results in a format easy to import into tools (like spreadsheets), use a custom delimiter (e.g., comma, tab, pipe):
# Use tab as delimiter grep -i "select" logfile | awk -F ' - |: ' '{printf "%s\t%s\t%s\n", $1, $2, $NF}' # Or use pipe as delimiter grep -i "select" logfile | awk -F ' - |: ' '{printf "%s|%s|%s\n", $1, $2, $NF}'
To make this easy to reuse in scripts, wrap it into a function with configurable delimiters:
parse_hive_select_queries() { local log_file="$1" local delimiter="${2:-|}" # Default to pipe if no delimiter is passed grep -i "select" "$log_file" | awk -v sep="$delimiter" -F ' - |: ' '{print $1 sep $2 sep $NF}' } # Example usage: parse_hive_select_queries "logfile" ","
内容的提问来源于stack exchange,提问作者Teju Priya

