You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Shell脚本中带分隔符的grep用法及日志grep操作咨询

Hey there! Let's break down how to parse your Hive logs with grep (plus some complementary tools) in a shell script, focusing on structured field separation using delimiters.

First, let's recap your log structure—each line follows this pattern (simplified):

2018-03-20T15:26:34,397 INFO [2da4e66f-6092-46a7-9542-60afc0611205 HiveServer2-Handler-Pool: Thread-32([])]: ql.Driver (Driver.java:compile(429)) - Compiling command(queryId=hive_20180320152634_a6ef02a1-e018-4085-8ceb-8a8d2733b427): select * from reportingperiod limit 5

Here are practical, script-friendly methods to handle this:

1. Extract & Split Log Fields with Delimiters

Grep alone doesn't handle field splitting, but pairing it with awk lets you split matching lines into clean, delimited fields. For example, splitting into timestamp, log level, and SQL statement:

# Filter lines with "select", then split using "- " and ": " as delimiters
grep -i "select" logfile | awk -F ' - |: ' '{print "Timestamp: "$1"\nLog Level: "$2"\nSQL Query: "$NF}'
  • grep -i "select": Matches all lines with "select" (case-insensitive; remove -i if you need exact case)
  • awk -F ' - |: ': Uses two delimiters (- and : ) to split the line into segments
  • $1 grabs the timestamp, $2 the log level (INFO), and $NF the last field (your full SQL query)
2. Pull Only the SELECT Query

If you just need to extract the SQL part without extra log noise, use grep's -o flag to output only the matched portion:

# Extract everything from "select" to the end of the line
grep -o 'select.*$' logfile
  • -o: Tells grep to print only the matched text, not the entire line
  • select.*$: Matches from "select" all the way to the end of the line

For more precision (to avoid accidental matches elsewhere), narrow it down to queries after the Compiling command marker:

grep -o 'Compiling command.*: select.*$' logfile | sed 's/Compiling command.*: //'

The sed command strips off the prefix, leaving just your SQL query.

3. Output Structured, Delimited Content

If you need the results in a format easy to import into tools (like spreadsheets), use a custom delimiter (e.g., comma, tab, pipe):

# Use tab as delimiter
grep -i "select" logfile | awk -F ' - |: ' '{printf "%s\t%s\t%s\n", $1, $2, $NF}'

# Or use pipe as delimiter
grep -i "select" logfile | awk -F ' - |: ' '{printf "%s|%s|%s\n", $1, $2, $NF}'
4. Wrap It in a Reusable Shell Function

To make this easy to reuse in scripts, wrap it into a function with configurable delimiters:

parse_hive_select_queries() {
    local log_file="$1"
    local delimiter="${2:-|}" # Default to pipe if no delimiter is passed
    grep -i "select" "$log_file" | awk -v sep="$delimiter" -F ' - |: ' '{print $1 sep $2 sep $NF}'
}

# Example usage:
parse_hive_select_queries "logfile" ","

内容的提问来源于stack exchange,提问作者Teju Priya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:50:24