You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用awk匹配词表文件内容并提取指定字段符合条件的记录?

用Awk基于词表文件筛选特定字段匹配的记录

我有一个含多字段的日志文件,还有一个存储关键词的词表文件,需要用awk命令从日志文件中提取**特定字段(这里是第3个以: 分隔的字段)**以词表中任意单词开头的记录。

示例日志文件(log.txt)

Feb 15 12:05:10 lcif adm.slm: root [23416]: cd /tmp
Feb 15 12:05:24 lcif adm.slm: root [23416]: cat tst.sh
Feb 15 12:05:44 lcif adm.slm: root [23416]: date
Feb 15 12:05:52 lcif adm.pse: root [23419]: rm -f file
Feb 15 12:05:58 lcif adm.pse: root [23419]: who
Feb 15 12:06:02 lcif adm.pse: root [23419]: uptime
Feb 15 12:06:56 lcif adm.pse: root [23419]: reboot
Feb 15 12:06:58 lcif adm.pse: root [23419]: ls -lrt

示例词表文件(words.txt)

rm
reboot
shutdown

期望输出

Feb 15 12:05:52 lcif adm.pse: root [23419]: rm -f file
Feb 15 12:06:56 lcif adm.pse: root [23419]: reboot

我之前写过硬编码多个正则的awk命令:

awk -F ": " '{if ($3 ~ "^rm" || $3 ~ "^reboot" || $3 ~ "^shutdown") print}' log.txt

但词表越来越长,硬编码太麻烦,希望直接用词表文件实现匹配。


解决方案

方案1:用数组存储词表匹配

通过awk的BEGIN块提前读取词表到数组,再遍历数组检查目标字段:

awk -F ": " '
BEGIN {
    # 读取词表文件,将每个关键词存入数组
    while (getline < "words.txt") {
        keywords[$0] = 1
    }
}
{
    # 遍历词表,检查第3字段是否以关键词开头
    for (word in keywords) {
        if ($3 ~ "^" word) {
            print
            break  # 匹配到即停止,避免重复输出
        }
    }
}
' log.txt

方案2:拼接正则表达式(更高效)

把词表拼接成一个复合正则,直接匹配目标字段:

awk -F ": " '
BEGIN {
    regex = ""
    # 拼接词表为正则:^rm|^reboot|^shutdown
    while (getline < "words.txt") {
        regex = regex (regex ? "|" : "") "^" $0
    }
}
# 匹配正则就打印该行
$3 ~ regex
' log.txt

内容的提问来源于stack exchange,提问作者Serge29

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 05:27:25