You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用awk提取XML标签值并替换标签值的日志处理需求

使用AWK处理含XML片段的日志脱敏需求

我明白你需要用AWK处理带XML片段的日志,完成特定字段的脱敏替换,虽然知道AWK不是解析XML的最优方案,但还是得基于它实现需求。下面是完整的解决方案:

完整AWK脚本

#!/usr/bin/awk -f

# 处理name标签:保留第一个字符,其余替换为*
function mask_name(str) {
    if (length(str) == 0) return str
    return substr(str, 1, 1) gensub(/./, "*", "g", substr(str, 2))
}

# 处理phone标签:保留第1、4、7位,其余替换为*
function mask_phone(str) {
    len = length(str)
    res = ""
    for (i=1; i<=len; i++) {
        if (i == 1 || i == 4 || i ==7) {
            res = res substr(str, i, 1)
        } else {
            res = res "*"
        }
    }
    return res
}

# 处理DOB标签:仅2022-02-08的日期脱敏为**-**-**,其余保留原内容
function mask_dob(str, date) {
    if (date == "2022-02-08") {
        return "**-**-**"
    } else {
        return str
    }
}

# 处理bankaccount标签:保留指定位置字符,其余替换为*
function mask_bank(str) {
    len = length(str)
    res = ""
    for (i=1; i<=len; i++) {
        if (i ==1 || (i>=4 && i<=5) || i==8 || i==11) {
            res = res substr(str, i, 1)
        } else {
            res = res "*"
        }
    }
    return res
}

# 主处理逻辑
{
    # 提取每行的日期字段
    date = $1
    line = $0

    # 替换name标签内容
    while (match(line, /<name>([^<]+)<\/name>/, arr)) {
        masked = mask_name(arr[1])
        line = substr(line, 1, RSTART-1) "<name>" masked "</name>" substr(line, RSTART+RLENGTH)
    }

    # 替换phone标签内容
    while (match(line, /<phone>([^<]+)<\/phone>/, arr)) {
        masked = mask_phone(arr[1])
        line = substr(line, 1, RSTART-1) "<phone>" masked "</phone>" substr(line, RSTART+RLENGTH)
    }

    # 替换DOB标签内容
    while (match(line, /<DOB>([^<]+)<\/DOB>/, arr)) {
        masked = mask_dob(arr[1], date)
        line = substr(line, 1, RSTART-1) "<DOB>" masked "</DOB>" substr(line, RSTART+RLENGTH)
    }

    # 替换bankaccount标签内容
    while (match(line, /<bankaccount>([^<]+)<\/bankaccount>/, arr)) {
        masked = mask_bank(arr[1])
        line = substr(line, 1, RSTART-1) "<bankaccount>" masked "</bankaccount>" substr(line, RSTART+RLENGTH)
    }

    # 输出处理后的行
    print line
}

脚本说明

自定义脱敏函数

  • mask_name:保留姓名的第一个字符,剩下的所有字符替换为*,比如sam会变成s**。
  • mask_phone:遍历手机号每一位,只保留第1、4、7位,其他位置替换为*,比如98762123处理后是9**6**2*。
  • mask_dob:根据日志的日期判断,如果是2022-02-08就把出生日期替换成**-**-**,其他日期的DOB保持原样。
  • mask_bank:保留银行账号的第1位、第4-5位、第8位、第11位,其余替换为*,比如4563728495847处理后是4**37**4**8**。

主处理流程

  1. 提取每行的第一个字段作为日期标识。
  2. 依次对name、phone、DOB、bankaccount四个标签进行匹配:用match函数捕获标签内的原始内容,调用对应脱敏函数处理后,替换掉原始的标签内容。
  3. 输出处理后的整行,没有XML片段的行会直接原样输出。

使用方法

把脚本保存为mask_log.awk,然后执行以下命令:

awk -f mask_log.awk input.log > output.log

输入输出示例

输入日志

2022-02-08 [this is the actual log file] BLAH - This is how my xml log file is &lt;sometag id=&quot;00000-00000&quot;&lt;name&gt;sam&lt;/name&gt;&lt;phone&gt;98762123&lt;/phone&gt;&lt;DOB&gt;12-09-77&lt;/DOB&gt;&lt;bankaccount&gt;4563728495847&lt;/bankaccount&gt;&lt;/sometag&gt;
2022-02-09 [this is the actual log file] BLAH - This is how my xml log file is &lt;sometag id=&quot;00000-00000&quot;&lt;name&gt;sam&lt;/name&gt;&lt;phone&gt;123456789&lt;/phone&gt;&lt;DOB&gt;12-09-77&lt;/DOB&gt;&lt;bankaccount&gt;4563728495847&lt;/bankaccount&gt;&lt;/sometag&gt;
2022-02-08 [this is the actual log file] BLAH
2022-02-08 [this is the actual log file] BLAH
2022-02-08 [this is the actual log file] BLAH

输出日志

2022-02-08 [this is the actual log file] BLAH - This is how my xml log file is &lt;sometag id=&quot;00000-00000&quot;&lt;name&gt;s**&lt;/name&gt;&lt;phone&gt;9**6**2*&lt;/phone&gt;&lt;DOB&gt;**-**-**&lt;/DOB&gt;&lt;bankaccount&gt;4**37**4**8**&lt;/bankaccount&gt;&lt;/sometag&gt;
2022-02-09 [this is the actual log file] BLAH - This is how my xml log file is &lt;sometag id=&quot;00000-00000&quot;&lt;name&gt;s**&lt;/name&gt;&lt;phone&gt;1**4**7**&lt;/phone&gt;&lt;DOB&gt;12-09-77&lt;/DOB&gt;&lt;bankaccount&gt;4**37**4**8**&lt;/bankaccount&gt;&lt;/sometag&gt;
2022-02-08 [this is the actual log file] BLAH
2022-02-08 [this is the actual log file] BLAH
2022-02-08 [this is the actual log file] BLAH

注意事项

  • 脚本假设XML标签都是单行的,且没有嵌套、跨行的情况,如果日志结构更复杂,需要调整正则匹配逻辑。
  • 正如你所说,AWK确实不是处理XML的最佳工具,这个脚本仅适配当前给定的日志格式,后续如果日志结构变化,可能需要修改代码。

内容的提问来源于stack exchange,提问作者felicity

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 13:07:37