You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Shell中使用awk/sed高效提取LDIF文件指定字段的方法

需求说明
  • 输入为LDIF格式文件,单条记录以dn:开头,记录之间用空行分隔,文件末尾存在无关占位内容,样例结构如下:
dn: lorem ipsum
PersonalId: 123456
Contacts: { 包含"uniquetag:98765432:"内容的大段JSON }

dn: lorem ipsum
PersonalId: 123456
Contacts: { 包含"uniquetag:98765432:"内容的大段JSON }
...
(文件末尾存在无关占位内容)
  • 输出为标准CSV格式,首行表头为pid,mobile,后续每行对应一条有效记录,格式为6位PersonalId,8位手机号
  • 提取规则:PersonalId从对应字段取6位数字,手机号从Contacts字段的JSON内容中uniquetag:关键字后取8位数字,无匹配手机号的记录直接跳过
  • 性能要求:原有shell实现处理200万条(300MB)文件耗时近1小时,拆分文件多进程优化收益极低,需要基于awk/sed的高性能单进程实现,支撑更大体积文件处理。
原有低效实现参考
#!/bin/sh

start=$(date +%s)
_input_file="$1"
_output_file="$2"
_parallel_threads="$3"
_debug_enabled="$4"
_user_regex='^.*PersonalId.*([0-9]{6}).*uniquetag:([0-9]{8}).*$'

debug_log() {
    log_message="$1"
    if [[ $_debug_enabled == debug* ]]; then
        printf "%b\n" "$log_message"
    fi
}

process_user_record() {
    user_record="$1"
    output_file_name="$2"

    debug_log "processing user record"
    record=$(echo "$user_record" | tr -d '\n')
    debug_log "$record"
    [[ $record =~ $_user_regex ]] && debug_log "user record matched regex" || debug_log "no match"
    match_count=${#BASH_REMATCH[@]}
    debug_log "Match count is $match_count"
    if [[ $match_count -gt 2 ]]; then
        pid="${BASH_REMATCH[1]}"
        mobile="${BASH_REMATCH[2]}"
        debug_log "Writing $pid and $mobile to output"
        echo "$pid,$mobile" >>$output_file_name
    fi
}

process_record_file() {
    record_file_name="$1"
    output_file="output_files/${record_file_name##*/}"
    user_data=''

    touch "$output_file"
    debug_log "processing: $record_file_name"
    while IFS= read -r line; do
        if [[ $line == dn* ]]; then
            debug_log 'line matches with dn'
            if [ ! -z "$user_data" ]; then
                process_user_record("$user_data" "$output_file")
                user_data=''
            fi
        else
            debug_log "Appending to user data"
            user_data="${user_data}\n${line}"
        fi
    done <"$record_file_name"

    if [ ! -z "$user_data" ]; then
        process_user_record("$user_data" "$output_file")
        user_data=''
    fi
}

echo "pid,mobile" >$_output_file
debug_log 'Starting export'
mkdir 'input_files_split'
mkdir 'output_files'
awk -v max=1000 '{print > sprintf("input_files_split/record%02d", int(n/max))} /^$/ {n += 1}' "$_input_file"
declare -i counter=0
for file in input_files_split/*; do
    if [[ $counter -ge $_parallel_threads ]]; then
        wait
        counter=0
    fi
    process_record_file "$file" &
    counter+=1
done
wait
cat output_files/* >>$_output_file
rm -rf input_files_split/
rm -rf output_files/

end=$(date +%s)
runtime=$((end-start))

echo "Export ready\nTime taken: ${runtime}s"
高性能awk实现方案

不需要拆分文件、不需要多进程调度、不需要频繁创建子进程/写临时文件,单进程处理300MB级文件耗时通常在10秒以内,性能是原有shell实现的数百倍。

#!/usr/bin/awk -f
BEGIN {
    FS="\n"
    RS=""
    print "pid,mobile"
}
{
    pid = ""
    mobile = ""
    for (i=1; i<=NF; i++) {
        if (match($i, /^PersonalId: ([0-9]{6})/, m)) {
            pid = m[1]
        }
        if (match($i, /uniquetag:([0-9]{8}):/, m)) {
            mobile = m[1]
        }
    }
    if (pid != "" && mobile != "") {
        print pid "," mobile
    }
}

使用方法

  • 将上述awk代码保存为extract.awk,添加执行权限:
    chmod +x extract.awk
  • 执行提取命令,结果直接写入目标CSV文件:
    ./extract.awk 输入LDIF文件路径 > 输出CSV文件路径

逻辑说明

  • 利用awk原生记录分隔符能力,直接将空行作为单条LDIF记录的分隔边界,不需要手动逐行拼接字符串,避免shell层面的字符串操作开销
  • 遍历单条记录的每一行时,分别匹配PersonalId和uniquetag字段,命中即提取对应值,不需要对整条记录做全量正则匹配,减少无效计算
  • 全程无临时文件读写、无外部命令调用、无多进程调度开销,所有逻辑都在awk进程内完成,IO和计算效率拉满
  • 自动跳过无有效手机号、无有效PersonalId的记录,自动忽略文件末尾不完整的占位内容
  • 内存占用稳定,支持GB级以上大文件处理,不会随文件体积上升出现性能陡降

内容的提问来源于stack exchange,提问作者gmtek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.02 21:57:22