Shell中使用awk/sed高效提取LDIF文件指定字段的方法
需求说明
- 输入为LDIF格式文件,单条记录以
dn:开头,记录之间用空行分隔,文件末尾存在无关占位内容,样例结构如下:
dn: lorem ipsum PersonalId: 123456 Contacts: { 包含"uniquetag:98765432:"内容的大段JSON } dn: lorem ipsum PersonalId: 123456 Contacts: { 包含"uniquetag:98765432:"内容的大段JSON } ... (文件末尾存在无关占位内容)
- 输出为标准CSV格式,首行表头为
pid,mobile,后续每行对应一条有效记录,格式为6位PersonalId,8位手机号 - 提取规则:PersonalId从对应字段取6位数字,手机号从Contacts字段的JSON内容中
uniquetag:关键字后取8位数字,无匹配手机号的记录直接跳过 - 性能要求:原有shell实现处理200万条(300MB)文件耗时近1小时,拆分文件多进程优化收益极低,需要基于awk/sed的高性能单进程实现,支撑更大体积文件处理。
原有低效实现参考
#!/bin/sh start=$(date +%s) _input_file="$1" _output_file="$2" _parallel_threads="$3" _debug_enabled="$4" _user_regex='^.*PersonalId.*([0-9]{6}).*uniquetag:([0-9]{8}).*$' debug_log() { log_message="$1" if [[ $_debug_enabled == debug* ]]; then printf "%b\n" "$log_message" fi } process_user_record() { user_record="$1" output_file_name="$2" debug_log "processing user record" record=$(echo "$user_record" | tr -d '\n') debug_log "$record" [[ $record =~ $_user_regex ]] && debug_log "user record matched regex" || debug_log "no match" match_count=${#BASH_REMATCH[@]} debug_log "Match count is $match_count" if [[ $match_count -gt 2 ]]; then pid="${BASH_REMATCH[1]}" mobile="${BASH_REMATCH[2]}" debug_log "Writing $pid and $mobile to output" echo "$pid,$mobile" >>$output_file_name fi } process_record_file() { record_file_name="$1" output_file="output_files/${record_file_name##*/}" user_data='' touch "$output_file" debug_log "processing: $record_file_name" while IFS= read -r line; do if [[ $line == dn* ]]; then debug_log 'line matches with dn' if [ ! -z "$user_data" ]; then process_user_record("$user_data" "$output_file") user_data='' fi else debug_log "Appending to user data" user_data="${user_data}\n${line}" fi done <"$record_file_name" if [ ! -z "$user_data" ]; then process_user_record("$user_data" "$output_file") user_data='' fi } echo "pid,mobile" >$_output_file debug_log 'Starting export' mkdir 'input_files_split' mkdir 'output_files' awk -v max=1000 '{print > sprintf("input_files_split/record%02d", int(n/max))} /^$/ {n += 1}' "$_input_file" declare -i counter=0 for file in input_files_split/*; do if [[ $counter -ge $_parallel_threads ]]; then wait counter=0 fi process_record_file "$file" & counter+=1 done wait cat output_files/* >>$_output_file rm -rf input_files_split/ rm -rf output_files/ end=$(date +%s) runtime=$((end-start)) echo "Export ready\nTime taken: ${runtime}s"
高性能awk实现方案
不需要拆分文件、不需要多进程调度、不需要频繁创建子进程/写临时文件,单进程处理300MB级文件耗时通常在10秒以内,性能是原有shell实现的数百倍。
#!/usr/bin/awk -f BEGIN { FS="\n" RS="" print "pid,mobile" } { pid = "" mobile = "" for (i=1; i<=NF; i++) { if (match($i, /^PersonalId: ([0-9]{6})/, m)) { pid = m[1] } if (match($i, /uniquetag:([0-9]{8}):/, m)) { mobile = m[1] } } if (pid != "" && mobile != "") { print pid "," mobile } }
使用方法
- 将上述awk代码保存为
extract.awk,添加执行权限:chmod +x extract.awk - 执行提取命令,结果直接写入目标CSV文件:
./extract.awk 输入LDIF文件路径 > 输出CSV文件路径
逻辑说明
- 利用awk原生记录分隔符能力,直接将空行作为单条LDIF记录的分隔边界,不需要手动逐行拼接字符串,避免shell层面的字符串操作开销
- 遍历单条记录的每一行时,分别匹配PersonalId和uniquetag字段,命中即提取对应值,不需要对整条记录做全量正则匹配,减少无效计算
- 全程无临时文件读写、无外部命令调用、无多进程调度开销,所有逻辑都在awk进程内完成,IO和计算效率拉满
- 自动跳过无有效手机号、无有效PersonalId的记录,自动忽略文件末尾不完整的占位内容
- 内存占用稳定,支持GB级以上大文件处理,不会随文件体积上升出现性能陡降
内容的提问来源于stack exchange,提问作者gmtek
相关产品推荐
相关产品推荐

