Awk中POSIX ERE负向预查的替代方案及链接筛选实现
替代嵌套if的更优方案
方案一:单Awk进程批量处理
原方案每个链接调用两次Awk,频繁创建子进程效率偏低,可将整个链接数组一次性传给Awk处理,减少进程开销:
links=("https://www.yahoo.com/" "https://www.yahoo.com/codeword/lalala1" "https://www.yahoo.com/nocodeword" "https://www.bing.com/" "https://www.yahoo.com/codeword" "https://www.yahoo.com/codeword/lalala2" "https://www.google.com/" "https://www.yahoo.com/foo/codeword" "https://www.yahoo.com/codewordbar") printf "%s\n" "${links[@]}" | awk -F/ ' $3 == "www.yahoo.com" && ($4 == "" || $4 != "codeword") { print } '
逻辑说明:
- 用
printf将所有链接逐行传递给Awk,避免循环内多次调用外部工具 - 按
/分割字符串后,$3对应域名部分,先校验是否为目标域名 $4对应路径的首个分段:若为空(即根路径https://www.yahoo.com/),或不等于codeword,则输出该链接
方案二:Shell原生正则匹配(无需Awk)
利用POSIX扩展正则,直接在[[ ]]中完成匹配判断,完全依赖shell原生能力:
links=("https://www.yahoo.com/" "https://www.yahoo.com/codeword/lalala1" "https://www.yahoo.com/nocodeword" "https://www.bing.com/" "https://www.yahoo.com/codeword" "https://www.yahoo.com/codeword/lalala2" "https://www.google.com/" "https://www.yahoo.com/foo/codeword" "https://www.yahoo.com/codewordbar") target_domain="www.yahoo.com" exclude_path="codeword" for link in "${links[@]}"; do if [[ $link =~ ^https://$target_domain(/|$) || ($link =~ ^https://$target_domain/[^/]+ && ! $link =~ ^https://$target_domain/$exclude_path(/|$)) ]]; then echo "$link" fi done
逻辑说明:
匹配两种符合要求的链接类型:
- 目标域名的根路径(含末尾带斜杠和不带斜杠的情况)
- 目标域名下有路径分段,但首个分段不是
codeword(直接排除以codeword开头的所有路径)
两种方案均可输出预期结果,相比原嵌套if方案,要么降低了子进程开销,要么逻辑更直观紧凑。
内容的提问来源于stack exchange,提问作者bullet03
相关产品推荐
相关产品推荐

