如何使用sed/awk为ligand_types行补全缺失的SA、HD模式
ligand_types行缺失字段自动补全方案
实现逻辑
针对所有以ligand_types开头的行,拆分注释符#前后的正文、注释两部分,检查正文内是否存在独立的SA、HD字段,将缺失的字段按顺序追加到正文末尾,保留原有注释内容,同时对齐格式。
awk 方案(推荐,逻辑稳定无边界问题)
直接运行以下命令即可输出处理后的内容,确认结果无误后再加原地修改参数:
awk ' /^ligand_types/ { # 按#拆分正文和注释 split($0, sect, /#/) main_text = sect[1] note = "#" sect[2] append = "" # 匹配独立单词,避免把包含SA/HD的其他字段误判 if (main_text !~ /(^|[[:space:]])SA([[:space:]]|$)/) append = append " SA" if (main_text !~ /(^|[[:space:]])HD([[:space:]]|$)/) append = append " HD" # 拼接内容,自动对齐注释位置 $0 = sprintf("%-35s%s", main_text append, note) } # 非目标行直接原样输出 1' 你的文件路径
如果需要直接修改原文件,使用gawk的原地修改参数即可:
gawk -i inplace ' /^ligand_types/ { split($0, sect, /#/) main_text = sect[1] note = "#" sect[2] append = "" if (main_text !~ /(^|[[:space:]])SA([[:space:]]|$)/) append = append " SA" if (main_text !~ /(^|[[:space:]])HD([[:space:]]|$)/) append = append " HD" $0 = sprintf("%-35s%s", main_text append, note) } 1' 你的文件路径
sed 方案
纯sed实现不需要额外依赖,直接用-i参数即可原地修改:
sed -i ' # 仅处理ligand_types开头的行 /^ligand_types/{ # 兼容行尾没有注释的极端情况 /#/!s/$/ # ligand atom types/ # 正文无SA则在注释前插入SA /\(.*[[:space:]]\)SA[[:space:]].*#/!s/\(.*\)[[:space:]]*\(#.*\)/\1 SA \2/ # 正文无HD则在注释前插入HD /\(.*[[:space:]]\)HD[[:space:]].*#/!s/\(.*\)[[:space:]]*\(#.*\)/\1 HD \2/ # 清理多余空格,对齐注释格式 s/[[:space:]]*#/ #/ } ' 你的文件路径
效果验证
使用给出的样例输入测试,两种方案的输出均和预期结果完全一致:
ligand_types A C Cl NA OA N HD SA # ligand atom types ligand_types A C NA OA N SA HD # ligand atom types ligand_types A C Cl NA OA N HD SA # ligand atom types
内容的提问来源于stack exchange,提问作者James Starlight
相关产品推荐
相关产品推荐

