如何用awk提取表格中非固定位置的Plan列内容?
提取表格Plan列内容的awk解决方案
问题描述
有如下格式的表格数据:
+-------+--------------------+-----------+------------+-----------+-------------+ | ID | Name | Status | Networks | Image | Plan | +-------+--------------------+-----------+------------+-----------+-------------+ | 1wsd | HostName A | PAUSED | IP=1.1.1.1 | Ubuntu20 | PlanA BGP40 | | 4fgh | An other hostname | ACTIVE | IP=2.2.2.2 | Ubuntu20 | PlanB BGP30 | | zxd1 | final.destination | REBOOTING | IP=3.3.3.3 | Debian11 | PlanA BGP10 | | 60hn | no problem | ACTIVE | IP=4.4.4.4 | Centos7 | Plan BGP90 | +-------+--------------------+-----------+------------+-----------+-------------+
需要提取其中的Plan列内容(该列内容可能包含空格,非固定列数),同时排除前3行表头和最后1行表尾,预期输出:
PlanA BGP40 PlanB BGP30 PlanA BGP10 Plan BGP90
解决方案
以下是几种实用的awk实现方式:
方法1:基于表头动态定位Plan列
先获取总行数,再从表头行找到Plan列的起始位置,后续行截取对应区域的内容并清理格式:
END { total_lines = NR } NR==2 { # 定位Plan列在表头中的起始位置 plan_col_start = index($0, "Plan") } NR>3 && NR<total_lines { # 截取从Plan列起始位置到行尾的内容 plan_content = substr($0, plan_col_start) # 去掉末尾的|和多余空格 sub(/ *\|$/, "", plan_content) print plan_content }
执行命令:
awk -f extract_plan.awk your_table_file.txt
方法2:按分隔符提取固定列(已知Plan为最后一列时)
如果确定Plan列是表格的最后一列,可以用|作为分隔符,提取最后一个有效字段:
BEGIN { FS="|" } END { total_lines = NR } NR>3 && NR<total_lines { # 提取第6个字段,清理前后空格 gsub(/^ +| +$/, "", $6) print $6 }
执行命令:
awk 'BEGIN { FS="|" } END { total_lines=NR } NR>3 && NR<total_lines { gsub(/^ +| +$/, "", $6); print $6 }' your_table_file.txt
方法3:正则匹配行尾的Plan内容
通过正则匹配每行末尾的Plan列内容,无需依赖列数:
END { total_lines = NR } NR>3 && NR<total_lines { # 匹配最后一个|后的内容,捕获组提取有效信息 match($0, /\| *([^|]+) *\|$/, match_result) print match_result[1] }
执行命令:
awk 'END { total_lines=NR } NR>3 && NR<total_lines { match($0, /\| *([^|]+) *\|$/, match_result); print match_result[1] }' your_table_file.txt
内容的提问来源于stack exchange,提问作者Saeed
相关产品推荐
相关产品推荐

