如何用awk/sed实现XML中continent与title信息关联输出
实现XML分类与对应标题行的关联输出
需求
将按分类排序的XML文件中,每个分类(continent)与该分类下含title元素的行信息打印在同一行。
输入示例
通过grep -E "Kundschaft|continent" filename得到的典型输入:
<landmarks continent="bagne" type="user"> <landmark type="3" x="595.21" y="-10981.47" title="~Ba-Nung Liangi (Kundschafter)"/> <landmark type="3" x="943.54" y="-10365.21" title="Ba-Nung Liangi (Kundschafterin, Sapsammler)"/> <landmarks continent="corrupted_moor" type="user"/> <landmarks continent="fyros" type="user"> <landmark type="3" x="17106.73" y="-25706.46" title="Xymus Tindix (Kundschafter)"/> <landmark type="3" x="17586.79" y="-25679.67" title="Apolus Abygrian (Kundschafter)"/> <landmark type="3" x="17018.25" y="-25306.73" title="Ba'Reiliam Breggi (Kundschafter)"/>
期望输出
bagne: ~Ba-Nung Liangi (Kundschafter) bagne: Ba-Nung Liangi (Kundschafterin, Sapsammler) fyros: Xymus Tindix (Kundschafter) fyros: Apolus Abygrian (Kundschafter) fyros: Ba'Reiliam Breggi (Kundschafter)
解决方案
方案1:结合grep与awk
先用grep过滤出目标行,再用awk记住continent值并输出对应内容:
grep -E "Kundschaft|continent" filename | awk ' /continent/ { # 提取continent字段的值,剥离引号和后续字符 split($0, parts, /"| /) curr_continent = parts[4] } /Kundschaft/ { # 提取title字段的值,剥离引号和结尾的/> split($0, parts, /"|\/>/) curr_title = parts[8] print curr_continent ": " curr_title } '
方案2:直接用awk处理原文件(更高效)
无需提前过滤,awk自身即可完成行匹配与字段提取:
awk ' # 匹配continent行,提取分类值 /landmarks continent=/ { match($0, /continent="([^"]+)"/, matches) curr_continent = matches[1] } # 匹配含Kundschaft的title行,提取标题并输出 /Kundschaft/ { match($0, /title="([^"]+)"/, matches) print curr_continent ": " matches[1] } ' filename
说明
- 两种方案都会先捕获当前的
continent值并保存,后续遇到含Kundschaft的行时,直接调用该值拼接输出 - 方案2省去了grep的额外步骤,处理大文件时效率更高
内容的提问来源于stack exchange,提问作者planetmaker
相关产品推荐
相关产品推荐

