You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用awk/sed实现XML中continent与title信息关联输出

实现XML分类与对应标题行的关联输出

需求

将按分类排序的XML文件中,每个分类(continent)与该分类下含title元素的行信息打印在同一行。

输入示例

通过grep -E "Kundschaft|continent" filename得到的典型输入:

<landmarks continent="bagne" type="user">
<landmark type="3" x="595.21" y="-10981.47" title="~Ba-Nung Liangi (Kundschafter)"/>
<landmark type="3" x="943.54" y="-10365.21" title="Ba-Nung Liangi (Kundschafterin, Sapsammler)"/>
<landmarks continent="corrupted_moor" type="user"/>
<landmarks continent="fyros" type="user">
<landmark type="3" x="17106.73" y="-25706.46" title="Xymus Tindix (Kundschafter)"/>
<landmark type="3" x="17586.79" y="-25679.67" title="Apolus Abygrian (Kundschafter)"/>
<landmark type="3" x="17018.25" y="-25306.73" title="Ba'Reiliam Breggi (Kundschafter)"/>

期望输出

bagne: ~Ba-Nung Liangi (Kundschafter)
bagne: Ba-Nung Liangi (Kundschafterin, Sapsammler)
fyros: Xymus Tindix (Kundschafter)
fyros: Apolus Abygrian (Kundschafter)
fyros: Ba'Reiliam Breggi (Kundschafter)

解决方案

方案1:结合grep与awk

先用grep过滤出目标行,再用awk记住continent值并输出对应内容:

grep -E "Kundschaft|continent" filename | awk '
/continent/ {
    # 提取continent字段的值,剥离引号和后续字符
    split($0, parts, /"| /)
    curr_continent = parts[4]
}
/Kundschaft/ {
    # 提取title字段的值,剥离引号和结尾的/>
    split($0, parts, /"|\/>/)
    curr_title = parts[8]
    print curr_continent ": " curr_title
}
'

方案2:直接用awk处理原文件(更高效)

无需提前过滤,awk自身即可完成行匹配与字段提取:

awk '
# 匹配continent行,提取分类值
/landmarks continent=/ {
    match($0, /continent="([^"]+)"/, matches)
    curr_continent = matches[1]
}
# 匹配含Kundschaft的title行,提取标题并输出
/Kundschaft/ {
    match($0, /title="([^"]+)"/, matches)
    print curr_continent ": " matches[1]
}
' filename

说明

  • 两种方案都会先捕获当前的continent值并保存,后续遇到含Kundschaft的行时,直接调用该值拼接输出
  • 方案2省去了grep的额外步骤,处理大文件时效率更高

内容的提问来源于stack exchange,提问作者planetmaker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 09:05:33