You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用XPath 1.0(xmllint)拼接XML中每个item的title与category

使用xmllint(XPath 1.0)提取RSS中item的标题与分类并拼接

需求

从以下RSS格式的XML文件中,通过仅支持XPath 1.0的xmllint工具,将每个<item>节点的<title>和<category>文本拼接为标题,分类格式,每行对应一个item:

<rss version="2.0">
    <channel>
        <language>en</language>
        <pubDate>Tue, 19 Mar 2024 06:06:35 GMT</pubDate>
        <item>
            <title>Title1</title>
            <category>91021</category>
        </item>
        <item>
            <title>Title2</title>
            <category>91022</category>
        </item>
        <item>
            <title>Title3</title>
            <category>91023</category>
        </item>
    </channel>
</rss>

期望输出:

Title1,91021
Title2,91022
Title3,91023

问题分析

你尝试的//item/(concat(title/text(), ".", category/text())表达式属于XPath 2.0语法,该语法支持对节点集中的每个节点单独执行表达式,但XPath 1.0不支持这种写法,因此无法生效。

可行解决方案

方案1:通用Shell循环(适配任意数量的item)

通过先统计item节点的数量,再逐个提取并拼接每个item的内容,最后添加换行:

# 统计item节点总数
item_count=$(xmllint --xpath 'count(//item)' input.xml)

# 遍历每个item节点,拼接内容并换行输出
for i in $(seq 1 $item_count); do
  xmllint --xpath '//item['$i']/concat(title/text(), ",", category/text())' input.xml
  echo
done

方案2:直接指定item位置(适合固定数量的item)

如果已知item的数量,可以直接通过索引拼接所有内容,手动添加换行符(XPath中用&#10;表示换行):

xmllint --xpath 'concat(
  //item[1]/title/text(), ",", //item[1]/category/text(), "&#10;",
  //item[2]/title/text(), ",", //item[2]/category/text(), "&#10;",
  //item[3]/title/text(), ",", //item[3]/category/text()
)' input.xml

方案3:利用xmllint Shell模式

通过xmllint的交互式Shell模式批量执行表达式,再过滤多余输出:

echo -e 'cat //item/concat(title/text(), ",", category/text())\nquit' | xmllint --shell input.xml | grep -v "^/ >" | grep -v "^----"

说明

以上方案均基于XPath 1.0语法,完全适配xmllint工具的特性,能够输出符合要求的每行拼接结果。

内容的提问来源于stack exchange,提问作者Jose L Martinez-Avial

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 03:13:12