在Google Sheets ImportXML中使用XPath谓语过滤标签的正确语法
Google Sheets ImportXML 过滤特定标签问题
问题场景
通过Google Sheets的ImportXML获取到如下XML结果:
<lfm status="ok"> <toptags artist="Madonna"> <tag> <count>100</count> <name>pop</name> <url>https://www.last.fm/tag/pop</url> </tag> <tag> <count>50</count> <name>dance</name> <url>https://www.last.fm/tag/dance</url> </tag> <tag> <count>50</count> <name>female vocalists</name> <url>https://www.last.fm/tag/female+vocalists</url> </tag> <tag> <count>35</count> <name>80s</name> <url>https://www.last.fm/tag/80s</url> </tag> <tag> <count>28</count> <name>electronic</name> <url>https://www.last.fm/tag/electronic</url> </tag> <tag> <count>9</count> <name>seen live</name> <url>https://www.last.fm/tag/seen+live</url> </tag> </toptags> </lfm>
需要过滤掉name为"seen live"的标签,仅获取其余标签的name值,预期结果为:pop、dance、female vocalists、electronic。尝试过"//tag[@name != 'seen live']"和"//tag['@name != seen live']"等写法,但都无法实现过滤效果。
解决方案
问题出在错误使用了@name(用于选择标签属性),但XML里的name是子节点而非属性。正确的XPath写法如下:
//tag[name != 'seen live']/name
如果需要更精确的节点定位,也可以写成:
/lfm/toptags/tag[name != 'seen live']/name
说明
@name是针对标签属性的语法(比如<tag name="pop">这种结构),但你的XML中name是独立的子节点,因此直接用name引用该子节点。name != 'seen live'用于筛选出子节点name的文本不等于"seen live"的tag元素,最后通过/name提取对应的name节点内容。
内容的提问来源于stack exchange,提问作者Taylor F
相关产品推荐
相关产品推荐

