R中使用rvest获取名称含.的XML节点返回空集的解决方法
问题背景
现有如下结构的XML文件:
<timeseries> <mrid>1</mrid> <businesstype>A96</businesstype> <flowdirection.direction>A02</flowdirection.direction> <quantity_measure_unit.name>MAW</quantity_measure_unit.name> </timeseries>
使用rvest包的函数查询节点时,查询名称带.的flowdirection.direction节点返回结果为{xml_nodeset (0)},但查询不含.的businesstype节点可正常返回结果,问题根源为节点名中的.字符,需要在沿用当前所用函数的前提下解决该问题。
当前使用的查询代码如下:
library(rvest) xml %>% html_elements("flowdirection.direction") %>% html_text() xml %>% html_nodes("flowdirection.direction")
故障原因
html_elements和html_nodes默认使用CSS选择器语法匹配节点,而在CSS选择器规则中,.是类选择器的特殊保留字符:写法flowdirection.direction会被解析为查找标签名为flowdirection、且携带direction类名的节点,而非查找标签名本身为flowdirection.direction的节点,因此匹配不到目标结果。
解决方案
沿用原有html_elements/html_nodes函数的前提下,只需要对节点名中的.做CSS转义即可。注意R语言字符串中反斜杠属于转义字符,因此表示CSS转义需要的单个反斜杠时,要在字符串内写双反斜杠:
library(rvest) # 转义节点名中的.即可正常匹配 xml %>% html_elements("flowdirection\\.direction") %>% html_text() xml %>% html_nodes("flowdirection\\.direction") # 其他带.的节点使用相同规则即可,例如查询quantity_measure_unit.name xml %>% html_elements("quantity_measure_unit\\.name") %>% html_text()
内容的提问来源于stack exchange,提问作者Priit Mets
相关产品推荐
相关产品推荐

