如何在R语言中从苹果健康导出的XML文档中筛选指定属性对应的value值
提取苹果健康XML数据的正确方法
看起来你之前的代码没拿到数据是因为没找对节点和属性的提取方式,用xml2包的话,我们需要先定位到目标Record节点,再提取对应的属性值,下面是具体的解决办法:
提取身高数据(HKQuantityTypeIdentifierHeight)
首先定位所有类型为身高的Record节点,然后提取value属性,还可以把相关的单位、日期等信息一起整理成数据框:
library(xml2) library(tibble) # 用来创建数据框,也可以用base R的data.frame # 你已经读取好的xml对象 xml <- read_xml('<?xml version="1.0" encoding="UTF-8"?> <HealthData locale="en_US"> <ExportDate value="2021-10-29 20:24:26 -0700"/> <Me HKCharacteristicTypeIdentifierDateOfBirth="1975-05-18" HKCharacteristicTypeIdentifierBiologicalSex="HKBiologicalSexMale" HKCharacteristicTypeIdentifierBloodType="HKBloodTypeNotSet" HKCharacteristicTypeIdentifierFitzpatrickSkinType="HKFitzpatrickSkinTypeNotSet" HKCharacteristicTypeIdentifierCardioFitnessMedicationsUse="None"/> <Record type="HKQuantityTypeIdentifierHeight" sourceName="Neil’s Apple Watch" sourceVersion="2.1" unit="ft" creationDate="2016-02-23 14:12:38 -0700" startDate="2016-02-23 14:12:38 -0700" endDate="2016-02-23 14:12:38 -0700" value="6.16667"/> <Record type="HKQuantityTypeIdentifierHeight" sourceName="Neil’s Apple Watch" sourceVersion="2.2.1" unit="ft" creationDate="2016-08-17 08:00:37 -0700" startDate="2016-08-17 08:00:37 -0700" endDate="2016-08-17 08:00:37 -0700" value="6.16667"/> <Record type="HKQuantityTypeIdentifierBodyMass" sourceName="Neil’s Apple Watch" sourceVersion="2.1" unit="lb" creationDate="2016-02-23 14:12:38 -0700" startDate="2016-02-23 14:12:38 -0700" endDate="2016-02-23 14:12:38 -0700" value="175"/> <Record type="HKQuantityTypeIdentifierBodyMass" sourceName="Neil’s Apple Watch" sourceVersion="2.2.1" unit="lb" creationDate="2016-08-17 08:00:36 -0700" startDate="2016-08-17 08:00:36 -0700" endDate="2016-08-17 08:00:36 -0700" value="180"/> <Record type="HKQuantityTypeIdentifierBodyMass" sourceName="Neils Apple Watch" sourceVersion="2.1" unit="lb" creationDate="2016-02-23 14:12:38 -0700" startDate="2016-02-23 14:12:38 -0700" endDate="2016-02-23 14:12:38 -0700" value="175"/> </HealthData>') # 定位所有身高Record节点 height_nodes <- xml_find_all(xml, "//Record[@type='HKQuantityTypeIdentifierHeight']") # 提取value属性并转为数值型 listOfAllMyHeights <- as.numeric(xml_attr(height_nodes, "value")) print(listOfAllMyHeights) # 更完整的做法:整理成数据框,包含更多信息 height_df <- tibble( height_ft = as.numeric(xml_attr(height_nodes, "value")), unit = xml_attr(height_nodes, "unit"), creation_date = as.POSIXct(xml_attr(height_nodes, "creationDate")), source = xml_attr(height_nodes, "sourceName") ) print(height_df)
提取体重数据(HKQuantityTypeIdentifierBodyMass)
和身高的处理逻辑完全一致,只需要修改XPath中的type值:
# 定位所有体重Record节点 weight_nodes <- xml_find_all(xml, "//Record[@type='HKQuantityTypeIdentifierBodyMass']") # 提取value属性 listOfAllMyWeights <- as.numeric(xml_attr(weight_nodes, "value")) print(listOfAllMyWeights) # 整理成数据框 weight_df <- tibble( weight_lb = as.numeric(xml_attr(weight_nodes, "value")), unit = xml_attr(weight_nodes, "unit"), creation_date = as.POSIXct(xml_attr(weight_nodes, "creationDate")), source = xml_attr(weight_nodes, "sourceName") ) print(weight_df)
为什么之前的代码返回空值?
你之前用xml_text()是提取节点的文本内容,但苹果健康的数值是存在Record节点的value属性里的,不是节点内部的文本。所以必须先用xml_find_all()定位到目标节点,再用xml_attr()提取对应属性的值。另外,XPath表达式//Record[@type='XXX']能帮我们精准筛选出所有类型匹配的记录,不管它在XML结构的哪个层级。
内容的提问来源于stack exchange,提问作者Neil Cronin
相关产品推荐
相关产品推荐

