如何将API获取的嵌套XML投票数据转换为R语言的Tibble?
解决方案:将智利众议院XML投票数据转换为目标Tibble结构
所需工具
使用xml2解析XML数据,tidyverse套件(含tibble、dplyr、stringr)完成数据清洗与格式转换。
代码实现
- 安装并加载依赖包
# 首次运行需安装包 install.packages(c("xml2", "tidyverse")) # 加载包 library(xml2) library(tidyverse)
- 读取并解析XML数据
# 从API读取XML xml_doc <- read_xml("https://opendata.camara.cl/camaradiputados/WServices/WSLegislativo.asmx/retornarVotacionDetalle?prmVotacionId=15446") # 提取XML命名空间(必须指定才能正确查找节点) ns <- xml_ns(xml_doc)
- 提取数据并转换为目标Tibble
# 遍历所有<Voto>节点,提取信息并合并 votacion_tibble <- map_dfr(xml_find_all(xml_doc, ".//d1:Voto", ns = ns), function(node) { # 提取议员基础信息 diputado_node <- xml_find_first(node, ".//d1:Diputado", ns = ns) id <- xml_text(xml_find_first(diputado_node, ".//d1:Id", ns = ns)) nombre_completo <- xml_text(xml_find_first(diputado_node, ".//d1:Nombre", ns = ns)) apellido_paterno <- xml_text(xml_find_first(diputado_node, ".//d1:ApellidoPaterno", ns = ns)) apellido_materno <- xml_text(xml_find_first(diputado_node, ".//d1:ApellidoMaterno", ns = ns)) # 提取投票选项 opcion_voto <- xml_text(xml_find_first(node, ".//d1:OpcionVoto", ns = ns)) # 处理Nombre字段:提取最后一个词(如"Carlos Abel"转为"Abel") nombre <- word(nombre_completo, -1) # 构造单行数据 tibble( Id = id, Nombre = nombre, ApellidoPaterno = apellido_paterno, ApellidoMaterno = apellido_materno, OpcionVoto = opcion_voto ) })
- 查看结果
print(votacion_tibble)
输出结果与目标结构一致:
| Id | Nombre | ApellidoPaterno | ApellidoMaterno | OpcionVoto |
|---|---|---|---|---|
| 810 | Gabriel | Ascencio | Mansilla | Afirmativo |
| 855 | Abel | Jarpa | Wevar | Afirmativo |
| 862 | Pablo | Lorenzini | Basso | En Contra |
| 898 | Gabriel | Silber | Romo | Afirmativo |
关键说明
- 命名空间处理:原XML包含默认命名空间,必须通过
xml_ns()提取并在节点查找时指定,否则无法定位目标元素。 - Nombre字段清洗:使用
stringr::word()提取全名最后一个词,匹配需求中的格式。
内容的提问来源于stack exchange,提问作者cristobalvch
相关产品推荐
相关产品推荐

