如何在R中展开XML转换后dataframe里的location_1字典列?
解决XML转DataFrame时location_1字段解析问题
针对你在将指定XML转换为R DataFrame时遇到的location_1列解析为NA的问题,可以通过针对性展开嵌套结构来解决。修正后的代码如下:
library(xml2) library(tidyverse) fileurl <- "https://d396qusza40orc.cloudfront.net/getdata%2Fdata%2Frestaurants.xml" xmllist <- as_list(read_xml(fileurl)) xml_df <- tibble::as_tibble(xmllist) %>% # 展开response列表为单独行 unnest_longer(response) %>% # 展开response中的字段为列 unnest_wider(response) %>% # 专门展开location_1的嵌套键值对,用names_sep避免列名冲突 unnest_wider(location_1, names_sep = "_") %>% # 自动转换各列数据类型 readr::type_convert()
关键说明
- 原代码中两次
unnest(cols = names(.))会粗暴展开所有列,但location_1是嵌套的键值对结构,需要用unnest_wider精准处理。 names_sep = "_"参数会将location_1内的子字段(如address、city等)命名为location_1_address、location_1_city,避免和其他同名字段冲突。- 执行完上述代码后,
location_1内的所有值都会被提取为独立列,不再出现NA。
内容的提问来源于stack exchange,提问作者DataGwynn
相关产品推荐
相关产品推荐

