在R中实现数据框每行转为独立列的通用解决方案
问题与解决方案
需求说明
我在政府部门工作,他们偏好杂乱的数据格式。我手里有多个整洁的数据集(示例如下),需要转换成这种杂乱格式来适配他们指定的Excel模板。这个需求有点像pivot_wider但又不完全一样。我自己写了示例代码实现了目标,但这个方法只能处理固定行数的数据,想找个支持任意行数的通用方法。
示例数据与原始代码
library(dplyr) #> #> Attaching package: 'dplyr' #> The following objects are masked from 'package:stats': #> #> filter, lag #> The following objects are masked from 'package:base': #> #> intersect, setdiff, setequal, union dat <- structure(list(`Industry (aggregate)` = c("Public Administration", "Health Care And Social Assistance", "Educational Services", "Professional, Scientific And Technical Services", "Repair, Personal And Non-Profit Services" ), jo = c(530, 70, 60, 10, 10), `%` = c(0.78, 0.1, 0.09, 0.01, 0.01)), row.names = c(NA, -5L), class = c("tbl_df", "tbl", "data.frame" )) # 原始仅支持固定行数的函数 splatten <- function(tbbl){ first <- tbbl[1,] second <- tbbl[2,] third <- tbbl[3,] fourth <- tbbl[4,] fifth <- tbbl[5,] bind_cols(first, second, third, fourth, fifth) } splatten(dat) #> New names: #> • `Industry (aggregate)` -> `Industry (aggregate)...1` #> • `jo` -> `jo...2` #> • `%` -> `%...3` #> • `Industry (aggregate)` -> `Industry (aggregate)...4` #> • `jo` -> `jo...5` #> • `%` -> `%...6` #> • `Industry (aggregate)` -> `Industry (aggregate)...7` #> • `jo` -> `jo...8` #> • `%` -> `%...9` #> • `Industry (aggregate)` -> `Industry (aggregate)...10` #> • `jo` -> `jo...11` #> • `%` -> `%...12` #> • `Industry (aggregate)` -> `Industry (aggregate)...13` #> • `jo` -> `jo...14` #> • `%` -> `%...15` #> # A tibble: 1 × 15 #> Industry (aggregate)...…¹ jo...2 `%...3` Industry (aggregate)…² jo...5 `%...6` #> <chr> <dbl> <dbl> <chr> <dbl> <dbl> #> 1 Public Administration 530 0.78 Health Care And Socia… 70 0.1 #> # ℹ abbreviated names: ¹`Industry (aggregate)...1`, ²`Industry (aggregate)...4` #> # ℹ 9 more variables: `Industry (aggregate)...7` <chr>, jo...8 <dbl>, #> # `%...9` <dbl>, `Industry (aggregate)...10` <chr>, jo...11 <dbl>, #> # `%...12` <dbl>, `Industry (aggregate)...13` <chr>, jo...14 <dbl>, #> # `%...15` <dbl>
通用解决方案
可以借助purrr包的遍历功能,动态处理任意行数的数据集。核心思路是把每一行单独提取为一个tibble,然后将所有tibble按列绑定:
library(dplyr) library(purrr) splatten_general <- function(tbbl){ # 遍历每一行,提取为单独的tibble row_list <- map(seq_len(nrow(tbbl)), ~tbbl[., ]) # 按列绑定所有行对应的tibble bind_cols(row_list) } # 测试通用函数 splatten_general(dat)
这个函数会自动处理任意行数的输入数据,生成和原始函数一致的杂乱格式输出,同时不需要硬编码行数。
内容的提问来源于stack exchange,提问作者Richard Martin
相关产品推荐
相关产品推荐

