如何在不等长嵌套列表场景下利用list_rbind()的names_to实现数据展开?
高效展开配对向量与不等长列表为数据框
我有两个从JSON获取的大型对象,原始数据顺序至关重要。其中species(对应真实场景的PG列)的长度和colors(对应真实场景的npi列表列)完全一致,允许新增列或使用行名。原本想借助list_rbind()的names_to参数处理,但遇到以下问题:
list_rbind()无法直接使用,因为colors不是数据框列表;as.data.frame()行不通,因为colors包含长度不一的子列表;- 直接调用
unlist()会丢失配对信息。
示例代码
species <- c("roses", "tulips", "lilies") colors <- list(list("red"), list("white", "yellow"), list("pink", "white"))
期望输出
species colors 1 roses red 2 tulips white 3 tulips yellow 4 lilies pink 5 lilies white
真实场景数据情况
实际场景涉及百万级行数据(示例中为1206169条),每个PG对应的npi列表至少包含1个元素,平均有8个元素,且数据并非字符串类型。数据结构如下:
> str(pg, max.level = 2) 'data.frame': 1206169 obs. of 3 variables: $ PG : int 1 2 3 4 5 6 7 8 9 10 ... $ npi:List of 1206169 ..$ : int 1376032029 1184159188 1629504501 1598703019 1487200408 1801443619 ..$ : int 1588809248 ..$ : int 1497791297
核心需求
用for循环暴力实现虽能得到结果,但面对百万级数据效率过低,需要更高效的处理方案。
内容的提问来源于stack exchange,提问作者Robert Hadow
相关产品推荐
相关产品推荐

