R语言中如何将数据框每行的嵌套列表展开为列?
问题
我从细菌生长条件数据库获取了一个约9万行的R数据框,每行对应一个细菌物种。但每行都存在多层嵌套列表(示例如下),我希望将这些嵌套列表“展开”为数据框的列,尝试过Tidyverse工具但未成功。
示例嵌套结构:
test_data[[1]][["Safety information"]] .. .. ..$ DSM-Number : int 4491 .. .. ..$ keywords :List of 4 .. .. .. ..$ : chr "Bacteria" .. .. .. ..$ : chr "16S sequence" .. .. .. ..$ : chr "genome sequence" .. .. .. ..$ : chr "mesophilic" .. .. ..$ description : chr "Acetobacter lovaniensis DSM 4491 is a mesophilic bacterium that was isolated from soil." .. .. ..$ NCBI tax id :List of 2 .. .. .. ..$ NCBI tax id : int 104100 .. .. .. ..$ Matching level: chr "species" .. .. ..$ strain history:List of 2 .. .. .. ..$ : chr "<- NCIMB <- W. Verhoeven <- J. Frateur" .. .. .. ..$ : chr "DSM 4491 <-- NCIMB 8620 <-- W. Verhoeven L 1024 <-- J. Frateur." .. .. ..$ doi : chr "10.13145/bacdive9.20220920.7" .. ..$ Name and taxonomic classification :List of 11 .. .. ..$ LPSN :List of 12 .. .. .. ..$ @ref : int 20215 .. .. .. ..$ description : chr "domain/bacteria" .. .. .. ..$ keyword : chr "phylum/pseudomonadota" .. .. .. ..$ domain : chr "Bacteria" .. .. .. ..$ phylum : chr "Pseudomonadota" .. .. .. ..$ class : chr "Alphaproteobacteria" .. .. .. ..$ order : chr "Rhodospirillales" .. .. .. ..$ family : chr "Acetobacteraceae" .. .. .. ..$ genus : chr "Acetobacter" .. .. .. ..$ species : chr "Acetobacter lovaniensis" .. .. .. ..$ full scientific name: chr "<I>Acetobacter</I> <I>lovaniensis</I> (Frateur 1950) Lisdiyanti et al. 2001" .. .. .. ..$ synonyms :List of 2 .. .. .. .. ..$ :List of 2 .. .. .. .. .. ..$ @ref : int 20215 .. .. .. .. .. ..$ synonym: chr "Acetobacter pasteurianus subsp. lovaniensis" .. .. .. .. ..$ :List of 2 .. .. .. .. .. ..$ @ref : int 20215 .. .. .. .. .. ..$ synonym: chr "Acetobacter lovaniense" .. .. ..$ @ref : int 1703 .. .. ..$ domain : chr "Bacteria" .. .. ..$ phylum : chr "Proteobacteria" .. .. ..$ class : chr "Alphaproteobacteria" .. .. ..$ order : chr "Rhizobiales" .. .. ..$ family : chr "Acetobacteraceae" .. .. ..$ genus : chr "Acetobacter" .. .. ..$ species : chr "Acetobacter lovaniensis" .. .. ..$ full scientific name: chr "Acetobacter lovaniensis (Frateur 1950) Lisdiyanti et al. 2001" .. .. ..$ type strain : chr "yes" .. ..$ Morphology : Named list() .. ..$ Culture and growth conditions :List of 2 .. .. ..$ culture medium:List of 2 .. .. .. ..$ :List of 5 .. .. .. .. ..$ @ref : int 1703 .. .. .. .. ..$ name : chr "YPM MEDIUM (DSMZ Medium 360)" .. .. .. .. ..$ growth : chr "yes" .. .. .. .. ..$ link : chr "https://bacmedia.dsmz.de/medium/360" .. .. .. .. ..$ composition: chr "Name: YPM MEDIUM (DSMZ Medium 360)\nComposition:\nMannitol 25.0 g/l\nAgar 12.0 g/l\nYeast extract 5.0 g/l\nPept"| __truncated__ .. .. .. ..$ :List of 5 .. .. .. .. ..$ @ref : int 1703 .. .. .. .. ..$ name : chr "GLUCONOBACTER OXYDANS MEDIUM (DSMZ Medium 105)" .. .. .. .. ..$ growth : chr "yes" .. .. .. .. ..$ link : chr "https://bacmedia.dsmz.de/medium/105" .. .. .. .. ..$ composition: chr "Name: GLUCONOBACTER OXYDANS MEDIUM (DSMZ Medium 105)\nComposition:\nGlucose 100.0 g/l\nCaCO3 20.0 g/l\nAgar 15."| __truncated__ .. .. ..$ culture temp :List of 2 .. .. .. ..$ :List of 5 .. .. .. .. ..$ @ref : int 1703 .. .. .. .. ..$ growth : chr "positive" .. .. .. .. ..$ type : chr "growth" .. .. .. .. ..$ temperature: chr "28" .. .. .. .. ..$ range : chr "mesophilic" .. .. .. ..$ :List of 5 .. .. .. .. ..$ @ref : int 67770 .. .. .. .. ..$ growth : chr "positive" .. .. .. .. ..$ type : chr "growth" .. .. .. .. ..$ temperature: chr "28" .. .. .. .. ..$ range : chr "mesophilic"
解决方案
1. 加载必要工具包
核心使用Tidyverse系列包处理嵌套结构:
library(tidyverse)
2. 读取数据
读取本地存储的测试数据:
test_data <- readRDS("test_data.rds")
3. 分层展开嵌套结构
多层嵌套需逐步处理,避免一次性展开导致数据混乱:
步骤1:展开顶层主列表项
先将数据框的顶层列表列展开为独立列,用names_sep避免列名冲突:
df <- test_data %>% unnest_wider(cols = everything(), names_sep = "_")
步骤2:处理无命名嵌套列表(如keywords、菌株历史)
这类无命名列表若拆分为多行会大幅增加数据量,建议合并为分隔符分隔的字符串:
df <- df %>% mutate( Safety_information_keywords = map_chr(Safety_information_keywords, ~paste(.x, collapse = "; ")), Safety_information_strain_history = map_chr(Safety_information_strain_history, ~paste(.x, collapse = "; ")) )
步骤3:展开命名嵌套列表(如NCBI分类ID)
对于有明确命名的嵌套子项,直接用unnest_wider展开为单独列:
df <- df %>% unnest_wider(cols = Safety_information_NCBI_tax_id, names_sep = "_")
步骤4:处理分类信息的深层嵌套
分类信息包含LPSN子列表和同义词,先展开LPSN,再合并同义词为字符串:
# 展开分类信息顶层 df <- df %>% unnest_wider(cols = Name_and_taxonomic_classification, names_sep = "_") # 展开LPSN字段 df <- df %>% unnest_wider(cols = Name_and_taxonomic_classification_LPSN, names_sep = "_") # 合并同义词为字符串 df <- df %>% mutate( Name_and_taxonomic_classification_LPSN_synonyms = map_chr(Name_and_taxonomic_classification_LPSN_synonyms, ~paste(map_chr(.x, "synonym"), collapse = "; ")) )
步骤5:处理培养条件的嵌套结构
培养条件包含培养基和温度列表,同样合并关键信息为字符串:
# 展开培养条件顶层 df <- df %>% unnest_wider(cols = Culture_and_growth_conditions, names_sep = "_") # 合并培养基名称和温度信息 df <- df %>% mutate( Culture_and_growth_conditions_culture_medium = map_chr(Culture_and_growth_conditions_culture_medium, ~paste(map_chr(.x, "name"), collapse = "; ")), Culture_and_growth_conditions_culture_temp = map_chr(Culture_and_growth_conditions_culture_temp, ~paste(map_chr(.x, "temperature"), collapse = "; ")) )
4. 整理最终数据
清理列名并移除全空列:
# 替换列名中的特殊字符 df <- df %>% rename_with(~gsub("@ref", "ref", .x)) %>% rename_with(~gsub("\\.| ", "_", .x)) # 移除全空列 df <- df %>% select(where(~!all(is.na(.x))))
关键提示
- 若需要保留无命名列表的每个元素为单独行,将
map_chr替换为unnest_longer即可,但9万行数据会大幅膨胀,需谨慎操作。 - 处理大样本时,建议先取10-20行测试代码逻辑,确认无误后再批量处理,避免内存溢出。
内容的提问来源于stack exchange,提问作者Patrick Kearns
相关产品推荐
相关产品推荐

