如何对full_join后的数据集进行拉长并合并对应列?
解决full_join后正确拉长数据集并合并对应列的问题
针对你遇到的合并后无法正确拉长列、补全缺失频率值的问题,这里提供两种高效的解决方案:
方法1:先合并原始数据集再补全缺失组合(推荐)
由于两个原始数据集结构完全一致,直接行合并后补全所有可能的组合,比先执行full_join再处理更简洁:
library(tidyverse) # 合并两个原始数据集 combined_df <- bind_rows(df1, df2) # 生成所有error-phase-instructions的组合,缺失的Freq填充0 final_df <- combined_df %>% complete(error, phase, instructions, fill = list(Freq = 0)) %>% arrange(error, phase, instructions)
说明
bind_rows直接将两个数据集的行合并,保留原有列结构;complete会自动识别error/phase/instructions的所有可能组合,为不存在的组合新增行,并通过fill参数将Freq设为0,完美覆盖你需要的缺失场景(比如C1 C2 C3 swap的implicit类型、CCVVC的explicit类型)。
方法2:处理已执行full_join后的数据集
如果已经得到带.x/.y后缀的合并数据集,可以通过pivot_longer的正则匹配功能拆分列名,再整理成目标格式:
final_df <- df3 %>% # 按列名后缀拆分并拉长 pivot_longer( cols = -error, names_to = c(".value", "source"), names_pattern = "(.*)\\.(x|y)" ) %>% # 移除临时的source列 select(-source) %>% # 补全所有组合并填充Freq为0 complete(error, phase, instructions, fill = list(Freq = 0)) %>% arrange(error, phase, instructions)
说明
names_pattern = "(.*)\\.(x|y)"通过正则表达式将列名拆分为两部分:比如phase.x会被解析为列名phase(对应.value,即保留该列的实际含义)和来源标识x(对应source,用于区分原始数据集);- 拉长后原本的
phase.x/phase.y会合并到同一phase列,Freq.x/Freq.y合并到同一Freq列,后续通过complete补全缺失组合即可。
两种方法最终都会生成你预期的输出格式,例如:
error phase Freq instructions C1 first 2 implicit C1 first 1 explicit C1 C2 C3 swap first 0 implicit C1 C2 C3 swap first 1 explicit C1 drop first 2 implicit C1 drop first 1 explicit ...
内容的提问来源于stack exchange,提问作者snd1101
相关产品推荐
相关产品推荐

