合并不同值类型RDS文件时dplyr::bind_rows报错求助
解决RDS文件合并时
bind_rows类型不匹配报错 问题原因
你脚本里的map_dfr(pluck, "performance")会调用dplyr::bind_rows合并所有提取出的performance数据框,但不同RDS文件中的performance数据框存在同一列类型不一致的情况(比如某列在部分文件里是整数,另一部分是字符/因子,或者数值类型不统一),导致合并失败。
解决步骤
1. 先定位类型不一致的列
先提取所有performance数据框,检查每列的类型分布:
# 提取所有performance数据框 performance_dfs <- iterative_run_ml_results %>% map(pluck, "performance") # 生成各文件列类型的统计表格 type_summary <- performance_dfs %>% map(~sapply(.x, class)) %>% bind_rows(.id = "rds_file") %>% pivot_longer(-rds_file, names_to = "column_name", values_to = "data_type") %>% group_by(column_name) %>% summarise(unique_types = n_distinct(data_type), types = paste(unique(data_type), collapse = "/")) %>% filter(unique_types > 1) print(type_summary)
运行这段代码后,就能看到哪些列存在多种数据类型,以及具体是哪些类型冲突。
2. 统一列类型后再合并
根据上面找到的问题列,统一转换为兼容的类型,再进行合并:
方案一:自动兼容转换(适合大部分场景)
利用type.convert自动将列转换为最合适的兼容类型:
# 统一所有数据框的列类型 unified_performance <- performance_dfs %>% map(~.x %>% mutate(across(everything(), type.convert, as.is = TRUE))) # 合并并输出 unified_performance %>% bind_rows() %>% write_tsv(glue("{root}_performance.tsv"))
方案二:手动指定类型(更可控,适合明确冲突类型的场景)
比如假设你发现accuracy列有整数/双精度冲突,model列有因子/字符冲突,就手动转换:
unified_performance <- performance_dfs %>% map(~.x %>% mutate( accuracy = as.double(accuracy), # 统一为双精度数值 model = as.character(model) # 统一为字符型 )) unified_performance %>% bind_rows() %>% write_tsv(glue("{root}_performance.tsv"))
额外提示
如果trained_model部分后续也出现类似问题,同样可以用这个思路先检查列类型,再统一转换后处理。
内容的提问来源于stack exchange,提问作者Seda Koldas
相关产品推荐
相关产品推荐

