tidyr::unnest报错:无法合并double与character类型的exists列
问题:unnest展开数据时出现类型不兼容错误
我编写了一段R代码遍历包含2500个文件的文件夹,生成的flight_info_1列表包含path和data两个变量。执行unnest(flight_info_1, data)展开数据时,出现以下错误:
! Can't combine
..1$existsand ..350$exists
完整代码
library(prettydoc) library(tidyverse) library(fs) library(purrr) library(haven) library(gghalves) library(tidygraph) library(ggraph) library(GGally) library(vtable) library(ggridges) library(scales) library(ggpubr) library(highcharter) library(plotly) library(patchwork) library(ggdark) library(ggthemes) library(rnaturalearth) `%||%` <- rlang::`%||%` flight_info_1 <- "Path/Path/" %>% dir_ls(recurse = TRUE, type = "file") %>% as_tibble_col("path") %>% mutate(data = map(path, function(path_i){ path_i <<- path_i data_i <- read_csv(path_i, show_col_types=FALSE) # if(nrow(data_i)==0) # {return(NULL)} data_i$date <- as.character(data_i[["date"]] %||% NA_character_) moved_vars <- tibble(.rows=nrow(data_i)) if(isTRUE(str_detect(data_i[["airline"]] %||% "", "€"))) { moved_vars <- bind_cols(moved_vars, select(data_i, price = airline)) } if(isTRUE(str_detect(data_i[["price"]] %||% "", "hr|min"))){ moved_vars <- bind_cols(moved_vars, select(data_i, duration = price)) } if(isTRUE(str_detect(data_i[["emissions"]] %||% "", "[included]"))){ moved_vars <- bind_cols(moved_vars, select(data_i, luggage = emissions)) } if(isTRUE(str_detect(data_i[["luggage"]] %||% "", "(?i)kg"))){ moved_vars <- bind_cols(moved_vars, select(data_i, emissions = luggage)) } if(isTRUE(str_detect(data_i[["duration"]] %||% "", "^[[:alpha:][:space:]]+$"))){ moved_vars <- bind_cols(moved_vars, select(data_i, airline = duration)) } data_i_mod <- bind_cols( moved_vars, select(data_i, -any_of(colnames(moved_vars))) ) return(data_i_mod) }))
报错回溯
data_csv <- unnest(flight_info_1, data) Error: ! Can't combine `..1$exists` <double> and `..350$exists` <character>. Backtrace: 1. tidyr::unnest(flight_info_1, data) 2. tidyr:::unnest.data.frame(flight_info_1, data) 3. tidyr::unchop(data, any_of(cols), keep_empty = keep_empty, ptype = ptype) 4. tidyr:::df_unchop(cols, ptype = ptype, keep_empty = keep_empty) 5. vctrs::vec_unchop(col, ptype = col_ptype) 6. vctrs (local) `<fn>`() 7. vctrs::vec_default_ptype2(...) 8. vctrs::stop_incompatible_type(...) 9. vctrs:::stop_incompatible(...) 10. vctrs:::stop_vctrs(...)
解决思路
1. 明确问题根源
报错核心是不同CSV文件中的exists列数据类型不一致:部分文件中该列是数值型(double),另一部分是字符型(character),unnest合并时无法兼容不同类型的列。
2. 统一列类型
在读取每个文件后,强制将exists列转换为统一类型,推荐两种方式:
方式一:读取后统一转换
在read_csv之后添加处理代码,将exists列转为字符型(或根据业务逻辑转为布尔型):
data_i <- read_csv(path_i, show_col_types=FALSE) # 新增:统一处理exists列类型 if ("exists" %in% colnames(data_i)) { # 转为字符型,兼容所有情况 data_i$exists <- as.character(data_i$exists) # 若确认该列实际是布尔值,也可转为逻辑型: # data_i$exists <- as.logical(data_i$exists) }
方式二:读取时指定列类型
直接在read_csv中通过col_types参数指定exists列的类型,避免自动识别错误:
data_i <- read_csv(path_i, show_col_types=FALSE, col_types = cols( exists = col_character() # 或者col_logical(),根据实际数据选择 ))
3. 排查所有列的类型差异
如果除了exists还有其他列可能出现类型不一致,可以先批量提取所有文件的列信息排查:
# 提取每个文件的列名和对应类型 col_info <- flight_info_1 %>% mutate(col_details = map(data, ~tibble( colname = colnames(.x), col_type = sapply(.x, class) ))) %>% unnest(col_details) # 查看各列的类型分布 col_info %>% group_by(colname) %>% count(col_type) %>% filter(n() > 1) # 筛选出存在多种类型的列
4. 处理空文件
取消注释原代码中的空文件判断,返回空tibble而非NULL,避免unnest时出现异常:
if(nrow(data_i)==0) { return(tibble()) # 返回空tibble,保证结构一致 }
内容的提问来源于stack exchange,提问作者Gaaaa
相关产品推荐
相关产品推荐

