使用map_df读取CSV时添加文件创建日期列的问题排查
问题:合并CSV文件时添加文件创建日期列及类型转换警告处理
背景
有一批CSV文件,单文件数据结构如下:
| participant | rt | response | stimulus | condition |
|---|---|---|---|---|
| 1 | 112 | 'e' | star.jpg | congruent |
| 1 | 150 | 'i' | diamond.jpg | congruent |
原本使用以下代码读取并合并所有文件:
df <- list.files(path = "C:/file_location", pattern = "*.csv") %>% map_df(~read_csv(.))
第一次尝试:使用imap_dfc报错
尝试添加文件创建日期列时,使用了以下代码:
df <- list.files(path = "C:/data location/", pattern = "*.csv") %>% imap_dfc(~read_csv(.)%>% mutate(Date_Created = file.info(.x)$ctime))
触发报错:
Error in 'dplyr::bind_cols()': ! Can't recycle '..1' (size 1400) to match '..2'(size 1425). Backtrace: 1. ... %>% map_df(bind_cols) 2. purrr::imap_dfc(., ~read_csv(.) %>% mutate(Date_Created = file.info(.x)$ctime)) 3. dplyr::bind_cols(res)
报错原因:imap_dfc是按列绑定数据框,但不同CSV文件的行数不一致,无法按列对齐合并,应使用按行合并的函数。
第二次尝试:map_df成功但出现类型转换警告
改用以下代码后数据框生成成功,但出现警告:
dfiles<- list.files(path = "C:/file location", pattern = "*.csv", full.names = TRUE) df<- purrr::map_df(L2ufiles, ~read_csv(.x)%>% mutate(EndDate = file.info(.x)$ctime, rt=as.double(rt))) #needed to be added but seems to do nothing?
警告信息:
Warning: There was 1 warning in mutate(). ℹ In argument: rt = as.double(rt). Caused by warning: ! NAs introduced by coercion
警告原因:rt列中存在非数值类型的内容,执行as.double(rt)时这些非数值内容被转换为NA。
解决方案
1. 正确合并并添加文件创建日期
使用map_df(或imap_dfr)按行合并,同时传入完整文件路径以确保file.info能正确获取创建时间:
# 获取带完整路径的CSV文件列表 dfiles <- list.files(path = "C:/file location", pattern = "*.csv", full.names = TRUE) # 按行合并所有文件,添加创建日期列 df <- purrr::map_df(dfiles, function(file_path) { read_csv(file_path) %>% mutate( EndDate = file.info(file_path)$ctime, # 可选:保留原始rt列,新增转换后的列避免覆盖 rt_num = as.double(rt) ) })
用imap_dfr简化代码的版本:
df <- list.files(path = "C:/file location", pattern = "*.csv", full.names = TRUE) %>% imap_dfr(~read_csv(.x) %>% mutate(EndDate = file.info(.x)$ctime))
2. 处理rt列的类型转换警告
- 定位异常数据:查看转换后产生NA的行,找到原文件中的异常值
# 查看rt_num为NA的行 df %>% filter(is.na(rt_num)) %>% select(participant, rt, stimulus) - 提前指定列类型:读取文件时直接指定
rt列为数值型,读取阶段就会暴露异常df <- purrr::map_df(dfiles, function(file_path) { read_csv( file_path, col_types = cols( participant = col_integer(), rt = col_double(), response = col_character(), stimulus = col_character(), condition = col_character() ) ) %>% mutate(EndDate = file.info(file_path)$ctime) }) - 批量修复异常值:如果确定
rt列的非数值是输入错误,可提取数字部分后再转换df <- purrr::map_df(dfiles, function(file_path) { read_csv(file_path) %>% mutate( EndDate = file.info(file_path)$ctime, # 只保留rt中的数字部分再转换 rt = as.double(stringr::str_extract(rt, "^\\d+")) ) })
内容的提问来源于stack exchange,提问作者fuskerqq
相关产品推荐
相关产品推荐

