如何读取包含未加引号换行符列的CSV文件?
解析含未引号换行符的CSV文件问题
我有一份存在问题的CSV文件,最后一列的字符串未添加引号,但该列可能包含换行符(示例中第2行无换行符)。
构造示例CSV的代码:
library(tidyverse) csv_file <- str_c( "a,b,c\n", "1,1,first\nrow\n", "1,1,second row\n", "1,1,third\nrow\n", collapse = "" )
尝试使用read_csv()读取时会出现解析错误:
read_csv(csv_file) #> Warning: One or more parsing issues, call `problems()` on your data frame for details, #> e.g.: #> dat <- vroom(...) #> problems(dat) #> Rows: 5 Columns: 3 #> ── Column specification ──────────────────────────────────────────────────────── #> Delimiter: "," #> chr (2): a, c #> dbl (1): b #> #> ℹ Use `spec()` to retrieve the full column specification for this data. #> ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message. #> # A tibble: 5 × 3 #> a b c #> <chr> <dbl> <chr> #> 1 1 1 first #> 2 row NA <NA> #> 3 1 1 second row #> 4 1 1 third #> 5 row NA <NA>
需要解析出如下预期结果:
tibble( a = c(1, 1, 1), b = c(1, 1, 1), c = c("first\nrow", "second row", "third\nrow") ) #> # A tibble: 3 × 3 #> a b c #> <dbl> <dbl> <chr> #> 1 1 1 "first\nrow" #> 2 1 1 "second row" #> 3 1 1 "third\nrow"
解决方案
由于前两列均为数值类型,无逗号分隔符,可通过识别数据行的起始特征(数字,数字,),将属于同一数据行的续行合并,再进行读取:
方法1:循环处理行
# 将CSV内容拆分为单行 lines <- str_split(csv_file, "\n")[[1]] # 判断哪些行是续行(非数据起始行) is_continuation <- !str_detect(lines, "^\\d+,\\d+,") # 合并续行到对应上一行 merged_lines <- lines for (i in seq_along(lines)) { if (is_continuation[i] && i > 1) { merged_lines[i-1] <- str_c(merged_lines[i-1], lines[i]) merged_lines[i] <- "" } } # 过滤空行并合并为完整CSV字符串 fixed_csv <- str_c(merged_lines[merged_lines != ""], collapse = "\n") # 读取修正后的CSV read_csv(fixed_csv)
方法2:tidyverse管道式处理
fixed_csv <- csv_file %>% str_split("\n") %>% pluck(1) %>% tibble(line = .) %>% # 按数据起始行分组 mutate(group = cumsum(str_detect(line, "^\\d+,\\d+,"))) %>% group_by(group) %>% # 合并同组内的所有行 summarize(line = str_c(line, collapse = ""), .groups = "drop") %>% pull(line) %>% str_c(collapse = "\n") read_csv(fixed_csv)
两种方法均可得到符合预期的解析结果。
内容的提问来源于stack exchange,提问作者Peter H.
相关产品推荐
相关产品推荐

