在R语言中提取指定行数据并转换为DataFrame的技术求助
问题描述
我有一个包含10万+行数据的txt文件,希望将其转换为DataFrame,但仅需提取以TI、AU、PD、AB开头的行,分别对应标题、作者、日期、摘要列。数据示例如下:
FN Clarivate Analytics Web of Science VR 1.0 PT J AU Yang, Qiang Liu, Yang Chen, Tianjian Tong, Yongxin TI Federated Machine Learning: Concept and Applications SO ACM TRANSACTIONS ON INTELLIGENT SYSTEMS AND TECHNOLOGY VL 10 IS 2 AR 12 DI 10.1145/3298981 DT Article PD FEB 2019 PY 2019 AB Today's artificial intelligence still faces two major challenges (...) etc.
我尝试使用以下read.table代码,但出现报错,无法实现需求,尤其不清楚如何让R将整句作为变量而非拆分单词,也不知如何筛选指定行:
read.table("groupprojectdatabase.txt", header = FALSE, sep = ",", quote = "", dec = ".", numerals = c("allow.loss"), row.names = c("TI", "AU", "PB","AB"), col.names = c('title_col','author_col','date_col','summary_col'), as.is = !stringsAsFactors, na.strings = "NA", colClasses = NA, nrows = -1, skip = 0, check.names = TRUE, fill = FALSE, strip.white = FALSE, blank.lines.skip = TRUE, comment.char = "#", allowEscapes = FALSE, flush = FALSE, stringsAsFactors = FALSE, fileEncoding = "", encoding = "unknown", text, skipNul = FALSE)
报错信息:
Error in scan(file = file, what = what, sep = sep, quote = quote, dec = dec, : line 1 did not have 4 elements
恳请提供可行方法或相关函数指导。
解决方案
报错原因说明
你使用的read.table是按指定分隔符拆分列的结构化读取工具,但你的数据是标签-值的键值对格式,且存在同标签多行续值(比如AU的多行作者),直接用该函数完全不匹配数据结构,因此报错。
方法1:基础R处理(内存友好,适合大文件)
核心逻辑是逐行读取文件,识别标签并合并同标签的多行内容:
# 初始化存储列表 result_list <- list(AU = character(), TI = character(), PD = character(), AB = character()) current_tag <- NULL current_content <- NULL # 打开文件连接(逐行读取,减少内存占用) con <- file("groupprojectdatabase.txt", "r") while (length(line <- readLines(con, n = 1)) > 0) { line_trim <- trimws(line) if (nchar(line_trim) == 0) next # 判断是否为目标标签行 tag <- substr(line_trim, 1, 2) if (tag %in% c("AU", "TI", "PD", "AB")) { # 保存上一组标签的内容 if (!is.null(current_tag)) { result_list[[current_tag]] <- c(result_list[[current_tag]], paste(current_content, collapse = "; ")) } # 更新当前标签和内容 current_tag <- tag current_content <- trimws(substr(line_trim, 4, nchar(line_trim))) } else if (!is.null(current_tag) && grepl("^\\s+", line_trim)) { # 处理同标签的续行(以空格开头) current_content <- c(current_content, trimws(line_trim)) } else { # 遇到非目标标签,保存当前内容 if (!is.null(current_tag)) { result_list[[current_tag]] <- c(result_list[[current_tag]], paste(current_content, collapse = "; ")) current_tag <- NULL current_content <- NULL } } } # 保存最后一组未处理的内容 if (!is.null(current_tag)) { result_list[[current_tag]] <- c(result_list[[current_tag]], paste(current_content, collapse = "; ")) } close(con) # 转换为DataFrame,补全长度不一致的列 max_len <- max(sapply(result_list, length)) df <- data.frame( author_col = if (length(result_list$AU) < max_len) c(result_list$AU, rep(NA, max_len - length(result_list$AU))) else result_list$AU, title_col = if (length(result_list$TI) < max_len) c(result_list$TI, rep(NA, max_len - length(result_list$TI))) else result_list$TI, date_col = if (length(result_list$PD) < max_len) c(result_list$PD, rep(NA, max_len - length(result_list$PD))) else result_list$PD, summary_col = if (length(result_list$AB) < max_len) c(result_list$AB, rep(NA, max_len - length(result_list$AB))) else result_list$AB, stringsAsFactors = FALSE )
方法2:tidyverse工具(代码简洁)
若已安装tidyverse包,可通过更简洁的代码实现:
library(tidyverse) # 读取所有行并预处理 lines <- read_lines("groupprojectdatabase.txt") %>% trimws() %>% .[nchar(.) > 0] # 转换为DataFrame df <- tibble(line = lines) %>% # 标记标签行,为续行填充对应标签 mutate(tag = ifelse(str_detect(line, "^[A-Z]{2} "), substr(line, 1, 2), NA)) %>% fill(tag, .direction = "down") %>% # 筛选目标标签,提取内容 filter(tag %in% c("AU", "TI", "PD", "AB")) %>% mutate(content = trimws(str_remove(line, "^[A-Z]{2} "))) %>% # 合并同标签的多行内容 group_by(tag) %>% summarise(content = str_c(content, collapse = "; ")) %>% # 转换为宽表并重命名列 pivot_wider(names_from = tag, values_from = content) %>% rename(author_col = AU, title_col = TI, date_col = PD, summary_col = AB)
关键提示
- 两种方法都处理了同标签续行问题,用分号分隔合并内容
- 基础R方法更适合10万+行的超大型文件,逐行读取不会占用过多内存
- tidyverse方法代码更易读,但需要加载相关包
内容的提问来源于stack exchange,提问作者Flynn Howl
相关产品推荐
相关产品推荐

