You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中提取指定行数据并转换为DataFrame的技术求助

问题描述

我有一个包含10万+行数据的txt文件,希望将其转换为DataFrame,但仅需提取以TI、AU、PD、AB开头的行,分别对应标题、作者、日期、摘要列。数据示例如下:

FN Clarivate Analytics Web of Science
VR 1.0
PT J
AU Yang, Qiang
   Liu, Yang
   Chen, Tianjian
   Tong, Yongxin
TI Federated Machine Learning: Concept and Applications
SO ACM TRANSACTIONS ON INTELLIGENT SYSTEMS AND TECHNOLOGY
VL 10
IS 2
AR 12
DI 10.1145/3298981
DT Article
PD FEB 2019
PY 2019
AB Today's artificial intelligence still faces two major challenges (...) etc. 

我尝试使用以下read.table代码,但出现报错,无法实现需求,尤其不清楚如何让R将整句作为变量而非拆分单词,也不知如何筛选指定行:

read.table("groupprojectdatabase.txt", header = FALSE, sep = ",", quote = "",
           dec = ".", numerals = c("allow.loss"),
           row.names = c("TI", "AU", "PB","AB"), col.names = c('title_col','author_col','date_col','summary_col'), as.is = !stringsAsFactors,
           na.strings = "NA", colClasses = NA, nrows = -1,
           skip = 0, check.names = TRUE, fill = FALSE,
           strip.white = FALSE, blank.lines.skip = TRUE,
           comment.char = "#",
           allowEscapes = FALSE, flush = FALSE,
           stringsAsFactors = FALSE,
           fileEncoding = "", encoding = "unknown", text, skipNul = FALSE)

报错信息:

Error in scan(file = file, what = what, sep = sep, quote = quote, dec = dec,  : 
  line 1 did not have 4 elements

恳请提供可行方法或相关函数指导。


解决方案

报错原因说明

你使用的read.table是按指定分隔符拆分列的结构化读取工具,但你的数据是标签-值的键值对格式,且存在同标签多行续值(比如AU的多行作者),直接用该函数完全不匹配数据结构,因此报错。

方法1:基础R处理(内存友好,适合大文件)

核心逻辑是逐行读取文件,识别标签并合并同标签的多行内容:

# 初始化存储列表
result_list <- list(AU = character(), TI = character(), PD = character(), AB = character())
current_tag <- NULL
current_content <- NULL

# 打开文件连接(逐行读取,减少内存占用)
con <- file("groupprojectdatabase.txt", "r")
while (length(line <- readLines(con, n = 1)) > 0) {
  line_trim <- trimws(line)
  if (nchar(line_trim) == 0) next
  
  # 判断是否为目标标签行
  tag <- substr(line_trim, 1, 2)
  if (tag %in% c("AU", "TI", "PD", "AB")) {
    # 保存上一组标签的内容
    if (!is.null(current_tag)) {
      result_list[[current_tag]] <- c(result_list[[current_tag]], paste(current_content, collapse = "; "))
    }
    # 更新当前标签和内容
    current_tag <- tag
    current_content <- trimws(substr(line_trim, 4, nchar(line_trim)))
  } else if (!is.null(current_tag) && grepl("^\\s+", line_trim)) {
    # 处理同标签的续行(以空格开头)
    current_content <- c(current_content, trimws(line_trim))
  } else {
    # 遇到非目标标签,保存当前内容
    if (!is.null(current_tag)) {
      result_list[[current_tag]] <- c(result_list[[current_tag]], paste(current_content, collapse = "; "))
      current_tag <- NULL
      current_content <- NULL
    }
  }
}
# 保存最后一组未处理的内容
if (!is.null(current_tag)) {
  result_list[[current_tag]] <- c(result_list[[current_tag]], paste(current_content, collapse = "; "))
}
close(con)

# 转换为DataFrame,补全长度不一致的列
max_len <- max(sapply(result_list, length))
df <- data.frame(
  author_col = if (length(result_list$AU) < max_len) c(result_list$AU, rep(NA, max_len - length(result_list$AU))) else result_list$AU,
  title_col = if (length(result_list$TI) < max_len) c(result_list$TI, rep(NA, max_len - length(result_list$TI))) else result_list$TI,
  date_col = if (length(result_list$PD) < max_len) c(result_list$PD, rep(NA, max_len - length(result_list$PD))) else result_list$PD,
  summary_col = if (length(result_list$AB) < max_len) c(result_list$AB, rep(NA, max_len - length(result_list$AB))) else result_list$AB,
  stringsAsFactors = FALSE
)

方法2:tidyverse工具(代码简洁)

若已安装tidyverse包,可通过更简洁的代码实现:

library(tidyverse)

# 读取所有行并预处理
lines <- read_lines("groupprojectdatabase.txt") %>% 
  trimws() %>% 
  .[nchar(.) > 0]

# 转换为DataFrame
df <- tibble(line = lines) %>%
  # 标记标签行,为续行填充对应标签
  mutate(tag = ifelse(str_detect(line, "^[A-Z]{2} "), substr(line, 1, 2), NA)) %>%
  fill(tag, .direction = "down") %>%
  # 筛选目标标签,提取内容
  filter(tag %in% c("AU", "TI", "PD", "AB")) %>%
  mutate(content = trimws(str_remove(line, "^[A-Z]{2} "))) %>%
  # 合并同标签的多行内容
  group_by(tag) %>%
  summarise(content = str_c(content, collapse = "; ")) %>%
  # 转换为宽表并重命名列
  pivot_wider(names_from = tag, values_from = content) %>%
  rename(author_col = AU, title_col = TI, date_col = PD, summary_col = AB)

关键提示

  • 两种方法都处理了同标签续行问题,用分号分隔合并内容
  • 基础R方法更适合10万+行的超大型文件,逐行读取不会占用过多内存
  • tidyverse方法代码更易读,但需要加载相关包

内容的提问来源于stack exchange,提问作者Flynn Howl

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 14:01:16