You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中从半结构化文本提取信息生成指定DataFrame?

无分隔符文本转R DataFrame的处理方案

核心思路

利用每个条目以空行结束的特点,先按空行分割所有条目,再针对两种条目结构分别用正则表达式提取目标字段。


步骤1:读取文本并分割条目

先把整个文本按行读取,再根据空行位置分割成单个条目(每个条目是多行内容的集合):

# 读取目标文本文件
text_lines <- readLines("your_file.txt")
# 定位所有空行的位置,作为条目分割标记
empty_line_pos <- which(text_lines == "")
# 按空行分割成独立条目
entries <- split(text_lines, cumsum(c(TRUE, diff(empty_line_pos) != 1)))
# 过滤掉空的无效条目
entries <- entries[sapply(entries, function(x) length(x) > 0)]

步骤2:编写条目解析函数

针对两种条目结构,分别提取标题、数量、位置、日期、时间、名称字段:

parse_single_entry <- function(entry_lines) {
  # 初始化结果容器
  entry_data <- list(标题 = NA, 数量 = NA, 位置 = NA, 日期 = NA, 时间 = NA, 名称 = NA)
  first_line <- entry_lines[1]
  
  # 处理书籍类条目
  if (!grepl("^Shelf Review", first_line)) {
    # 匹配「书名 (数量) 国家」格式
    book_pattern <- regexec("^(.*?)\\s\\((\\d+)\\)\\s.*$", first_line)
    match_result <- regmatches(first_line, book_pattern)[[1]]
    if (length(match_result) >= 3) {
      entry_data$标题 <- match_result[2]
      entry_data$数量 <- as.integer(match_result[3])
    }
    # 提取后续行的日期、时间、作者信息
    for (line in entry_lines[-1]) {
      entry_data$日期 <- if (grepl("^Date:", line)) sub("^Date:\\s*", "", line) else entry_data$日期
      entry_data$时间 <- if (grepl("^Time:", line)) sub("^Time:\\s*", "", line) else entry_data$时间
      entry_data$名称 <- if (grepl("^Author:", line)) sub("^Author:\\s*", "", line) else entry_data$名称
    }
  } 
  # 处理Shelf Review类条目
  else {
    # 匹配「Shelf Review () 图书馆位置」格式
    shelf_pattern <- regexec("^Shelf Review\\s\\(\\)\\s(.*)$", first_line)
    match_result <- regmatches(first_line, shelf_pattern)[[1]]
    if (length(match_result) >= 2) {
      entry_data$标题 <- "Shelf Review"
      entry_data$位置 <- match_result[2]
    }
    # 提取后续行的日期、时间、工作人员、价格信息
    for (line in entry_lines[-1]) {
      entry_data$日期 <- if (grepl("^Date:", line)) sub("^Date:\\s*", "", line) else entry_data$日期
      entry_data$时间 <- if (grepl("^Time:", line)) sub("^Time:\\s*", "", line) else entry_data$时间
      entry_data$名称 <- if (grepl("^Staff:", line)) sub("^Staff:\\s*", "", line) else entry_data$名称
      # 提取价格/价值的数值部分
      if (grepl("^(Value|Price):", line)) {
        entry_data$数量 <- as.numeric(gsub("[^0-9.]", "", line))
      }
    }
  }
  return(as.data.frame(entry_data))
}

步骤3:批量解析并生成DataFrame

遍历所有条目执行解析,最后合并成统一的DataFrame:

# 批量处理所有条目
df_list <- lapply(entries, parse_single_entry)
# 合并为最终DataFrame
final_df <- do.call(rbind, df_list)

注意事项

  • 正则表达式需根据实际文本的字段前缀、格式变体调整,比如如果日期前缀是「日期:」而非「Date:」,要修改匹配规则;
  • 若存在字段缺失的条目,结果中会保留NA,可后续用tidyr::fill()或base::replace()处理;
  • 若条目开头格式有细微差异(比如括号内有空格),需微调正则表达式的匹配逻辑。

内容的提问来源于stack exchange,提问作者Andres

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 09:55:25