如何在R中从半结构化文本提取信息生成指定DataFrame?
无分隔符文本转R DataFrame的处理方案
核心思路
利用每个条目以空行结束的特点,先按空行分割所有条目,再针对两种条目结构分别用正则表达式提取目标字段。
步骤1:读取文本并分割条目
先把整个文本按行读取,再根据空行位置分割成单个条目(每个条目是多行内容的集合):
# 读取目标文本文件 text_lines <- readLines("your_file.txt") # 定位所有空行的位置,作为条目分割标记 empty_line_pos <- which(text_lines == "") # 按空行分割成独立条目 entries <- split(text_lines, cumsum(c(TRUE, diff(empty_line_pos) != 1))) # 过滤掉空的无效条目 entries <- entries[sapply(entries, function(x) length(x) > 0)]
步骤2:编写条目解析函数
针对两种条目结构,分别提取标题、数量、位置、日期、时间、名称字段:
parse_single_entry <- function(entry_lines) { # 初始化结果容器 entry_data <- list(标题 = NA, 数量 = NA, 位置 = NA, 日期 = NA, 时间 = NA, 名称 = NA) first_line <- entry_lines[1] # 处理书籍类条目 if (!grepl("^Shelf Review", first_line)) { # 匹配「书名 (数量) 国家」格式 book_pattern <- regexec("^(.*?)\\s\\((\\d+)\\)\\s.*$", first_line) match_result <- regmatches(first_line, book_pattern)[[1]] if (length(match_result) >= 3) { entry_data$标题 <- match_result[2] entry_data$数量 <- as.integer(match_result[3]) } # 提取后续行的日期、时间、作者信息 for (line in entry_lines[-1]) { entry_data$日期 <- if (grepl("^Date:", line)) sub("^Date:\\s*", "", line) else entry_data$日期 entry_data$时间 <- if (grepl("^Time:", line)) sub("^Time:\\s*", "", line) else entry_data$时间 entry_data$名称 <- if (grepl("^Author:", line)) sub("^Author:\\s*", "", line) else entry_data$名称 } } # 处理Shelf Review类条目 else { # 匹配「Shelf Review () 图书馆位置」格式 shelf_pattern <- regexec("^Shelf Review\\s\\(\\)\\s(.*)$", first_line) match_result <- regmatches(first_line, shelf_pattern)[[1]] if (length(match_result) >= 2) { entry_data$标题 <- "Shelf Review" entry_data$位置 <- match_result[2] } # 提取后续行的日期、时间、工作人员、价格信息 for (line in entry_lines[-1]) { entry_data$日期 <- if (grepl("^Date:", line)) sub("^Date:\\s*", "", line) else entry_data$日期 entry_data$时间 <- if (grepl("^Time:", line)) sub("^Time:\\s*", "", line) else entry_data$时间 entry_data$名称 <- if (grepl("^Staff:", line)) sub("^Staff:\\s*", "", line) else entry_data$名称 # 提取价格/价值的数值部分 if (grepl("^(Value|Price):", line)) { entry_data$数量 <- as.numeric(gsub("[^0-9.]", "", line)) } } } return(as.data.frame(entry_data)) }
步骤3:批量解析并生成DataFrame
遍历所有条目执行解析,最后合并成统一的DataFrame:
# 批量处理所有条目 df_list <- lapply(entries, parse_single_entry) # 合并为最终DataFrame final_df <- do.call(rbind, df_list)
注意事项
- 正则表达式需根据实际文本的字段前缀、格式变体调整,比如如果日期前缀是「日期:」而非「Date:」,要修改匹配规则;
- 若存在字段缺失的条目,结果中会保留
NA,可后续用tidyr::fill()或base::replace()处理; - 若条目开头格式有细微差异(比如括号内有空格),需微调正则表达式的匹配逻辑。
内容的提问来源于stack exchange,提问作者Andres
相关产品推荐
相关产品推荐

