如何在R中自动将10000份Word文档数据转为Excel或DataFrame?
基于R的Word文档批量患者信息提取方案
1. 安装并加载依赖包
我们将使用officer读取Word文档内容,dplyr整理输出数据,stringr处理文本匹配:
install.packages(c("officer", "dplyr", "stringr")) library(officer) library(dplyr) library(stringr)
2. 编写单文档信息提取函数
假设你的Word文档采用键值对格式(例如ID: P001、Patient Name: Jane Smith),以下函数可提取指定字段:
extract_patient_data <- function(file_path) { # 读取docx文档 doc <- read_docx(file_path) # 提取所有文本并合并为单字符串 full_text <- paste(body_extract_all_text(doc)$text, collapse = "\n") # 定义需要提取的目标字段 target_fields <- c("ID", "Patient Name", "Age", "Gender", "Diagnosis") output <- list() # 循环匹配每个字段内容 for (field in target_fields) { # 正则匹配:字段名+冒号+任意空白+内容(直到换行) match_result <- str_match(full_text, paste0("^", field, ":\\s*(.*)$")) output[[field]] <- ifelse(!is.na(match_result[1,2]), match_result[1,2], NA) } # 转换为数据框返回 return(as.data.frame(output)) }
如果文档信息是表格形式,可替换为body_extract_tbl(doc)直接提取表格数据,再筛选对应列
3. 批量处理所有文档
假设所有Word文档都存放在./patient_docs/目录下,执行以下代码批量提取并导出结果:
# 获取所有docx文件的完整路径 all_doc_paths <- list.files(path = "./patient_docs/", pattern = "\\.docx$", full.names = TRUE) # 批量应用提取函数,合并为总数据框 all_patient_data <- bind_rows(lapply(all_doc_paths, extract_patient_data)) # 导出为CSV文件(方便后续分析) write.csv(all_patient_data, "./all_patient_records.csv", row.names = FALSE)
4. 适配特殊情况的调整方案
- 若文档是
.doc格式:改用antiword包读取文本,安装后替换read_docx为antiword::antiword(file_path) - 若字段格式不统一(如大小写差异、分隔符不是冒号):修改正则表达式,例如将
paste0("^", field, ":\\s*(.*)$")改为paste0("^", str_to_lower(field), ":?\\s*(.*)$", ignore.case = TRUE) - 跳过损坏文档:添加安全处理函数避免中断批量任务
# 包装为安全执行函数 safe_extract <- purrr::safely(extract_patient_data) # 批量执行并过滤成功结果 raw_results <- lapply(all_doc_paths, safe_extract) valid_results <- purrr::keep(raw_results, ~!is.null(.x$result)) all_patient_data <- bind_rows(purrr::map(valid_results, "result"))
内容的提问来源于stack exchange,提问作者DatNguyenOph
相关产品推荐
相关产品推荐

