You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Officer和Docxtractr读取Docx文件时遇Path_To_Connection错误

解决R语言读取法律数据库docx文件时的xml2 path_to_connection错误

问题背景

有数百个来自法律数据库的.docx文件需要做NLP处理,但用R语言的officer和docxtractr包读取时,一直触发xml2的path_to_connection错误,因版权限制无法提供测试文档,导致工作停滞。手动创建的测试docx能正常读取,但数据库下载的文件执行相同步骤就失败。

已做排查

  • 修改文件名(改为test.docx)后问题依旧
  • 分别在Windows办公电脑和Mac家用电脑测试,错误完全一致
  • 追踪底层函数发现,临时解压路径中存在不存在的子目录,其中word子目录在两个系统中均缺失

复现代码(Mac环境)

library(tidyverse)
library(officer)
#> Warning: package 'officer' was built under R version 4.4.1

test <- officer::read_docx("~/Downloads/test.docx")
#> Warning in read_core_properties(package_dir): No properties found. Using
#> properties from template file in officer.
#> Error in `path_to_connection()`:
#> ! 
#>   '/var/folders/q2/8qcfd1qd5rqfvhr2x9l2k24m0000gp/T//RtmpZ6jZ5U/file95017ae662ea/word/document.xml'
#>   does not exist.
reprex::reprex()
#> ℹ Non-interactive session, setting `html_preview = FALSE`.
#> CLIPR_ALLOW has not been set, so clipr will not run interactively
#> Error in switch(where, expr = stringify_expression(x_expr), clipboard = ingest_clipboard(), : EXPR must be a length 1 vector

错误回溯

11. signal_abort(cnd, .file)
10. rlang::abort(message, ..., call = call, use_cli_format = TRUE,.frame = .frame)
9. cli::cli_abort(msg, call = call)
8. check_path(path)
7. path_to_connection(x)
6. read_xml.character(file)
5. read_xml(file)
4. super$feed(file.path(private$package_dir, "word", main_file))
3. initialize(...)
2. docx_part$new(package_dir, main_file = "document.xml", cursor = "/w:document/w:body/*[1]",body_xpath = "/w:document/w:body")
1. officer::read_docx("~/Downloads/test.docx")

解决方案建议

1. 先确认docx文件的实际结构

docx本质是压缩包,手动解压任意一个问题文件,检查里面是否存在word/document.xml:

  • 如果原文件本身就没有这个结构,说明文件可能是伪装成docx的其他格式,或者数据库导出时损坏
  • 如果有这个结构,说明officer的解压/读取逻辑对这类文件不兼容

2. 尝试其他R包读取

  • readtext包:底层依赖python的python-docx,兼容性较好
    library(readtext)
    doc_content <- readtext("~/Downloads/test.docx")
    # 提取文本内容
    doc_text <- doc_content$text
    
  • textreadr包:专门处理文本提取,支持多种格式
    library(textreadr)
    doc_text <- read_docx("~/Downloads/test.docx")
    

3. 手动解压后读取xml

如果上述包仍失败,可以手动解压docx后直接读取xml:

# 设置临时解压目录
temp_dir <- tempdir()
# 解压docx文件
unzip("~/Downloads/test.docx", exdir = temp_dir)
# 检查目标文件是否存在
if (file.exists(file.path(temp_dir, "word", "document.xml"))) {
  # 读取xml内容
  doc_xml <- xml2::read_xml(file.path(temp_dir, "word", "document.xml"))
  # 提取所有文本节点内容
  doc_text <- xml2::xml_text(xml2::xml_find_all(doc_xml, "//w:t"))
  # 合并文本
  full_text <- paste(doc_text, collapse = " ")
} else {
  warning("文件结构异常,缺失word/document.xml")
}

4. 修复损坏的docx文件

如果文件本身结构损坏,可尝试:

  • 用Microsoft Word打开文件,直接重新保存为docx格式
  • 用LibreOffice命令行批量转换修复(需安装LibreOffice):
    libreoffice --headless --convert-to docx ~/Downloads/*.docx
    

5. 检查数据库导出设置

确认法律数据库导出docx时,是否使用了非标准格式、加密压缩,或者附加了特殊元数据,尝试调整导出参数(如选择“标准docx格式”)重新导出文件。

内容的提问来源于stack exchange,提问作者wdefreit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 20:10:57