使用documents4j转换xlsx/pptx为PDF提示无可用转换器如何解决
问题说明
开发Office转PDF轻量应用时,实测Documents4j处理docx转PDF效果符合预期,计划复用该组件实现xlsx、pptx格式转PDF能力。参考实现代码如下:
class OfficeToPdfConverter { fun convert(input: File, output: File) = try { LocalConverter.builder().build().convert(input).`as`(determineInput(input)).to(output) .`as`(DocumentType.PDF) .execute() } catch (e: Exception) { e.printStackTrace() } private fun determineInput(input: File) = when (input.extension) { "docx" -> DocumentType.DOCX "pptx" -> DocumentType.PPTX "xlsx" -> DocumentType.XLSX else -> DocumentType.DOCX } }
上述代码转换docx文件正常运行,转换xlsx、pptx时抛出无对应转换器的异常,异常信息如下:
- No converter for conversion of application/vnd.openxmlformats-officedocument.presentationml.presentation to application/pdf available
- No converter for conversion of application/vnd.openxmlformats-officedocument.spreadsheetml.sheet to application/pdf available
(注:原问题中两条异常均标注为PPTX对应MIME类型,属于笔误,第二条为XLSX转换的报错信息)
根本原因
Documents4j本身不内置文档解析转换能力,本质是通过JNI调用Windows系统下本地安装的Microsoft Office COM接口完成格式转换,默认仅引入Word组件的转换桥接包时,只会加载DOC/DOCX转PDF的转换链路,Excel、PowerPoint的转换能力需要额外引入对应适配依赖,同时满足环境要求才能正常运行。
解决步骤
补全依赖包
绝大多数场景下报错是因为只引入了Word转换的适配依赖,缺少Excel、PowerPoint的桥接包。以Gradle项目为例,需要补充引入以下依赖(版本号和已引入的documents4j核心包保持一致即可,目前稳定版为1.1.7):// PPT/PPTX转换适配 implementation 'com.documents4j:documents4j-transformer-msoffice-powerpoint:1.1.7' // XLS/XLSX转换适配 implementation 'com.documents4j:documents4j-transformer-msoffice-excel:1.1.7'如果是Maven项目,对应替换为Maven依赖格式即可。
校验本地运行环境
Documents4j的本地转换模式对环境有明确要求,不满足则会持续报无可用转换器错误:- 仅支持Windows系统,Linux、macOS没有Office COM组件,无法直接使用LocalConverter做本地转换,需要单独部署Windows节点部署documents4j-server,通过RemoteConverter远程调用
- 本地必须安装完整版Microsoft Office 2010及以上版本,WPS、Office精简版都不兼容;安装完成后需要手动启动一次Word、Excel、PowerPoint,完成组件初始化和正版激活,否则COM接口调用会失败
- 运行程序的账号需要有Office COM组件的访问权限,不要用SYSTEM等高权限系统账号运行服务,普通用户权限即可
优化转换代码逻辑
原代码每次转换都新建LocalConverter实例,会重复创建转换进程,浪费资源也容易触发进程锁问题,建议全局维护单例Converter实例,同时补充基础的文件校验,调整后代码参考:import com.documents4j.api.DocumentType import com.documents4j.api.IConverter import com.documents4j.job.LocalConverter import java.io.File import java.util.concurrent.TimeUnit // 全局单例初始化转换器,应用启动时创建,关闭时调用officeConverter.shutdown()释放资源 val officeConverter: IConverter = LocalConverter.builder() .workerPool(2, 5) // 根据业务并发量调整核心、最大工作线程数 .processTimeout(60, TimeUnit.SECONDS) // 大文件转换可以适当调高超时时间 .build() class OfficeToPdfConverter { fun convert(input: File, output: File) = runCatching { require(input.exists()) { "待转换文件不存在: ${input.absolutePath}" } // 目标文件已存在时先删除,避免写入失败 if (output.exists()) output.delete() officeConverter.convert(input) .`as`(resolveInputType(input)) .to(output) .`as`(DocumentType.PDF) .execute() }.onFailure { it.printStackTrace() } private fun resolveInputType(input: File): DocumentType { return when (input.extension.lowercase()) { "doc" -> DocumentType.DOC "docx" -> DocumentType.DOCX "ppt" -> DocumentType.PPT "pptx" -> DocumentType.PPTX "xls" -> DocumentType.XLS "xlsx" -> DocumentType.XLSX else -> throw IllegalArgumentException("不支持的输入文件格式: ${input.extension}") } } }
补充说明:如果是服务端批量转换场景,建议不要把转换能力和业务服务部署在同一台Linux服务器上,单独部署Windows转换节点做远程调用稳定性更高,也不会出现Office COM组件的权限兼容问题。
内容的提问来源于stack exchange,提问作者Ricardo de Vries

