如何修复不一致DOCX文件,实现内存中DOCX转PDF?
DOCX合并与内存转PDF的格式修复问题
我正面临DOCX转PDF的技术难题:从数据库读取DOCX文件的字节数据,使用Apache POI合并后,尝试通过Docx4j在内存中转换为PDF(生产环境无法使用本地存储,仅测试阶段可用)。问题在于数据库中的DOCX文件格式不一致,可能缺失部分元数据或属性。
合并DOCX的实现代码
public XWPFDocument mergeDocx(List<String> docxNames) throws Exception { List<FileData> fileData = repository.getDocxs(docxNames); ZipSecureFile.setMinInflateRatio(0); InputStream inputS = new ByteArrayInputStream(fileData.get(0).getData()); OPCPackage opcPackage = OPCPackage.open(inputS); XWPFDocument xwpfDocument = new XWPFDocument(opcPackage); fileData.remove(0); if (!fileData.isEmpty()) { for (FileData fd : fileData) { inputS = new ByteArrayInputStream(fd.getData()); opcPackage = OPCPackage.open(inputS); XWPFDocument xwpf = new XWPFDocument(opcPackage); CTBody bodyToAppend = xwpf.getDocument().getBody(); xwpfDocument.getDocument().addNewBody().set(bodyToAppend); } } inputS.close(); opcPackage.close(); return xwpfDocument; }
转PDF的实现代码
public void toPdf(XWPFDocument docxDocument) throws Exception { //in ByteArrayOutputStream baos = new ByteArrayOutputStream(); docxDocument.write(baos); docxDocument.close(); byte[] bytes = baos.toByteArray(); //this is basically a ByteArray of an XML file, not a consistent DOCX one WordprocessingMLPackage ml = Docx4J.load(new ByteArrayInputStream(bytes)); //out OutputStream output = new FileOutputStream("/Users/Santiago/Documents/test.pdf"); Docx4J.toPDF(ml, output); output.flush(); output.close(); }
目前的问题是:最终合并后的DOCX文件及数据库中读取的原始文件均存在损坏,无法在转PDF方法中正常运行;仅当将合并后的文件保存为本地文件后,通过DOCX转换器处理才能正常工作。
现咨询:是否可在进入转PDF方法前,通过添加属性或应用格式等方式修复DOCX文件使其格式一致?且不依赖外部工具(如当前使用的在线转换工具)
内容的提问来源于stack exchange,提问作者Santiago Cometto
相关产品推荐
相关产品推荐

