使用docx4j插入段落导致docx文件损坏,请求协助排查
问题:docx4j添加段落生成的Word文档损坏
我尝试使用docx4j在一个仅包含一行测试文本的现有docx文件中添加一段内容为“Hello from Java test!”的文本段落。但添加该段落后生成的docx文件在使用MS Word打开时显示“损坏”,Word提示:
Word found unreadable content in generated.docx. Do you want to recover the contents of this document? If you trust the source of this document, click Yes.
移除mainDocumentPart.addParagraphOfText("Hello from Java test!");这一行后,生成的docx文件可正常用Word打开。
使用的版本信息
- Word版本:Microsoft 365 Apps for enterprise, Version 2505 (Build 18827, 20202 Click-to-Run)
- Docx4j版本:8.3.3
- Java版本:1.8
相关代码
public ByteBuffer addingDummyParagraphOnDocxFile(ByteBuffer original) throws Exception { byte[] originalData = original.array(); WordprocessingMLPackage wordPackage = WordprocessingMLPackage.load(new ByteArrayInputStream(originalData)); MainDocumentPart mainDocumentPart = wordPackage.getMainDocumentPart(); // Adding a simple paragraph of text to test mainDocumentPart.addParagraphOfText("Hello from Java test!"); ByteArrayOutputStream out = new ByteArrayOutputStream(); wordPackage.save(out); //================================ SAVE LOCALLY!!! ========================================== // DEBUG: Save to local temp folder for inspection try { File tempDir = new File("C:\\temp"); if (!tempDir.exists()) { boolean created = tempDir.mkdirs(); if (!created) { log.warn("Could not create temp directory: {}", tempDir.getAbsolutePath()); } } File tempFile = new File(tempDir, "doc-" + System.currentTimeMillis() + ".docx"); try (FileOutputStream fos = new FileOutputStream(tempFile)) { fos.write(out.toByteArray()); log.info("Saved document for inspection at: {}", tempFile.getAbsolutePath()); } } catch (IOException e) { log.error("Failed to save temporary document", e); } //========================================================================== return ByteBuffer.wrap(out.toByteArray()); }
可能的解决方法
1. 修正ByteBuffer数据读取方式
original.array()会获取ByteBuffer的整个底层数组,可能包含未使用的冗余字节,导致加载的文档数据异常。改为读取实际有效数据:
// 读取ByteBuffer中实际有效的字节 byte[] originalData = new byte[original.remaining()]; original.get(originalData); // 重置ByteBuffer的位置,避免后续操作出错 original.position(original.position() - originalData.length); WordprocessingMLPackage wordPackage = WordprocessingMLPackage.load(new ByteArrayInputStream(originalData));
2. 手动构建段落结构
addParagraphOfText生成的段落可能缺少Word要求的必要元素,比如样式引用。手动构建完整的段落结构:
// 创建段落对象 P paragraph = new P(); // 创建文本运行对象 R run = new R(); // 创建文本内容 Text text = new Text(); text.setValue("Hello from Java test!"); // 将文本添加到运行 run.getContent().add(text); // 将运行添加到段落 paragraph.getContent().add(run); // 将段落添加到文档主内容 mainDocumentPart.getContent().add(paragraph);
3. 确保输出流完整写入
保存文档后刷新输出流,确保所有数据都被写入:
ByteArrayOutputStream out = new ByteArrayOutputStream(); wordPackage.save(out); out.flush(); // 强制刷新输出流
4. 验证原始文档兼容性
将原始docx文档用Word另存为标准的“Word Document (*.docx)”格式,排除原始文档本身的格式异常后再测试。
内容的提问来源于stack exchange,提问作者ipapag
相关产品推荐
相关产品推荐

