You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

docx4j替换占位符时WordML标记显示异常求助

问题

使用docx4j和docx4j-search-and-replace-util工具替换纯文本占位符(如${NAME})正常,但替换带格式HTML生成的WordML到${APP_ADDITIONAL_INFO}时,生成的docx文档直接显示WordML标记而非格式化文本。尝试过两种方式均失败:

  • 用JTidy转HTML为XHTML,再通过docx4j-ImportXHTML转成WordML字符串放入替换映射;
  • 直接使用WordML字符串替换。

相关代码如下:

WordprocessingMLPackage wordMLPackage = WordprocessingMLPackage.createPackage();
XHTMLImporterImpl XHTMLImporter = new XHTMLImporterImpl(wordMLPackage);

BufferedReader br = new BufferedReader(new StringReader(doc.getDescription()));

StringWriter sw = new StringWriter();
Tidy t = new Tidy();
t.setDropEmptyParas(true);
t.setShowWarnings(false); //to hide errors
t.setQuiet(true); //to hide warning
t.setUpperCaseAttrs(false);
t.setXmlOut(true);
t.setUpperCaseTags(false);
t.setInputEncoding("UTF-8");
t.setOutputEncoding("UTF-8");
t.setXmlOut(true);
t.parse(br,sw);
StringBuffer sb = sw.getBuffer();
String strClean = sb.toString();
br.close();
sw.close();

wordMLPackage.getMainDocumentPart().getContent().addAll(XHTMLImporter.convert( strClean, null) );
// the variable that should contain WordML markup
String description = XmlUtils.marshaltoString(wordMLPackage.getMainDocumentPart().getJaxbElement(), true, true);

// map with placeholders and replacing data
Map<String, String> replaceMap = new HashMap<String, String>() {{
    put("${APP_EMPLOYEE}",              doc.getEmployeeName());
    put("${APP_JOB_TITLE}",             doc.getJobtitle());
    put("${APP_ADDITIONAL_INFO}",       description);
}};

byte[] cos = gt.generateDocXDocument(filePath, replaceMap, masterId);

return ResponseEntity.ok()
        .header(HttpHeaders.CONTENT_DISPOSITION, "attachment; filename=\"test.docx\"").body(cos);

generateDocXDocument方法代码:

generateDocXDocument(String filePath, Map <String, String> replaceMap){
  byte[] decryptedBytesOfFile = storageService.loadFile(filePath);
  WordprocessingMLPackage wordMLPackage = WordprocessingMLPackage.load(new ByteArrayInputStream(decryptedBytesOfFile));
  Docx4JSRUtil.searchAndReplace(wordMLPackage, replaceMap);
  OutputStream outputStream = new ByteArrayOutputStream();
  Save saver = new Save(wordMLPackage);
  saver.save(outputStream);
  return ((ByteArrayOutputStream) outputStream).toByteArray();
}

直接替换WordML的尝试代码:

put("${APP_ADDITIONAL_INFO}", "<w:r><w:rPr><w:rFonts w:ascii=\"Times New Roman\" w:hAnsi=\"Times New Roman\"/><w:b/><w:i w:val=\"false\"/><w:color w:val=\"000000\"/><w:sz w:val=\"22\"/></w:rPr><w:t>Formatted text</w:t></w:r>");
解决方案

问题根源

docx4j-search-and-replace-util的searchAndReplace是纯文本替换工具,它会把传入的WordML字符串当作普通文本处理,转义后直接插入文档,所以最终显示的是原始标签而非解析后的格式内容。

正确处理步骤

要实现带格式内容替换,不能用纯文本替换逻辑,需要定位占位符位置后,用docx4j API直接插入格式化的WordML内容:

  1. 保留HTML转换后的原生WordML对象
    不需要把转换结果序列化为字符串,直接保留XHTMLImporter.convert返回的List<Object>(docx4j内部的JAXB元素集合):

    // 直接用docx4j处理Quill生成的HTML,无需JTidy转换
    XHTMLImporterImpl xhtmlImporter = new XHTMLImporterImpl(templatePackage);
    List<Object> formattedContent = xhtmlImporter.convert(doc.getDescription(), null);
    
  2. 遍历文档替换占位符为格式化内容
    放弃Docx4JSRUtil.searchAndReplace,直接遍历文档结构找到占位符所在的文本节点,替换为格式化内容:

    // 加载模板文档
    WordprocessingMLPackage templatePackage = WordprocessingMLPackage.load(new ByteArrayInputStream(decryptedBytesOfFile));
    MainDocumentPart mainPart = templatePackage.getMainDocumentPart();
    List<Object> content = mainPart.getContent();
    
    // 遍历段落查找占位符
    for (int i = 0; i < content.size(); i++) {
        Object obj = content.get(i);
        if (!(obj instanceof P)) continue;
        P paragraph = (P) obj;
    
        // 遍历段落内的文本块
        for (int j = 0; j < paragraph.getContent().size(); j++) {
            Object pObj = paragraph.getContent().get(j);
            if (!(pObj instanceof R)) continue;
            R run = (R) pObj;
    
            // 遍历文本节点
            for (int k = 0; k < run.getContent().size(); k++) {
                Object rObj = run.getContent().get(k);
                if (!(rObj instanceof Text)) continue;
                Text text = (Text) rObj;
    
                if ("${APP_ADDITIONAL_INFO}".equals(text.getValue())) {
                    // 移除原占位符段落
                    content.remove(i);
                    // 插入格式化内容
                    content.addAll(i, formattedContent);
                    // 跳出所有循环避免重复处理
                    i = content.size();
                    j = paragraph.getContent().size();
                    k = run.getContent().size();
                }
            }
        }
    }
    
  3. 纯文本占位符单独处理
    对于${APP_EMPLOYEE}这类纯文本占位符,可以继续用Docx4JSRUtil.searchAndReplace,但要注意顺序:先处理格式化内容替换,再处理纯文本替换,避免占位符被提前覆盖。

额外注意事项

  • Quill生成的HTML基本符合XHTML规范,无需JTidy额外转换;
  • 若格式化内容包含图片等资源,需为XHTMLImporter配置资源处理逻辑;
  • 替换时需保证插入的WordML元素(段落、文本块等)符合docx的结构规范。

内容的提问来源于stack exchange,提问作者MichaelSun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 19:14:53