You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Mammoth与html-docx-js实现Docx-Html互转时遇逆向转换异常

问题:Mammoth无法转换html-docx-js生成的Docx文件

我在项目中实现了这样的流程:使用Mammoth将Docx转Html,经Angular编辑器编辑后,用html-docx-js把Html转回Docx。但用这个流程生成的Docx文件,再次用Mammoth转换为Html时,无法输出有效内容,仅返回异常;而微软Office生成的Docx则能正常转换。

以下是相关Node.js API代码:

Docx转Html代码

exports.convertDocxToHtml = catchAsync(async (req, res, next) => {
  let url = 'https://teamee-drive-bucket.s3.amazonaws.com/drive/untitle.docx';
  const response = await axios.get(url, { responseType: 'arraybuffer' })
  const result = await mammoth.convertToHtml({ buffer: Buffer.from(response.data, "utf-8") })

  // let url = 'src/controllers/drive/default.docx';
  // // let url = 'src/controllers/drive/Imran4.docx';
  // const result = await mammoth.convertToHtml({ path: url })
  console.log(result)

  res.status(200).json({
    status: 'success',
    data: result.value,
    message: "Data loaded success!"
  });
});

Html转Docx代码

exports.saveHtmlToDocx = catchAsync(async (req, res, next) => {
  let body = req.body
  // console.log(body)
  const htmlString = `<!DOCTYPE html>
  <html>
  <head>
  <title></title>
  </head>
  <body>${body.html}</body>
  </html>`
  const opt = {
    margin: {
      top: 100
    },
    orientation: 'landscape'
  }
  var converted = htmlDocx.asBlob(htmlString, opt);
  console.log(converted)

  const s3 = new AWS.S3()
  const param = {
    Bucket: process.env.AWS_BUCKET_NAME + '/drive',
    Key: `${body.document_title}.docx`,
    Body: converted
  }
  let data =  await s3.upload(param).promise()

  const update = await DriveFile.updateOne(
    {_id:body.id},
    {$set:{
      url:data.Location,
      folderId:body.folderId,
      fileName:`${body.document_title}.docx`
    }}
  )
  if (!update) {
      return next(new AppError("Document update faile", 400))
  }
  
  res.status(200).json({
    status: 'success',
    data: update,
    message: "Data save success!"
  });
});

核心原因

  1. 文档结构差异:html-docx-js基于HTML生成Docx,其内部的XML结构、样式标记逻辑和微软Office原生生成的标准Docx不完全匹配,Mammoth的解析逻辑更偏向兼容原生Office格式,对非标准结构的支持有限。
  2. 二进制数据处理错误:当前Docx转Html代码中,用Buffer.from(response.data, "utf-8")处理arraybuffer是错误的,会破坏二进制文件结构。

解决方案

1. 修复二进制数据处理逻辑

直接使用axios返回的arraybuffer作为Mammoth的输入,不需要额外编码转换:

exports.convertDocxToHtml = catchAsync(async (req, res, next) => {
  let url = 'https://teamee-drive-bucket.s3.amazonaws.com/drive/untitle.docx';
  const response = await axios.get(url, { responseType: 'arraybuffer' })
  // 直接传入response.data,无需utf-8转换
  const result = await mammoth.convertToHtml({ buffer: response.data })

  console.log(result)

  res.status(200).json({
    status: 'success',
    data: result.value,
    message: "Data loaded success!"
  });
});

2. 优化html-docx-js的生成兼容性

  • 尽量使用简单的HTML标签(如<p>、<h1>-<h6>、<ul>/<ol>),避免Angular编辑器生成的冗余自定义标签或复杂样式。
  • 添加贴近原生Office的配置,提升生成Docx的标准性:
const opt = {
  margin: { top: 100, right: 50, bottom: 100, left: 50 },
  orientation: 'landscape',
  // 配置Office默认字体,减少结构差异
  font: {
    name: 'Calibri',
    size: 11
  }
}

3. 捕获Mammoth异常定位问题

添加异常捕获,查看具体错误信息,精准定位问题:

exports.convertDocxToHtml = catchAsync(async (req, res, next) => {
  let url = 'https://teamee-drive-bucket.s3.amazonaws.com/drive/untitle.docx';
  try {
    const response = await axios.get(url, { responseType: 'arraybuffer' })
    const result = await mammoth.convertToHtml({ buffer: response.data })

    res.status(200).json({
      status: 'success',
      data: result.value,
      message: "Data loaded success!"
    });
  } catch (error) {
    console.error("Mammoth转换错误详情:", error)
    return next(new AppError(`Docx转换失败: ${error.message}`, 500))
  }
});

4. 备选方案:替换生成库

如果上述优化仍无法解决,可更换为更贴近原生Docx标准的生成库:

  • docx:通过API直接构建标准Docx文档,兼容性拉满
  • pandoc:命令行工具,支持HTML到Docx的高质量转换,可通过Node.js调用

内容的提问来源于stack exchange,提问作者Al Imran

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 21:27:12