使用Mammoth与html-docx-js实现Docx-Html互转时遇逆向转换异常
问题:Mammoth无法转换html-docx-js生成的Docx文件
我在项目中实现了这样的流程:使用Mammoth将Docx转Html,经Angular编辑器编辑后,用html-docx-js把Html转回Docx。但用这个流程生成的Docx文件,再次用Mammoth转换为Html时,无法输出有效内容,仅返回异常;而微软Office生成的Docx则能正常转换。
以下是相关Node.js API代码:
Docx转Html代码
exports.convertDocxToHtml = catchAsync(async (req, res, next) => { let url = 'https://teamee-drive-bucket.s3.amazonaws.com/drive/untitle.docx'; const response = await axios.get(url, { responseType: 'arraybuffer' }) const result = await mammoth.convertToHtml({ buffer: Buffer.from(response.data, "utf-8") }) // let url = 'src/controllers/drive/default.docx'; // // let url = 'src/controllers/drive/Imran4.docx'; // const result = await mammoth.convertToHtml({ path: url }) console.log(result) res.status(200).json({ status: 'success', data: result.value, message: "Data loaded success!" }); });
Html转Docx代码
exports.saveHtmlToDocx = catchAsync(async (req, res, next) => { let body = req.body // console.log(body) const htmlString = `<!DOCTYPE html> <html> <head> <title></title> </head> <body>${body.html}</body> </html>` const opt = { margin: { top: 100 }, orientation: 'landscape' } var converted = htmlDocx.asBlob(htmlString, opt); console.log(converted) const s3 = new AWS.S3() const param = { Bucket: process.env.AWS_BUCKET_NAME + '/drive', Key: `${body.document_title}.docx`, Body: converted } let data = await s3.upload(param).promise() const update = await DriveFile.updateOne( {_id:body.id}, {$set:{ url:data.Location, folderId:body.folderId, fileName:`${body.document_title}.docx` }} ) if (!update) { return next(new AppError("Document update faile", 400)) } res.status(200).json({ status: 'success', data: update, message: "Data save success!" }); });
核心原因
- 文档结构差异:html-docx-js基于HTML生成Docx,其内部的XML结构、样式标记逻辑和微软Office原生生成的标准Docx不完全匹配,Mammoth的解析逻辑更偏向兼容原生Office格式,对非标准结构的支持有限。
- 二进制数据处理错误:当前Docx转Html代码中,用
Buffer.from(response.data, "utf-8")处理arraybuffer是错误的,会破坏二进制文件结构。
解决方案
1. 修复二进制数据处理逻辑
直接使用axios返回的arraybuffer作为Mammoth的输入,不需要额外编码转换:
exports.convertDocxToHtml = catchAsync(async (req, res, next) => { let url = 'https://teamee-drive-bucket.s3.amazonaws.com/drive/untitle.docx'; const response = await axios.get(url, { responseType: 'arraybuffer' }) // 直接传入response.data,无需utf-8转换 const result = await mammoth.convertToHtml({ buffer: response.data }) console.log(result) res.status(200).json({ status: 'success', data: result.value, message: "Data loaded success!" }); });
2. 优化html-docx-js的生成兼容性
- 尽量使用简单的HTML标签(如
<p>、<h1>-<h6>、<ul>/<ol>),避免Angular编辑器生成的冗余自定义标签或复杂样式。 - 添加贴近原生Office的配置,提升生成Docx的标准性:
const opt = { margin: { top: 100, right: 50, bottom: 100, left: 50 }, orientation: 'landscape', // 配置Office默认字体,减少结构差异 font: { name: 'Calibri', size: 11 } }
3. 捕获Mammoth异常定位问题
添加异常捕获,查看具体错误信息,精准定位问题:
exports.convertDocxToHtml = catchAsync(async (req, res, next) => { let url = 'https://teamee-drive-bucket.s3.amazonaws.com/drive/untitle.docx'; try { const response = await axios.get(url, { responseType: 'arraybuffer' }) const result = await mammoth.convertToHtml({ buffer: response.data }) res.status(200).json({ status: 'success', data: result.value, message: "Data loaded success!" }); } catch (error) { console.error("Mammoth转换错误详情:", error) return next(new AppError(`Docx转换失败: ${error.message}`, 500)) } });
4. 备选方案:替换生成库
如果上述优化仍无法解决,可更换为更贴近原生Docx标准的生成库:
docx:通过API直接构建标准Docx文档,兼容性拉满pandoc:命令行工具,支持HTML到Docx的高质量转换,可通过Node.js调用
内容的提问来源于stack exchange,提问作者Al Imran
相关产品推荐
相关产品推荐

