如何使用DOCX.js库将HTML标签转换为WordDoc文本格式?
问题原因
docx.js原生不支持直接解析渲染HTML字符串,直接将HTML内容作为文本传入时,所有标签会被默认当作普通字符串处理,自然不会渲染对应样式。
解决方案
你需要先将HTML解析为DOM节点,再将节点映射为docx.js支持的内置元素,具体实现逻辑如下:
- 浏览器环境直接用原生DOM API解析HTML字符串,Node.js环境可借助jsdom完成解析
- 按标签类型映射到docx.js对应元素:
<p>标签对应Paragraph实例<strong>标签对应设置bold: true属性的TextRun实例- 普通文本对应默认属性的
TextRun实例
- 所有转换完成的元素传入
Document实例后再导出Word文件即可
示例代码
import { Document, Paragraph, TextRun, Packer } from 'docx'; import { saveAs } from 'file-saver'; // 待转换的HTML字符串 const htmlStr = `<p> <strong>voice:</strong> During an address to Congress, George Washington said “If you can’t send money, send cigars.” [pause] Or something like that. [pause] There’s no question that Washington was a wise man. He knew that premium quality cigars are worth their weight in gold. And if he were alive today, he’d be running the country from the lounge at <strong>No Name Cigar Company.</strong> Where the conversation never comes to an end – and neither does their premier cigar selection. Mingle with legends – at <strong>No Name Cigar Company.</strong> </p> <p> </p> <p>Visit them online at <strong>NoNameWebsite.com</strong></p>`; // 解析HTML const tempDiv = document.createElement('div'); tempDiv.innerHTML = htmlStr; const childNodes = tempDiv.childNodes; const paragraphs = []; // 遍历p标签 childNodes.forEach(node => { if (node.nodeName === 'P') { const textRuns = []; // 遍历p标签内的子节点 node.childNodes.forEach(child => { if (child.nodeName === 'STRONG') { textRuns.push(new TextRun({ text: child.textContent, bold: true })); } else if (child.nodeType === Node.TEXT_NODE) { textRuns.push(new TextRun(child.textContent.replace(/\u00A0/g, ' '))); } }); paragraphs.push(new Paragraph({ children: textRuns })); } }); // 生成文档 const doc = new Document({ sections: [{ properties: {}, children: paragraphs }] }); // 导出文件 Packer.toBlob(doc).then(blob => { saveAs(blob, "输出文件.docx"); });
如果需要处理更多HTML标签(比如链接、斜体、列表等),只需要在转换逻辑中增加对应标签的映射规则即可。
内容的提问来源于stack exchange,提问作者Muhammad Umar
相关产品推荐
相关产品推荐

