You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python将HTML文件内容替换到docx模板指定位置并保存?

如何将带格式的HTML内容插入到docx模板指定位置?

问题背景

需要把data_1.html中包含粗体、斜体、下划线等格式的内容,插入到template.docx模板的指定占位符位置,最终导出为output.docx。直接将HTML文件对象传入docxtpl上下文替换时,生成的文档无法被Word正常打开。

解决思路

docxtpl无法直接解析HTML格式内容,也不能直接处理文件对象,因此需要先将HTML转换成docx原生支持的富文本结构,再结合模板完成替换。我们可以用htmldocx处理HTML到docx格式的转换,再通过python-docx的API把转换后的富文本内容插入到模板的占位符位置。

完整实现代码

from htmldocx import HtmlToDocx
from docxtpl import DocxTemplate
from docx import Document

# 1. 把HTML内容转换成带格式的docx文档对象
html_content = open("data_1.html", "r", encoding="utf-8").read()
temp_doc = Document()
html_parser = HtmlToDocx()
html_parser.add_html_to_document(html_content, temp_doc)

# 2. 加载目标模板
template_doc = DocxTemplate("template.docx")

# 3. 自定义替换逻辑:将HTML转换后的富文本插入到占位符位置
def insert_formatted_content(template, placeholder, content_doc):
    # 处理普通段落中的占位符
    for para in template.paragraphs:
        if placeholder in para.text:
            # 清空原段落的内容
            para.clear()
            # 把HTML转换后的每一段内容带格式复制过来
            for content_para in content_doc.paragraphs:
                for run in content_para.runs:
                    new_run = para.add_run(run.text)
                    # 复制格式属性
                    new_run.bold = run.bold
                    new_run.italic = run.italic
                    new_run.underline = run.underline
                    new_run.font.size = run.font.size
                    new_run.font.name = run.font.name
            break  # 若占位符仅出现一次,处理后退出循环

    # 处理表格单元格中的占位符(按需启用)
    for table in template.tables:
        for row in table.rows:
            for cell in row.cells:
                if placeholder in cell.text:
                    cell.clear()
                    for content_para in content_doc.paragraphs:
                        cell_para = cell.add_paragraph()
                        for run in content_para.runs:
                            new_run = cell_para.add_run(run.text)
                            new_run.bold = run.bold
                            new_run.italic = run.italic
                            new_run.underline = run.underline
                            new_run.font.size = run.font.size
                            new_run.font.name = run.font.name

# 执行替换,注意占位符要和模板中的一致,比如模板里是{{ MONTAG }}就传这个字符串
insert_formatted_content(template_doc, "{{ MONTAG }}", temp_doc)

# 保存最终文档
template_doc.save("output.docx")

错误原因分析

之前直接传入文件对象到docxtpl的上下文,docxtpl会将文件对象的字符串表示(如<_io.TextIOWrapper name='data_1.html' mode='r' encoding='UTF-8'>)写入文档,这种内容不符合docx的格式规范,导致Word无法解析,最终出现文件损坏无法打开的问题。


内容的提问来源于stack exchange,提问作者marlon_python

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 18:51:22