如何用Python将HTML文件内容替换到docx模板指定位置并保存?
如何将带格式的HTML内容插入到docx模板指定位置?
问题背景
需要把data_1.html中包含粗体、斜体、下划线等格式的内容,插入到template.docx模板的指定占位符位置,最终导出为output.docx。直接将HTML文件对象传入docxtpl上下文替换时,生成的文档无法被Word正常打开。
解决思路
docxtpl无法直接解析HTML格式内容,也不能直接处理文件对象,因此需要先将HTML转换成docx原生支持的富文本结构,再结合模板完成替换。我们可以用htmldocx处理HTML到docx格式的转换,再通过python-docx的API把转换后的富文本内容插入到模板的占位符位置。
完整实现代码
from htmldocx import HtmlToDocx from docxtpl import DocxTemplate from docx import Document # 1. 把HTML内容转换成带格式的docx文档对象 html_content = open("data_1.html", "r", encoding="utf-8").read() temp_doc = Document() html_parser = HtmlToDocx() html_parser.add_html_to_document(html_content, temp_doc) # 2. 加载目标模板 template_doc = DocxTemplate("template.docx") # 3. 自定义替换逻辑:将HTML转换后的富文本插入到占位符位置 def insert_formatted_content(template, placeholder, content_doc): # 处理普通段落中的占位符 for para in template.paragraphs: if placeholder in para.text: # 清空原段落的内容 para.clear() # 把HTML转换后的每一段内容带格式复制过来 for content_para in content_doc.paragraphs: for run in content_para.runs: new_run = para.add_run(run.text) # 复制格式属性 new_run.bold = run.bold new_run.italic = run.italic new_run.underline = run.underline new_run.font.size = run.font.size new_run.font.name = run.font.name break # 若占位符仅出现一次,处理后退出循环 # 处理表格单元格中的占位符(按需启用) for table in template.tables: for row in table.rows: for cell in row.cells: if placeholder in cell.text: cell.clear() for content_para in content_doc.paragraphs: cell_para = cell.add_paragraph() for run in content_para.runs: new_run = cell_para.add_run(run.text) new_run.bold = run.bold new_run.italic = run.italic new_run.underline = run.underline new_run.font.size = run.font.size new_run.font.name = run.font.name # 执行替换,注意占位符要和模板中的一致,比如模板里是{{ MONTAG }}就传这个字符串 insert_formatted_content(template_doc, "{{ MONTAG }}", temp_doc) # 保存最终文档 template_doc.save("output.docx")
错误原因分析
之前直接传入文件对象到docxtpl的上下文,docxtpl会将文件对象的字符串表示(如<_io.TextIOWrapper name='data_1.html' mode='r' encoding='UTF-8'>)写入文档,这种内容不符合docx的格式规范,导致Word无法解析,最终出现文件损坏无法打开的问题。
内容的提问来源于stack exchange,提问作者marlon_python
相关产品推荐
相关产品推荐

