如何使用python-docx替换Word文档页眉中的占位符内容?
使用python-docx替换Word文档页眉中的占位符
问题概述
需要用Python动态替换Word文档页眉中的<<company_name>>、<<company_number>>等占位符,替换值存储在JSON文件中,原代码未能成功替换页眉内容。
解决方案
以下是完整的实现代码,包含正文和页眉的占位符替换,同时解决了占位符可能被拆分到多个文本段(run)的问题:
from docx import Document import json def replace_placeholders(doc, data): # 替换正文内容 for paragraph in doc.paragraphs: full_text = paragraph.text for key, value in data.items(): full_text = full_text.replace(key, str(value)) paragraph.text = full_text # 替换页眉内容 for section in doc.sections: header = section.header for paragraph in header.paragraphs: full_text = paragraph.text for key, value in data.items(): full_text = full_text.replace(key, str(value)) paragraph.text = full_text # 加载Word文档 doc_path = "template2 (1).docx" doc = Document(doc_path) # 加载JSON中的替换值 json_file_path = "new.json" with open(json_file_path, 'r') as json_file: data = json.load(json_file) # 执行替换 replace_placeholders(doc, data) # 保存修改后的文档 modified_doc_path = "modified_document.docx" doc.save(modified_doc_path) print(f"文档 '{doc_path}' 已更新并保存为 '{modified_doc_path}'。")
关键说明
- 处理多Run问题:原代码遍历单个run进行替换,但Word可能会将一个占位符拆分成多个run(比如
<<、company_name、>>分别在不同run中),导致单个run内找不到完整占位符。直接操作段落的text属性可以避免这个问题,因为它获取的是段落的完整文本。 - 遍历所有节:Word文档的每个节(section)可以设置独立的页眉,因此需要遍历所有节的页眉进行处理。
- 格式保留方案:直接修改段落
text会重置该段落的格式(如字体、字号等)。如果需要保留页眉原有格式,可以使用以下精细处理逻辑:
def replace_placeholders_with_format(doc, data): # 处理页眉并保留格式 for section in doc.sections: header = section.header for paragraph in header.paragraphs: # 合并当前段落的所有run文本 full_text = ''.join(run.text for run in paragraph.runs) # 替换占位符 for key, value in data.items(): full_text = full_text.replace(key, str(value)) # 清空现有run,重新添加一个包含替换后文本的run for run in paragraph.runs: run.text = '' paragraph.runs[0].text = full_text # 正文替换同理可修改以保留格式 # ...
内容的提问来源于stack exchange,提问作者FFFFFFFFF
相关产品推荐
相关产品推荐

