使用Python Pandas替换Docx文本时如何保留行格式?
问题描述
需要自动化替换Word模板中的占位符Sample Customer和对应数值,替换后需保留原段落的格式,但当前修改paragraph.text的方式会导致整行格式丢失。
解决方案
核心问题是直接修改paragraph.text会重置段落内所有run(Word中连续同格式的文本片段)的格式,正确做法是遍历段落中的每个run,在run内部替换文本,以此保留原run的格式。
修改后代码
from docx import Document from docx.shared import Pt import os word_doc_path = r"\placeholder\placeholder.docx" # 数据清洗(保留原逻辑) valid_customers = excel_data.iloc[1:-1, 1].dropna().astype(str) valid_customers = valid_customers[valid_customers != "Customer"] valid_customers = valid_customers[valid_customers != "Customer "] valid_customers = valid_customers[valid_customers != "Intensity"] for customer in valid_customers: print(f"Processing customer: {customer}") customer_data = excel_data[excel_data.iloc[:, 1].astype(str) == customer] print(customer_data) if not customer_data.empty: value_2023 = round(customer_data.iloc[0, 4], 2) doc = Document(word_doc_path) print(f"Processing {customer} - Document Loaded") for paragraph in doc.paragraphs: # 替换Sample Customer:针对单个run操作 if "Sample Customer" in paragraph.text: for run in paragraph.runs: if "Sample Customer" in run.text: run.text = run.text.replace("Sample Customer", customer) # 替换tonnes前的数值:适配数值与tonnes同run/异run两种情况 if "tonnes" in paragraph.text: for run in paragraph.runs: # 情况1:数值和tonnes在同一个run里 if "tonnes" in run.text: text_parts = run.text.rsplit(maxsplit=1) if len(text_parts) == 2 and text_parts[1] == "tonnes": run.text = f"{value_2023} tonnes" # 情况2:数值是单独的run,紧跟tonnes的run elif run.text.strip().replace('.','').isdigit(): run.text = str(value_2023) # 每个客户生成独立文件,避免覆盖 word_save_path = fr'\placeholder\{customer}_document.docx' if os.path.exists(word_save_path): os.remove(word_save_path) doc.save(word_save_path) print(f"{customer} - Word Document Saved at: {word_save_path}")
关键说明
- 禁止直接修改
paragraph.text:该操作会将段落内所有run合并为一个,彻底丢失原有的字体、大小、加粗等格式。 - 基于run操作:Word的格式是绑定在run上的,仅修改run内的文本,会自动继承该run的原有格式。
- 适配多场景替换:针对数值与"tonnes"是否在同一run的两种情况做处理,覆盖模板的常见排版方式。
- 避免文件覆盖:原代码会循环覆盖同一个文件,修改后按客户名命名输出文件,避免数据丢失。
额外提示
如果模板中Sample Customer跨多个run(比如部分文本加粗、部分普通),可以先拼接所有run的文本定位占位符位置,再拆分对应run的文本进行替换;或者在模板中给占位符添加书签,通过书签定位替换会更精准。
内容的提问来源于stack exchange,提问作者just_losing_it
相关产品推荐
相关产品推荐

