You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Pandas替换Docx文本时如何保留行格式?

问题描述

需要自动化替换Word模板中的占位符Sample Customer和对应数值,替换后需保留原段落的格式,但当前修改paragraph.text的方式会导致整行格式丢失。

解决方案

核心问题是直接修改paragraph.text会重置段落内所有run(Word中连续同格式的文本片段)的格式,正确做法是遍历段落中的每个run,在run内部替换文本,以此保留原run的格式。

修改后代码

from docx import Document
from docx.shared import Pt
import os

word_doc_path = r"\placeholder\placeholder.docx"

# 数据清洗(保留原逻辑)
valid_customers = excel_data.iloc[1:-1, 1].dropna().astype(str)
valid_customers = valid_customers[valid_customers != "Customer"]
valid_customers = valid_customers[valid_customers != "Customer "]
valid_customers = valid_customers[valid_customers != "Intensity"]

for customer in valid_customers:
    print(f"Processing customer: {customer}")
    customer_data = excel_data[excel_data.iloc[:, 1].astype(str) == customer]
    print(customer_data)
    
    if not customer_data.empty:
        value_2023 = round(customer_data.iloc[0, 4], 2)
        doc = Document(word_doc_path)
        print(f"Processing {customer} - Document Loaded")
        
        for paragraph in doc.paragraphs:
            # 替换Sample Customer:针对单个run操作
            if "Sample Customer" in paragraph.text:
                for run in paragraph.runs:
                    if "Sample Customer" in run.text:
                        run.text = run.text.replace("Sample Customer", customer)
            
            # 替换tonnes前的数值:适配数值与tonnes同run/异run两种情况
            if "tonnes" in paragraph.text:
                for run in paragraph.runs:
                    # 情况1:数值和tonnes在同一个run里
                    if "tonnes" in run.text:
                        text_parts = run.text.rsplit(maxsplit=1)
                        if len(text_parts) == 2 and text_parts[1] == "tonnes":
                            run.text = f"{value_2023} tonnes"
                    # 情况2:数值是单独的run,紧跟tonnes的run
                    elif run.text.strip().replace('.','').isdigit():
                        run.text = str(value_2023)
        
        # 每个客户生成独立文件,避免覆盖
        word_save_path = fr'\placeholder\{customer}_document.docx'
        if os.path.exists(word_save_path):
            os.remove(word_save_path)
        doc.save(word_save_path)
        print(f"{customer} - Word Document Saved at: {word_save_path}")

关键说明

  1. 禁止直接修改paragraph.text:该操作会将段落内所有run合并为一个,彻底丢失原有的字体、大小、加粗等格式。
  2. 基于run操作:Word的格式是绑定在run上的,仅修改run内的文本,会自动继承该run的原有格式。
  3. 适配多场景替换:针对数值与"tonnes"是否在同一run的两种情况做处理,覆盖模板的常见排版方式。
  4. 避免文件覆盖:原代码会循环覆盖同一个文件,修改后按客户名命名输出文件,避免数据丢失。
额外提示

如果模板中Sample Customer跨多个run(比如部分文本加粗、部分普通),可以先拼接所有run的文本定位占位符位置,再拆分对应run的文本进行替换;或者在模板中给占位符添加书签,通过书签定位替换会更精准。

内容的提问来源于stack exchange,提问作者just_losing_it

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 12:22:23