You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python修改Word文档中指定单词的字体颜色?

解决Word文档指定单词/短语的字体高亮问题

原代码存在的核心问题

  1. 直接操作底层XML元素,导致整个run变色而非仅目标单词/短语
  2. 文本替换逻辑错误,会丢失原内容且无法准确定位匹配项
  3. 未处理多词短语匹配,也未做单词边界校验,易出现子串误匹配(比如高亮"sand"里的"and")

修正后的完整代码

from docx import Document
from docx.shared import RGBColor
import re

def highlight_words(document, words):
    # 区分多词短语和单个单词,构建正则匹配规则
    # 优先匹配长短语,避免被拆分为单个单词
    phrase_patterns = [re.escape(phrase) for phrase in words if ' ' in phrase]
    word_patterns = [fr'\b{re.escape(word)}\b' for word in words if ' ' not in word]
    combined_pattern = re.compile('|'.join(sorted(phrase_patterns + word_patterns, key=lambda x: -len(x))))

    for paragraph in document.paragraphs:
        # 从后往前遍历run,避免拆分时索引偏移
        for run in reversed(paragraph.runs):
            text = run.text
            if not text:
                continue
            matches = list(combined_pattern.finditer(text))
            if not matches:
                continue
            
            current_pos = len(text)
            # 从后往前处理匹配项,确保拆分位置正确
            for match in reversed(matches):
                match_start, match_end = match.span()
                # 拆分匹配内容之后的文本为新run
                if match_end < current_pos:
                    new_run = run._insert_run(match_end)
                    new_run.text = text[match_end:current_pos]
                    # 复制原run的格式属性
                    new_run.font.bold = run.font.bold
                    new_run.font.italic = run.font.italic
                    new_run.font.size = run.font.size
                # 拆分匹配内容为独立run并设置高亮颜色
                highlight_run = run._insert_run(match_start)
                highlight_run.text = text[match_start:match_end]
                highlight_run.font.color.rgb = RGBColor(255, 0, 0)
                current_pos = match_start
            # 更新原run的文本为剩余未处理部分
            run.text = text[:current_pos]

def analyze(filename):
    causal_words_to_highlight = ["and", "then"]
    weak_words_to_highlight = ["got", "gots", "put", "really", "very", "said", "good", "bad", "go", "going", "went", "come", "comes", "came", "say", "get", "see", "saw", "nice", "mean", "pretty", "ugly", "big", "a lot", "fun", "well", "little", "fast", "slow", "small", "interesting", "fetid"]
    there_words_to_highlight = ["There is", "There was", "There are", "There were"]
    story = import_story(filename)
    if story is not None:
        doc = Document()

        # 按换行拆分文本并添加为段落
        for paragraph_text in story.split("\n"):
            paragraph = doc.add_paragraph()
            paragraph.add_run(paragraph_text)

        # 合并高亮列表并去重,避免重复处理
        words_to_highlight = list(set(
            causal_words_to_highlight 
            + weak_words_to_highlight 
            + there_words_to_highlight 
            + ["just"]
        ))
        highlight_words(doc, words_to_highlight)

        # 保存高亮后的文档
        doc.save(f"{filename[:-5]}_highlighted.docx")

关键改进说明

  • 精准匹配:通过正则表达式的单词边界\b确保只高亮完整单词,同时优先匹配多词短语,避免拆分错误
  • 安全拆分run:从后往前遍历和拆分run,防止因run数量变化导致的索引错位
  • 格式保留:拆分新run时复制原run的格式(加粗、斜体、字号等),仅修改目标内容的颜色
  • 去重优化:对高亮列表去重,减少不必要的重复处理

内容的提问来源于stack exchange,提问作者walle ras

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 03:23:17