如何用Python修改Word文档中指定单词的字体颜色?
解决Word文档指定单词/短语的字体高亮问题
原代码存在的核心问题
- 直接操作底层XML元素,导致整个run变色而非仅目标单词/短语
- 文本替换逻辑错误,会丢失原内容且无法准确定位匹配项
- 未处理多词短语匹配,也未做单词边界校验,易出现子串误匹配(比如高亮"sand"里的"and")
修正后的完整代码
from docx import Document from docx.shared import RGBColor import re def highlight_words(document, words): # 区分多词短语和单个单词,构建正则匹配规则 # 优先匹配长短语,避免被拆分为单个单词 phrase_patterns = [re.escape(phrase) for phrase in words if ' ' in phrase] word_patterns = [fr'\b{re.escape(word)}\b' for word in words if ' ' not in word] combined_pattern = re.compile('|'.join(sorted(phrase_patterns + word_patterns, key=lambda x: -len(x)))) for paragraph in document.paragraphs: # 从后往前遍历run,避免拆分时索引偏移 for run in reversed(paragraph.runs): text = run.text if not text: continue matches = list(combined_pattern.finditer(text)) if not matches: continue current_pos = len(text) # 从后往前处理匹配项,确保拆分位置正确 for match in reversed(matches): match_start, match_end = match.span() # 拆分匹配内容之后的文本为新run if match_end < current_pos: new_run = run._insert_run(match_end) new_run.text = text[match_end:current_pos] # 复制原run的格式属性 new_run.font.bold = run.font.bold new_run.font.italic = run.font.italic new_run.font.size = run.font.size # 拆分匹配内容为独立run并设置高亮颜色 highlight_run = run._insert_run(match_start) highlight_run.text = text[match_start:match_end] highlight_run.font.color.rgb = RGBColor(255, 0, 0) current_pos = match_start # 更新原run的文本为剩余未处理部分 run.text = text[:current_pos] def analyze(filename): causal_words_to_highlight = ["and", "then"] weak_words_to_highlight = ["got", "gots", "put", "really", "very", "said", "good", "bad", "go", "going", "went", "come", "comes", "came", "say", "get", "see", "saw", "nice", "mean", "pretty", "ugly", "big", "a lot", "fun", "well", "little", "fast", "slow", "small", "interesting", "fetid"] there_words_to_highlight = ["There is", "There was", "There are", "There were"] story = import_story(filename) if story is not None: doc = Document() # 按换行拆分文本并添加为段落 for paragraph_text in story.split("\n"): paragraph = doc.add_paragraph() paragraph.add_run(paragraph_text) # 合并高亮列表并去重,避免重复处理 words_to_highlight = list(set( causal_words_to_highlight + weak_words_to_highlight + there_words_to_highlight + ["just"] )) highlight_words(doc, words_to_highlight) # 保存高亮后的文档 doc.save(f"{filename[:-5]}_highlighted.docx")
关键改进说明
- 精准匹配:通过正则表达式的单词边界
\b确保只高亮完整单词,同时优先匹配多词短语,避免拆分错误 - 安全拆分run:从后往前遍历和拆分run,防止因run数量变化导致的索引错位
- 格式保留:拆分新run时复制原run的格式(加粗、斜体、字号等),仅修改目标内容的颜色
- 去重优化:对高亮列表去重,减少不必要的重复处理
内容的提问来源于stack exchange,提问作者walle ras
相关产品推荐
相关产品推荐

