You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python实现Word文件选中文本高亮?求精准可用API/库

Word文档指定文本高亮方案推荐

微软官方方案

1. Open XML SDK

这是处理docx格式的官方底层库,能精准定位文本并设置高亮,支持跨段落、带格式文本匹配等复杂场景。核心思路是遍历文档中的文本节点,匹配目标内容后设置Highlight属性:

// C#示例片段
using DocumentFormat.OpenXml.Wordprocessing;
using DocumentFormat.OpenXml.Packaging;

using (WordprocessingDocument doc = WordprocessingDocument.Open("test.docx", true))
{
    foreach (var text in doc.MainDocumentPart.Document.Descendants<Text>())
    {
        if (text.Text.Contains("目标文本"))
        {
            var run = text.Parent as Run;
            run.RunProperties ??= new RunProperties();
            run.RunProperties.Highlight ??= new Highlight { Val = HighlightColorValues.Yellow };
        }
    }
    doc.MainDocumentPart.Document.Save();
}

2. Office Interop

直接调用本地Word应用程序,完全复刻手动高亮的逻辑,精度最高,但仅支持Windows环境,需提前安装Office:

# Python示例(依赖pywin32)
import win32com.client as win32

word = win32.Dispatch("Word.Application")
doc = word.Documents.Open("test.docx")
selection = word.Selection

# 循环查找并高亮所有匹配项
while selection.Find.Execute("目标文本", Forward=True):
    selection.Range.HighlightColorIndex = 7  # 7对应黄色高亮
doc.Save()
doc.Close()
word.Quit()

关于Microsoft Graph API

如果之前使用效果不佳,大概率是未精准处理文本定位——Graph API的update操作需要明确指定文本范围的ID,可结合search接口获取匹配位置后再执行高亮,适合云端文档批量处理,但对复杂格式文本的支持仍有局限。

第三方库推荐

1. Aspose.Words

商业级库,支持Python、C#等多语言,能处理几乎所有Word格式场景,高亮逻辑简单直接,无需依赖Office:

# Python示例
import aspose.words as aw

doc = aw.Document("test.docx")
find_options = aw.replacing.FindReplaceOptions()
find_options.apply_highlight = aw.Color.YELLOW

doc.range.replace("目标文本", "", find_options)
doc.save("highlighted.docx")

2. python-docx进阶用法

如果仍想使用python-docx,可通过遍历所有Run节点,拆分匹配的文本片段来实现精准高亮(原库默认不支持跨Run匹配,这是之前效果差的核心原因):

from docx import Document
from docx.enum.text import WD_COLOR_INDEX

def highlight_text(doc, target):
    for para in doc.paragraphs:
        runs = para.runs
        i = 0
        while i < len(runs):
            run = runs[i]
            if target in run.text:
                # 拆分Run为匹配前、匹配中、匹配后三部分
                split_idx = run.text.index(target)
                pre_text = run.text[:split_idx]
                match_text = run.text[split_idx:split_idx+len(target)]
                post_text = run.text[split_idx+len(target):]
                
                # 修改当前Run为匹配前文本
                run.text = pre_text
                # 插入高亮的匹配Run
                highlight_run = para.add_run(match_text)
                highlight_run.font.highlight_color = WD_COLOR_INDEX.YELLOW
                # 插入匹配后文本的Run
                if post_text:
                    para.add_run(post_text)
                # 更新Run列表
                runs = para.runs
            i += 1

doc = Document("test.docx")
highlight_text(doc, "目标文本")
doc.save("highlighted.docx")

内容的提问来源于stack exchange,提问作者KARTIK CHOUDHARY

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 22:43:01