You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python调用微软截图工具实现高精度OCR文本提取?

用Python调用微软内置OCR替代PyTesseract的解决方案

微软截图工具的文本提取依赖Windows内置的OCR引擎,准确率高且无需额外安装Tesseract,完全可以通过Python实现调用,以下是两种可行方案:

方案1:通过PowerShell调用Windows OCR API

直接用Python的subprocess执行PowerShell脚本,调用系统内置的OCR引擎处理图片:

import subprocess
import json

def get_windows_ocr_result(image_path):
    ps_cmd = f"""
    $ocrEngine = [Windows.Media.Ocr.OcrEngine]::TryCreateFromUserProfileLanguages()
    $bitmapDecoder = [Windows.Media.Imaging.BitmapDecoder]::CreateAsync([System.Uri]::new('{image_path}')).GetAwaiter().GetResult()
    $ocrResult = $ocrEngine.RecognizeAsync($bitmapDecoder.Frames[0]).GetAwaiter().GetResult()
    $ocrResult.Text | ConvertTo-Json
    """
    proc = subprocess.run(["powershell", "-Command", ps_cmd], capture_output=True, text=True)
    if proc.returncode == 0:
        return json.loads(proc.stdout)
    raise RuntimeError(f"OCR执行失败: {proc.stderr}")

注意事项:

  • 仅支持Windows 10 1903及以上版本
  • 传入的image_path必须是绝对路径

方案2:使用第三方封装库winocr

winocr是专门封装Windows OCR API的Python库,使用更简洁:

  1. 安装库:
pip install winocr
  1. 调用示例:
from winocr import OcrEngine

# 初始化引擎
ocr_engine = OcrEngine()
# 识别本地图片
ocr_result = ocr_engine.recognize_file("你的截图路径.png")
# 获取识别文本
extracted_text = ocr_result.text

适配你的单词搜索项目

替换原代码中PyTesseract的调用逻辑即可,结合自动截图工具(如pyautogui)可实现完整流程:

import pyautogui

# 截取屏幕指定区域(按需调整坐标和尺寸)
screenshot = pyautogui.screenshot(region=(100, 100, 500, 300))
temp_img_path = "temp_snip.png"
screenshot.save(temp_img_path)

# 调用OCR获取文本
extracted_text = get_windows_ocr_result(temp_img_path)  # 或用winocr的方法
# 生成待搜索单词列表
word_list = [word.strip() for word in extracted_text.split("\n") if word.strip()]

此方案完全替代PyTesseract,既提升识别准确率,又无需用户额外安装Tesseract-OCR。

内容的提问来源于stack exchange,提问作者Johnny Bravo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 05:05:19