You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Karate UI与Karate Robot高亮窗口并识别图像文本及坐标?

Karate UI + Karate Robot 实现窗口高亮、图像文本识别及坐标获取方案

一、高亮指定窗口

通过Karate Robot的窗口操作API定位目标窗口,结合矩形绘制实现高亮:

  • 定位目标窗口:根据标题关键字、进程名等属性锁定窗口
    // 按标题模糊匹配定位
    let targetWindow = robot.window('目标窗口标题关键字');
    // 或按进程名+标题片段组合定位
    // let targetWindow = robot.window({ process: 'xxx.exe', titleContains: '窗口标题片段' });
    
  • 获取窗口边界信息:提取窗口的屏幕坐标与尺寸
    let windowBounds = targetWindow.bounds();
    // windowBounds包含x(起始横坐标)、y(起始纵坐标)、width(宽度)、height(高度)四个属性
    
  • 绘制高亮边框:用红色粗边框标记窗口区域,配合延迟让高亮可见
    // 绘制红色边框,厚度2px
    robot.rectangle(windowBounds.x, windowBounds.y, windowBounds.width, windowBounds.height, { color: 'red', thickness: 2 });
    // 暂停3秒,方便观察高亮效果
    robot.delay(3000);
    

二、识别窗口内图像文本

Karate本身无内置OCR能力,需结合Tess4J(Tesseract的Java封装)实现文本识别:

  1. 依赖准备:在项目构建文件(如pom.xml)中引入Tess4J依赖

    <dependency>
        <groupId>net.sourceforge.tess4j</groupId>
        <artifactId>tess4j</artifactId>
        <version>5.6.0</version>
    </dependency>
    

    同时下载对应语言的tessdata数据包(如中文chi_sim.traineddata),放置到项目指定目录。

  2. 截取窗口/区域截图:

    // 截取整个目标窗口的截图,保存为本地文件
    robot.screenshot({ window: targetWindow, file: 'window-snapshot.png' });
    // 或截取窗口内指定区域(示例:从窗口左上角偏移10px,截取200x50的区域)
    // robot.screenshot({ x: windowBounds.x + 10, y: windowBounds.y + 10, width: 200, height: 50, file: 'region-snapshot.png' });
    
  3. 调用OCR识别文本:通过Karate的Java互操作能力调用Tess4J

    var Tesseract = Java.type('net.sourceforge.tess4j.Tesseract');
    var tess = new Tesseract();
    // 设置tessdata数据包目录路径
    tess.setDatapath('./tessdata');
    // 设置识别语言,中文用'chi_sim',英文用'eng'
    tess.setLanguage('chi_sim');
    // 识别截图中的文本
    let recognizedText = tess.doOCR(new Java.io.File('window-snapshot.png'));
    karate.log('识别到的文本内容:', recognizedText);
    

三、获取目标文本的对应坐标

利用Tess4J的getWords方法获取每个识别文本的边界框,转换为屏幕绝对坐标:

var Word = Java.type('net.sourceforge.tess4j.Word');
// 获取所有识别到的单词及边界框信息
let wordList = tess.getWords(new Java.io.File('window-snapshot.png'), 0);
// 遍历查找目标文本
let targetText = '需要定位的文本内容';
for (let i = 0; i < wordList.size(); i++) {
    let currentWord = wordList.get(i);
    if (currentWord.getText().trim() === targetText) {
        let textRect = currentWord.getBoundingBox();
        // 转换为屏幕绝对坐标(截图坐标 + 窗口起始坐标)
        let absX = windowBounds.x + textRect.x;
        let absY = windowBounds.y + textRect.y;
        let textWidth = textRect.width;
        let textHeight = textRect.height;
        // 高亮定位到的文本区域
        robot.rectangle(absX, absY, textWidth, textHeight, { color: 'green', thickness: 2 });
        karate.log(`目标文本"${targetText}"的屏幕坐标:(${absX}, ${absY}),尺寸:${textWidth}x${textHeight}`);
    }
}

注意事项

  • 窗口定位需确保属性唯一,避免匹配到多个窗口
  • OCR识别准确率受图像清晰度、字体、背景干扰影响,可通过预处理截图(如灰度化、降噪)优化
  • 需提前配置tessdata路径及对应语言包,否则会识别失败

内容的提问来源于stack exchange,提问作者Lavanya Yarasu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 10:50:47