You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Protractor实现光学字符识别(OCR)从图片提取文本?

Extracting Text from Images in Protractor (OCR for Automated Tests)

Hey there! Great question—yes, you absolutely can pull text from images in your Protractor automated tests using Optical Character Recognition (OCR). Let’s break down the most practical, self-contained approach using a JavaScript-based OCR library that fits right into your Protractor workflow.

The Go-To Tool: Tesseract.js

We’ll use Tesseract.js—a pure JavaScript OCR library that works seamlessly with Node.js (the environment Protractor runs in). No need for external installations like Tesseract CLI; it’s all handled via npm.

Step 1: Install Dependencies

First, add the required packages to your project:

npm install tesseract.js --save-dev

If you plan to crop screenshots to target specific image elements (recommended for accuracy), also install jimp for image processing:

npm install jimp --save-dev

Step 2: Basic OCR with Full Page Screenshot

Here’s a simple example that captures the entire page, runs OCR on it, and asserts the extracted text matches your expectations:

const { createWorker } = require('tesseract.js');

describe('OCR Test: Full Page Text Extraction', () => {
  it('should extract and verify text from the page screenshot', async () => {
    // Navigate to your test page
    await browser.get('https://your-test-application.com');
    
    // Capture the full page as a base64 string
    const screenshotBase64 = await browser.takeScreenshot();
    
    // Convert base64 to a buffer (Tesseract can handle base64 directly too)
    const imageBuffer = Buffer.from(screenshotBase64, 'base64');
    
    // Initialize Tesseract worker with English language pack
    const worker = await createWorker('eng');
    
    // Run OCR on the image buffer
    const { data: { text } } = await worker.recognize(imageBuffer);
    
    // Log and assert the extracted text
    console.log('Full Page Extracted Text:', text);
    expect(text).toContain('Your Expected Text Here');
    
    // Clean up the worker to free resources
    await worker.terminate();
  });
});

Step 3: Targeted OCR for Specific Image Elements

If you only need text from a particular image (like a captcha, logo with text, or dynamic image), crop the full screenshot to the element’s bounds first. Here’s how:

const { createWorker } = require('tesseract.js');
const Jimp = require('jimp');

describe('OCR Test: Targeted Image Text Extraction', () => {
  it('should extract text from a specific image element', async () => {
    await browser.get('https://your-test-application.com');
    
    // Locate the target image element
    const targetImage = element(by.css('#your-image-selector'));
    
    // Get the element's position and size on the page
    const elementLocation = await targetImage.getLocation();
    const elementSize = await targetImage.getSize();
    
    // Capture full page screenshot
    const screenshotBase64 = await browser.takeScreenshot();
    const imageBuffer = Buffer.from(screenshotBase64, 'base64');
    
    // Crop the screenshot to the element's area using Jimp
    const fullImage = await Jimp.read(imageBuffer);
    const croppedImage = fullImage.crop(
      elementLocation.x,
      elementLocation.y,
      elementSize.width,
      elementSize.height
    );
    
    // Convert cropped image to buffer for Tesseract
    const croppedBuffer = await croppedImage.getBufferAsync(Jimp.MIME_PNG);
    
    // Run OCR on the cropped image
    const worker = await createWorker('eng');
    const { data: { text } } = await worker.recognize(croppedBuffer);
    
    // Verify the extracted text (trim to remove extra whitespace)
    const cleanedText = text.trim();
    console.log('Extracted Text from Target Image:', cleanedText);
    expect(cleanedText).toEqual('Exact Text You Expect');
    
    await worker.terminate();
  });
});

Pro Tips for Better OCR Accuracy

  • Use the right language pack: Replace 'eng' in createWorker() with other language codes (e.g., 'spa' for Spanish, 'fra' for French) if your text isn’t in English.
  • Preprocess images: Improve accuracy by converting images to grayscale or applying thresholding before OCR:
    const processedImage = fullImage.grayscale().threshold(128);
    
  • Ensure clear images: Test with high-resolution screenshots, avoid blurry or low-contrast images—OCR works best with crisp, well-lit text.

Alternative: Cloud OCR APIs (Less Ideal for Local Tests)

While you could use cloud APIs like Google Cloud Vision or AWS Textract, these require API keys, network calls, and add latency to your tests. Tesseract.js is far better for fast, self-contained automated testing workflows.

内容的提问来源于stack exchange,提问作者Jonny

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:37:44