You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于图像定位Selenium元素?求子图坐标检测方案

Reliable Libraries for Sub-Image Matching in Your Selenium Custom Locator

Great idea adapting Appium's image-based locator approach for Selenium! You’re absolutely right that standard locators (like ID, XPath) can fall short in scenarios where elements are dynamic or lack unique attributes. Your ByImage class is a solid start—let’s fill in the gap with robust sub-image matching libraries that will give you the precise coordinates you need.

OpenCV is the gold standard for computer vision tasks, including template matching. It’s fast, accurate, and handles edge cases like minor color variations or scaling better than lightweight alternatives.

Step 1: Add Dependency

For Maven, include this in your pom.xml:

<dependency>
    <groupId>org.openpnp</groupId>
    <artifactId>opencv</artifactId>
    <version>4.5.5-2</version>
</dependency>

Step 2: Integrate into Your findElements Method

Replace the placeholder coordinate logic with this OpenCV implementation:

@Override
public List<WebElement> findElements(SearchContext context) {
    try {
        // Get current screenshot
        byte[] screenshotByte = ((TakesScreenshot) context).getScreenshotAs(OutputType.BYTES);
        byte[] subImgToFindByte = DatatypeConverter.parseBase64Binary(imageBase64String);

        // Convert byte arrays to OpenCV Mat objects
        Mat screenshotMat = Imgcodecs.imdecode(new MatOfByte(screenshotByte), Imgcodecs.IMREAD_COLOR);
        Mat subImgMat = Imgcodecs.imdecode(new MatOfByte(subImgToFindByte), Imgcodecs.IMREAD_COLOR);

        // Perform template matching (TM_CCOEFF_NORMED is great for most cases)
        Mat result = new Mat();
        Imgproc.matchTemplate(screenshotMat, subImgMat, result, Imgproc.TM_CCOEFF_NORMED);

        // Find the best match location
        Core.MinMaxLocResult mmr = Core.minMaxLoc(result);
        Point matchLoc = mmr.maxLoc;

        // Set a threshold to filter weak matches (adjust based on your needs)
        double matchThreshold = 0.8;
        if (mmr.maxVal < matchThreshold) {
            throw new NoSuchElementException("Sub-image match score (" + mmr.maxVal + ") below threshold (" + matchThreshold + ")");
        }

        // Calculate center coordinates
        int centerX = (int) (matchLoc.x + (subImgMat.cols() / 2.0));
        int centerY = (int) (matchLoc.y + (subImgMat.rows() / 2.0));

        // Get element at center point
        JavascriptExecutor js = ((JavascriptExecutor) context);
        return (List<WebElement>) js.executeScript("return document.elementsFromPoint(arguments[0], arguments[1]);", centerX, centerY);
    } catch (Exception e) {
        throw new NoSuchElementException("Failed to find element via image match", e);
    }
}

Pros & Cons

  • Pros: Fast, handles scaling/color variations, widely supported.
  • Cons: Slightly larger dependency, requires basic familiarity with OpenCV concepts.

2. Imgscalr (Lightweight Alternative)

If you don’t need full computer vision power and want a simpler, lighter library, Imgscalr is a great choice. It’s focused on image manipulation but can be adapted for sub-image matching with a sliding window approach.

Step 1: Add Dependency

<dependency>
    <groupId>org.imgscalr</groupId>
    <artifactId>imgscalr-lib</artifactId>
    <version>4.2</version>
</dependency>

Step 2: Integrate Matching Logic

@Override
public List<WebElement> findElements(SearchContext context) {
    try {
        byte[] screenshotByte = ((TakesScreenshot) context).getScreenshotAs(OutputType.BYTES);
        byte[] subImgToFindByte = DatatypeConverter.parseBase64Binary(imageBase64String);

        BufferedImage screenshotImg = ImageIO.read(new ByteArrayInputStream(screenshotByte));
        BufferedImage subImg = ImageIO.read(new ByteArrayInputStream(subImgToFindByte));

        int subWidth = subImg.getWidth();
        int subHeight = subImg.getHeight();
        int screenshotWidth = screenshotImg.getWidth();
        int screenshotHeight = screenshotImg.getHeight();

        ImageDiffer differ = new ImageDiffer();
        double tolerance = 0.1; // Allow minor pixel variations

        double x = -1;
        double y = -1;

        // Sliding window to find matching sub-section
        for (int xPos = 0; xPos <= screenshotWidth - subWidth; xPos++) {
            for (int yPos = 0; yPos <= screenshotHeight - subHeight; yPos++) {
                BufferedImage subSection = screenshotImg.getSubimage(xPos, yPos, subWidth, subHeight);
                ImageDiff diff = differ.makeDiff(subSection, subImg, tolerance);
                if (!diff.isDiff()) {
                    x = xPos;
                    y = yPos;
                    break;
                }
            }
            if (x != -1) break;
        }

        if (x == -1) {
            throw new NoSuchElementException("Sub-image not found in screenshot");
        }

        int centerX = (int) (x + (subWidth / 2.0));
        int centerY = (int) (y + (subHeight / 2.0));

        JavascriptExecutor js = ((JavascriptExecutor) context);
        return (List<WebElement>) js.executeScript("return document.elementsFromPoint(arguments[0], arguments[1]);", centerX, centerY);
    } catch (Exception e) {
        throw new NoSuchElementException("Failed to find element via image match", e);
    }
}

Pros & Cons

  • Pros: Lightweight, easy to integrate, no complex CV setup.
  • Cons: Slower for large images, less accurate with scaling/color changes.

3. Apache Commons Imaging (Basic Pixel Matching)

If you want to avoid external CV libraries entirely, Apache Commons Imaging provides basic image utilities you can use for manual pixel-by-pixel matching. This is best for simple, static UI elements with no variation.

Step 1: Add Dependency

<dependency>
    <groupId>org.apache.commons</groupId>
    <artifactId>commons-imaging</artifactId>
    <version>1.0-alpha3</version>
</dependency>

Step 2: Integrate Pixel Comparison

@Override
public List<WebElement> findElements(SearchContext context) {
    try {
        byte[] screenshotByte = ((TakesScreenshot) context).getScreenshotAs(OutputType.BYTES);
        byte[] subImgToFindByte = DatatypeConverter.parseBase64Binary(imageBase64String);

        BufferedImage screenshotImg = Imaging.getBufferedImage(screenshotByte);
        BufferedImage subImg = Imaging.getBufferedImage(subImgToFindByte);

        int subWidth = subImg.getWidth();
        int subHeight = subImg.getHeight();
        int screenshotWidth = screenshotImg.getWidth();
        int screenshotHeight = screenshotImg.getHeight();

        double x = -1;
        double y = -1;
        int colorTolerance = 10; // Allow small color differences

        // Sliding window pixel comparison
        outerLoop:
        for (int xPos = 0; xPos <= screenshotWidth - subWidth; xPos++) {
            for (int yPos = 0; yPos <= screenshotHeight - subHeight; yPos++) {
                boolean match = true;
                for (int i = 0; i < subWidth; i++) {
                    for (int j = 0; j < subHeight; j++) {
                        int screenshotPixel = screenshotImg.getRGB(xPos + i, yPos + j);
                        int subPixel = subImg.getRGB(i, j);

                        // Check RGB channels with tolerance
                        int r1 = (screenshotPixel >> 16) & 0xFF;
                        int g1 = (screenshotPixel >> 8) & 0xFF;
                        int b1 = screenshotPixel & 0xFF;
                        int r2 = (subPixel >> 16) & 0xFF;
                        int g2 = (subPixel >> 8) & 0xFF;
                        int b2 = subPixel & 0xFF;

                        if (Math.abs(r1 - r2) > colorTolerance
                                || Math.abs(g1 - g2) > colorTolerance
                                || Math.abs(b1 - b2) > colorTolerance) {
                            match = false;
                            break;
                        }
                    }
                    if (!match) break;
                }
                if (match) {
                    x = xPos;
                    y = yPos;
                    break outerLoop;
                }
            }
        }

        if (x == -1) {
            throw new NoSuchElementException("Sub-image not found in screenshot");
        }

        int centerX = (int) (x + (subWidth / 2.0));
        int centerY = (int) (y + (subHeight / 2.0));

        JavascriptExecutor js = ((JavascriptExecutor) context);
        return (List<WebElement>) js.executeScript("return document.elementsFromPoint(arguments[0], arguments[1]);", centerX, centerY);
    } catch (Exception e) {
        throw new NoSuchElementException("Failed to find element via image match", e);
    }
}

Pros & Cons

  • Pros: No external CV dependencies, full control over matching logic.
  • Cons: Slowest option, prone to false negatives with any UI variation.

Additional Optimization Tips

  • Match Thresholds: Always use a threshold (like the 0.8 value in OpenCV) to filter out weak matches and avoid false positives.
  • Resolution Normalization: If your tests run on different screen sizes, resize the screenshot and sub-image to a consistent resolution before matching.
  • Element Validation: After retrieving the element via elementsFromPoint, add checks (e.g., verifying the element’s tag name, text, or attributes) to ensure you’ve found the correct one.
  • Caching: If you’re matching the same sub-image multiple times, cache the converted Mat or BufferedImage to improve performance.

内容的提问来源于stack exchange,提问作者Saurabh Gaur

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:57:45