如何基于图像定位Selenium元素?求子图坐标检测方案
Great idea adapting Appium's image-based locator approach for Selenium! You’re absolutely right that standard locators (like ID, XPath) can fall short in scenarios where elements are dynamic or lack unique attributes. Your ByImage class is a solid start—let’s fill in the gap with robust sub-image matching libraries that will give you the precise coordinates you need.
1. OpenCV for Java (Most Robust & Recommended)
OpenCV is the gold standard for computer vision tasks, including template matching. It’s fast, accurate, and handles edge cases like minor color variations or scaling better than lightweight alternatives.
Step 1: Add Dependency
For Maven, include this in your pom.xml:
<dependency> <groupId>org.openpnp</groupId> <artifactId>opencv</artifactId> <version>4.5.5-2</version> </dependency>
Step 2: Integrate into Your findElements Method
Replace the placeholder coordinate logic with this OpenCV implementation:
@Override public List<WebElement> findElements(SearchContext context) { try { // Get current screenshot byte[] screenshotByte = ((TakesScreenshot) context).getScreenshotAs(OutputType.BYTES); byte[] subImgToFindByte = DatatypeConverter.parseBase64Binary(imageBase64String); // Convert byte arrays to OpenCV Mat objects Mat screenshotMat = Imgcodecs.imdecode(new MatOfByte(screenshotByte), Imgcodecs.IMREAD_COLOR); Mat subImgMat = Imgcodecs.imdecode(new MatOfByte(subImgToFindByte), Imgcodecs.IMREAD_COLOR); // Perform template matching (TM_CCOEFF_NORMED is great for most cases) Mat result = new Mat(); Imgproc.matchTemplate(screenshotMat, subImgMat, result, Imgproc.TM_CCOEFF_NORMED); // Find the best match location Core.MinMaxLocResult mmr = Core.minMaxLoc(result); Point matchLoc = mmr.maxLoc; // Set a threshold to filter weak matches (adjust based on your needs) double matchThreshold = 0.8; if (mmr.maxVal < matchThreshold) { throw new NoSuchElementException("Sub-image match score (" + mmr.maxVal + ") below threshold (" + matchThreshold + ")"); } // Calculate center coordinates int centerX = (int) (matchLoc.x + (subImgMat.cols() / 2.0)); int centerY = (int) (matchLoc.y + (subImgMat.rows() / 2.0)); // Get element at center point JavascriptExecutor js = ((JavascriptExecutor) context); return (List<WebElement>) js.executeScript("return document.elementsFromPoint(arguments[0], arguments[1]);", centerX, centerY); } catch (Exception e) { throw new NoSuchElementException("Failed to find element via image match", e); } }
Pros & Cons
- Pros: Fast, handles scaling/color variations, widely supported.
- Cons: Slightly larger dependency, requires basic familiarity with OpenCV concepts.
2. Imgscalr (Lightweight Alternative)
If you don’t need full computer vision power and want a simpler, lighter library, Imgscalr is a great choice. It’s focused on image manipulation but can be adapted for sub-image matching with a sliding window approach.
Step 1: Add Dependency
<dependency> <groupId>org.imgscalr</groupId> <artifactId>imgscalr-lib</artifactId> <version>4.2</version> </dependency>
Step 2: Integrate Matching Logic
@Override public List<WebElement> findElements(SearchContext context) { try { byte[] screenshotByte = ((TakesScreenshot) context).getScreenshotAs(OutputType.BYTES); byte[] subImgToFindByte = DatatypeConverter.parseBase64Binary(imageBase64String); BufferedImage screenshotImg = ImageIO.read(new ByteArrayInputStream(screenshotByte)); BufferedImage subImg = ImageIO.read(new ByteArrayInputStream(subImgToFindByte)); int subWidth = subImg.getWidth(); int subHeight = subImg.getHeight(); int screenshotWidth = screenshotImg.getWidth(); int screenshotHeight = screenshotImg.getHeight(); ImageDiffer differ = new ImageDiffer(); double tolerance = 0.1; // Allow minor pixel variations double x = -1; double y = -1; // Sliding window to find matching sub-section for (int xPos = 0; xPos <= screenshotWidth - subWidth; xPos++) { for (int yPos = 0; yPos <= screenshotHeight - subHeight; yPos++) { BufferedImage subSection = screenshotImg.getSubimage(xPos, yPos, subWidth, subHeight); ImageDiff diff = differ.makeDiff(subSection, subImg, tolerance); if (!diff.isDiff()) { x = xPos; y = yPos; break; } } if (x != -1) break; } if (x == -1) { throw new NoSuchElementException("Sub-image not found in screenshot"); } int centerX = (int) (x + (subWidth / 2.0)); int centerY = (int) (y + (subHeight / 2.0)); JavascriptExecutor js = ((JavascriptExecutor) context); return (List<WebElement>) js.executeScript("return document.elementsFromPoint(arguments[0], arguments[1]);", centerX, centerY); } catch (Exception e) { throw new NoSuchElementException("Failed to find element via image match", e); } }
Pros & Cons
- Pros: Lightweight, easy to integrate, no complex CV setup.
- Cons: Slower for large images, less accurate with scaling/color changes.
3. Apache Commons Imaging (Basic Pixel Matching)
If you want to avoid external CV libraries entirely, Apache Commons Imaging provides basic image utilities you can use for manual pixel-by-pixel matching. This is best for simple, static UI elements with no variation.
Step 1: Add Dependency
<dependency> <groupId>org.apache.commons</groupId> <artifactId>commons-imaging</artifactId> <version>1.0-alpha3</version> </dependency>
Step 2: Integrate Pixel Comparison
@Override public List<WebElement> findElements(SearchContext context) { try { byte[] screenshotByte = ((TakesScreenshot) context).getScreenshotAs(OutputType.BYTES); byte[] subImgToFindByte = DatatypeConverter.parseBase64Binary(imageBase64String); BufferedImage screenshotImg = Imaging.getBufferedImage(screenshotByte); BufferedImage subImg = Imaging.getBufferedImage(subImgToFindByte); int subWidth = subImg.getWidth(); int subHeight = subImg.getHeight(); int screenshotWidth = screenshotImg.getWidth(); int screenshotHeight = screenshotImg.getHeight(); double x = -1; double y = -1; int colorTolerance = 10; // Allow small color differences // Sliding window pixel comparison outerLoop: for (int xPos = 0; xPos <= screenshotWidth - subWidth; xPos++) { for (int yPos = 0; yPos <= screenshotHeight - subHeight; yPos++) { boolean match = true; for (int i = 0; i < subWidth; i++) { for (int j = 0; j < subHeight; j++) { int screenshotPixel = screenshotImg.getRGB(xPos + i, yPos + j); int subPixel = subImg.getRGB(i, j); // Check RGB channels with tolerance int r1 = (screenshotPixel >> 16) & 0xFF; int g1 = (screenshotPixel >> 8) & 0xFF; int b1 = screenshotPixel & 0xFF; int r2 = (subPixel >> 16) & 0xFF; int g2 = (subPixel >> 8) & 0xFF; int b2 = subPixel & 0xFF; if (Math.abs(r1 - r2) > colorTolerance || Math.abs(g1 - g2) > colorTolerance || Math.abs(b1 - b2) > colorTolerance) { match = false; break; } } if (!match) break; } if (match) { x = xPos; y = yPos; break outerLoop; } } } if (x == -1) { throw new NoSuchElementException("Sub-image not found in screenshot"); } int centerX = (int) (x + (subWidth / 2.0)); int centerY = (int) (y + (subHeight / 2.0)); JavascriptExecutor js = ((JavascriptExecutor) context); return (List<WebElement>) js.executeScript("return document.elementsFromPoint(arguments[0], arguments[1]);", centerX, centerY); } catch (Exception e) { throw new NoSuchElementException("Failed to find element via image match", e); } }
Pros & Cons
- Pros: No external CV dependencies, full control over matching logic.
- Cons: Slowest option, prone to false negatives with any UI variation.
Additional Optimization Tips
- Match Thresholds: Always use a threshold (like the 0.8 value in OpenCV) to filter out weak matches and avoid false positives.
- Resolution Normalization: If your tests run on different screen sizes, resize the screenshot and sub-image to a consistent resolution before matching.
- Element Validation: After retrieving the element via
elementsFromPoint, add checks (e.g., verifying the element’s tag name, text, or attributes) to ensure you’ve found the correct one. - Caching: If you’re matching the same sub-image multiple times, cache the converted
MatorBufferedImageto improve performance.
内容的提问来源于stack exchange,提问作者Saurabh Gaur

