You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python获取屏幕文本位置并实现屏幕按钮定位点击?

Yes, You Can Absolutely Build This!

Great question—this is totally achievable with Python, and you’re already halfway there knowing pytesseract. Let’s walk through how to get from text recognition to locating those buttons and clicking them.

Step 1: Get Text Positions with pytesseract

You might not realize this, but pytesseract can return more than just raw text—it gives you the bounding box coordinates for every piece of text it recognizes. This is exactly what you need to find where your target button text lives on the screen.

The image_to_data() method is your best bet here: it returns structured data (like a dictionary) with fields like left, top, width, height, and the recognized text. Here’s how to use it:

import pytesseract
import pyautogui
from PIL import Image

# Take a full-screen screenshot
screenshot = pyautogui.screenshot()
img = Image.frombytes("RGB", screenshot.size, screenshot.tobytes())

# Extract OCR data including position details
ocr_data = pytesseract.image_to_data(img, output_type=pytesseract.Output.DICT)

# Search for your target button text
target_text = "Submit"
for i in range(len(ocr_data["text"])):
    cleaned_text = ocr_data["text"][i].strip()
    if cleaned_text == target_text:
        # Pull the coordinates of the text block
        left = ocr_data["left"][i]
        top = ocr_data["top"][i]
        width = ocr_data["width"][i]
        height = ocr_data["height"][i]
        # Calculate the center of the button (for accurate clicking)
        click_x = left + width // 2
        click_y = top + height // 2
        print(f"Found button at ({click_x}, {click_y})")
        break

Step 2: Simulate the Mouse Click

Once you have the button’s center coordinates, pyautogui makes clicking trivial. Just add this line right after calculating click_x and click_y:

pyautogui.click(click_x, click_y)

Pro Tips to Boost Reliability

  • Preprocess your screenshot: OCR works way better with high-contrast images. Try converting to grayscale or applying a threshold to make text pop:
    from PIL import ImageOps
    img_gray = ImageOps.grayscale(img)
    # Apply threshold to turn text black and background white
    img_threshold = img_gray.point(lambda x: 0 if x < 128 else 255, '1')
    
  • Handle screen scaling: If your display uses scaling (like 125% or 150%), pyautogui’s coordinates might be off. Use pyautogui.displayMousePosition() to hover over the button and verify coordinates manually, or adjust for scaling in code.
  • Alternative tools for specific apps: If you’re automating a desktop GUI app (not just random screen text), libraries like pywinauto can target buttons directly without OCR (more reliable if the app is accessible). For web buttons, selenium is the go-to choice.

Full Working Example

Putting all the pieces together into a reusable function:

import pytesseract
import pyautogui
from PIL import Image, ImageOps

# Set tesseract path if it's not in your system PATH (Windows example)
# pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe'

def find_and_click_button(target_text):
    # Capture screen and preprocess
    screenshot = pyautogui.screenshot()
    img = Image.frombytes("RGB", screenshot.size, screenshot.tobytes())
    img_gray = ImageOps.grayscale(img)
    
    # Extract OCR data
    ocr_data = pytesseract.image_to_data(img_gray, output_type=pytesseract.Output.DICT)
    
    # Search for target text (case-insensitive)
    for i in range(len(ocr_data["text"])):
        cleaned_text = ocr_data["text"][i].strip()
        if cleaned_text.lower() == target_text.lower():
            left = ocr_data["left"][i]
            top = ocr_data["top"][i]
            width = ocr_data["width"][i]
            height = ocr_data["height"][i]
            click_x = left + width // 2
            click_y = top + height // 2
            
            # Optional: Smoothly move mouse to the button first
            pyautogui.moveTo(click_x, click_y, duration=0.5)
            # Click the button
            pyautogui.click()
            print(f"Successfully clicked '{target_text}' at ({click_x}, {click_y})")
            return True
    print(f"Button '{target_text}' not found on screen")
    return False

# Test it out
find_and_click_button("Save Changes")

内容的提问来源于stack exchange,提问作者Tadei Alessandro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 14:03:15