如何用Python(Selenium)识别该数字验证码并提取纯文本?
Python + Selenium 识别数字验证码实操方案
1. 提取页面中的验证码图片
用Selenium定位验证码元素,通过JS将其转为base64编码,再转换为PIL图片方便后续处理:
from selenium import webdriver from selenium.webdriver.common.by import By import base64 import io from PIL import Image # 初始化浏览器并打开目标页面 driver = webdriver.Chrome() driver.get("你的目标页面地址") # 定位验证码图片元素(根据实际页面结构调整定位规则) captcha_img = driver.find_element(By.XPATH, "//img[contains(@src, 'captcha')]") # 通过JS获取图片的base64编码(避免直接截图的尺寸偏差问题) captcha_base64 = driver.execute_script(""" const canvas = document.createElement('canvas'); const ctx = canvas.getContext('2d'); const img = arguments[0]; canvas.width = img.naturalWidth; canvas.height = img.naturalHeight; ctx.drawImage(img, 0, 0); return canvas.toDataURL('image/png').split(',')[1]; """, captcha_img) # 转换为PIL Image对象 img_data = base64.b64decode(captcha_base64) img = Image.open(io.BytesIO(img_data))
2. 预处理图片(降低干扰,提升识别准确率)
数字验证码通常带有干扰线或噪点,先做灰度化和二值化处理:
# 转换为灰度图 gray_img = img.convert('L') # 二值化处理(阈值可根据验证码实际深浅调整,示例用127) binary_img = gray_img.point(lambda x: 0 if x < 127 else 255, '1')
3. OCR识别数字
使用Tesseract进行识别,先安装依赖:
pip install pytesseract
注意:需在系统中安装Tesseract本体,Windows用户若未配置环境变量,可在代码中手动指定路径
识别代码,限制仅识别0-9数字:
import pytesseract # 若未配置环境变量,手动指定Tesseract路径 # pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe' # 配置OCR参数:仅识别数字 ocr_config = r'--oem 3 --psm 8 -c tessedit_char_whitelist=0123456789' captcha_num = pytesseract.image_to_string(binary_img, config=ocr_config).strip() print(f"识别出的验证码:{captcha_num}")
4. 将识别结果输入页面
# 定位验证码输入框(根据实际页面结构调整) captcha_input = driver.find_element(By.ID, "captcha-input") captcha_input.send_keys(captcha_num)
额外提示
- 若验证码干扰较强,可尝试添加噪点去除、腐蚀膨胀等进阶图像处理步骤
- 若Tesseract识别效果不佳,可考虑训练自定义数字识别模型,或使用专业验证码识别服务
内容的提问来源于stack exchange,提问作者pro221
相关产品推荐
相关产品推荐

