You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的OpenCV提取幻灯片边界框并适配光照与拍摄角度?

解决方案:幻灯片边界框检测与编号提取优化

一、幻灯片边界框检测的预处理优化

你之前套用的车牌识别流程不适合幻灯片的大矩形特征,调整如下步骤可提升边界框检测效果:

  • 替换单一方向的Sobel边缘检测为Canny边缘检测,能更全面捕捉幻灯片的四个边框
  • 加入形态学闭运算,连接断裂边缘并消除小噪点
  • 通过轮廓的面积、顶点数量筛选出幻灯片区域

修改后的预处理代码:

import cv2
import numpy as np

image = cv2.imread("extract_num2.jpg")
# 1. 基础预处理:高斯模糊+灰度化
blurred = cv2.GaussianBlur(image, (5,5), 0)
gray = cv2.cvtColor(blurred, cv2.COLOR_BGR2GRAY)

# 2. Canny边缘检测,阈值可根据实际光照微调
canny = cv2.Canny(gray, 50, 150)

# 3. 形态学闭运算,连接离散边缘
kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (7,7))
closed = cv2.morphologyEx(canny, cv2.MORPH_CLOSE, kernel)

# 4. 筛选幻灯片轮廓
contours, _ = cv2.findContours(closed.copy(), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
for cnt in contours:
    # 过滤小面积噪点轮廓
    area = cv2.contourArea(cnt)
    if area < 10000:  # 可根据拍摄距离调整阈值
        continue
    # 多边形逼近,提取矩形顶点
    peri = cv2.arcLength(cnt, True)
    approx = cv2.approxPolyDP(cnt, 0.02 * peri, True)
    # 幻灯片是矩形,逼近后应为4个顶点
    if len(approx) == 4:
        cv2.drawContours(image, [approx], -1, (0,255,0), 3)
        # 提取幻灯片区域(用于后续编号识别)
        x,y,w,h = cv2.boundingRect(approx)
        slide_roi = image[y:y+h, x:x+w]
        break

cv2.imshow("Slide with Bounding Box", image)
cv2.waitKey(0)
cv2.destroyAllWindows()

二、角度、光照、拍摄距离的适配方案

1. 倾斜角度校正:透视变换

如果幻灯片存在倾斜,用透视变换将其校正为正矩形:

# 对四个顶点排序(左上、右上、右下、左下)
def order_points(pts):
    rect = np.zeros((4,2), dtype="float32")
    s = pts.sum(axis=1)
    rect[0] = pts[np.argmin(s)]
    rect[2] = pts[np.argmax(s)]
    diff = np.diff(pts, axis=1)
    rect[1] = pts[np.argmin(diff)]
    rect[3] = pts[np.argmax(diff)]
    return rect

# 计算透视变换矩阵并校正
rect = order_points(approx.reshape(4,2))
(tl, tr, br, bl) = rect
maxWidth = max(int(np.linalg.norm(br-bl)), int(np.linalg.norm(tr-tl)))
maxHeight = max(int(np.linalg.norm(tr-br)), int(np.linalg.norm(tl-bl)))

dst = np.array([
    [0, 0],
    [maxWidth - 1, 0],
    [maxWidth - 1, maxHeight - 1],
    [0, maxHeight - 1]], dtype="float32")

M = cv2.getPerspectiveTransform(rect, dst)
warped = cv2.warpPerspective(image, M, (maxWidth, maxHeight))

2. 光照不均适配

  • 用自适应阈值替代OTSU阈值,应对局部明暗差异:
gray_warped = cv2.cvtColor(warped, cv2.COLOR_BGR2GRAY)
thresh = cv2.adaptiveThreshold(gray_warped, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2)
  • 直方图均衡化提升低光照场景对比度:
clahe = cv2.createCLAHE(clipLimit=2.0, tileGridSize=(8,8))
equalized = clahe.apply(gray_warped)

3. 拍摄距离适配

将轮廓面积阈值设置为图像总像素的比例,避免固定阈值失效:

total_area = image.shape[0] * image.shape[1]
min_contour_area = int(total_area * 0.3)  # 幻灯片至少占图像30%

三、Tesseract编号提取优化

  1. 提取校正后幻灯片的右下角ROI:
h, w = warped.shape[:2]
# 取右下角10%区域,可根据实际编号大小调整比例
num_roi = warped[int(h*0.9):h, int(w*0.9):w]
  1. 针对数字识别优化Tesseract参数:
import pytesseract

# 预处理编号区域
gray_num = cv2.cvtColor(num_roi, cv2.COLOR_BGR2GRAY)
thresh_num = cv2.threshold(gray_num, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)[1]

# 配置Tesseract:只识别数字,针对短文本优化
custom_config = r'--oem 3 --psm 10 -c tessedit_char_whitelist=0123456789'
number = pytesseract.image_to_string(thresh_num, config=custom_config)
print("幻灯片编号:", number.strip())

内容的提问来源于stack exchange,提问作者marssarso

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 15:27:06