寻求通用算法:检测图像手写内容及视频幻灯片手写填充占比
问题描述
请问是否存在可检测图像中是否包含手写内容的算法?我无需识别手写文字内容,仅需判断是否存在手写内容。我拥有一段手写填充幻灯片的视频,目标是确定幻灯片已被手写内容填充的比例。此前在《Quantify how much a slide has been filled with handwriting》中已有基于手写特定颜色占比求和的解决方案,但如果手写颜色并非蓝色,且该颜色也存在于非手写区域时,该方案将失效。因此我想了解是否存在更通用的图像手写内容检测方案。我目前的思路:考虑提取图像轮廓,再基于轮廓的弯曲程度检测手写部分,但不知如何实现,且该思路可能并非最优方案,准确率无法保证。附上已编写的代码:
import cv2 import matplotlib.pyplot as plt img = cv2.imread(PATH TO IMAGE) print("img shape=", img.shape) gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) cv2.imshow("image", gray) cv2.waitKey(1) #### extract all contours # Find Canny edges edged = cv2.Canny(gray, 30, 200) cv2.waitKey(0) # Finding Contours # Use a copy of the image e.g. edged.copy() # since findContours alters the image contours, hierarchy = cv2.findContours(edged, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_NONE) cv2.imshow('Canny Edges After Contouring', edged) cv2.waitKey(0) print("Number of Contours found = " + str(len(contours))) # Draw all contours # -1 signifies drawing all contours cv2.drawContours(img, contours, -1, (0, 255, 0), 3) cv2.imshow('Contours', img) cv2.waitKey(0)
解决方案
你的轮廓思路方向是对的,但单纯提取所有轮廓太宽泛——幻灯片里的印刷元素(比如边框、图标、文字)也会产生轮廓,得想办法区分手写和印刷的特征。以下是几个更通用的方案,结合你的需求做了优化:
1. 改进版轮廓检测(快速实现)
在你现有代码的基础上,通过轮廓形状特征过滤掉非手写区域:手写线条通常细长、弯曲度高,而印刷体轮廓更规整。我们可以计算轮廓的长宽比、弯曲度来筛选手写区域。
优化后的代码:
import cv2 import numpy as np def calculate_contour_features(contour): # 计算外接矩形长宽比 x, y, w, h = cv2.boundingRect(contour) aspect_ratio = float(w) / h if h != 0 else 0 # 计算弯曲度:周长平方/(4π*面积),值越大线条越不规则 area = cv2.contourArea(contour) perimeter = cv2.arcLength(contour, True) circularity = (perimeter ** 2) / (4 * np.pi * area) if area != 0 else 0 return aspect_ratio, circularity def detect_handwriting_ratio(img): gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) # 先做高斯模糊减少噪声干扰 blur = cv2.GaussianBlur(gray, (5, 5), 0) edged = cv2.Canny(blur, 50, 150) contours, _ = cv2.findContours(edged, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) handwriting_area = 0 total_slide_area = img.shape[0] * img.shape[1] for cnt in contours: area = cv2.contourArea(cnt) # 过滤过小的噪声轮廓 if area < 20: continue aspect_ratio, circularity = calculate_contour_features(cnt) # 筛选细长且弯曲的轮廓(可根据实际情况调整阈值) if (aspect_ratio > 2 or aspect_ratio < 0.5) and circularity > 1.5: handwriting_area += area return handwriting_area / total_slide_area # 使用示例 img = cv2.imread("your_slide_frame.jpg") fill_ratio = detect_handwriting_ratio(img) print(f"手写填充比例: {fill_ratio:.2%}")
2. 基于纹理特征的分类器(通用场景)
手写内容的纹理和印刷体/空白区域差异极大:手写线条粗细不均、边缘不规则,而印刷体纹理规整。可以用**LBP(局部二值模式)**提取纹理特征,搭配SVM分类器区分手写区域,适合颜色多变、背景复杂的场景。
核心代码示例:
import cv2 import numpy as np from sklearn.svm import SVC from sklearn.model_selection import train_test_split # 提取LBP纹理特征 def get_lbp_features(img): gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) if len(img.shape) == 3 else img lbp = cv2.LBP_create(8, 1) # 8个采样点,半径1 lbp_img = lbp.compute(gray) # 计算特征直方图并归一化 hist, _ = np.histogram(lbp_img, bins=np.arange(0, 256)) hist = hist.astype(np.float32) cv2.normalize(hist, hist) return hist # 假设你已经有标注好的手写/非手写样本集 # handwriting_samples = [手写区域图像列表] # non_handwriting_samples = [非手写区域图像列表] X = [] y = [] for img in handwriting_samples: X.append(get_lbp_features(img)) y.append(1) for img in non_handwriting_samples: X.append(get_lbp_features(img)) y.append(0) # 训练SVM分类器 X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2) svm = SVC(kernel='linear') svm.fit(X_train, y_train) # 对单帧图像分窗口检测 def calculate_fill_ratio(img, svm, window_size=(20,20)): h, w = img.shape[:2] step_h, step_w = window_size total_windows = 0 handwriting_windows = 0 for y in range(0, h-step_h, step_h): for x in range(0, w-step_w, step_w): window = img[y:y+step_h, x:x+step_w] feat = get_lbp_features(window) if svm.predict([feat])[0] == 1: handwriting_windows +=1 total_windows +=1 return handwriting_windows / total_windows if total_windows >0 else 0
3. 轻量级深度学习方案(鲁棒性最强)
如果有足够的标注样本,用小体积CNN做语义分割是最可靠的方案——直接分割出手写区域,计算占比即可。比如用TensorFlow搭建一个简单的U-Net模型,输入幻灯片图像,输出二值掩码(1表示手写区域)。
方案选择建议
- 场景简单(背景干净、手写与印刷差异明显):优先用改进版轮廓检测,快速落地
- 场景复杂(手写颜色多变、背景有干扰元素):选择纹理特征+分类器方案
- 追求最高鲁棒性:尝试轻量级深度学习分割模型
内容的提问来源于stack exchange,提问作者henry

