牙科病历扫描件网格去除优化及手写数字识别率提升咨询
Hey there! As someone who’s tackled similar document preprocessing challenges, I totally get how frustrating those persistent scan grids can be—especially when you’re trying to extract handwritten digits accurately. Let’s break down your questions and explore some better approaches.
Better Grid Removal Strategies
Your current template subtraction + erosion/dilation works, but the need for constant calibration is a pain. Here are more elegant, robust alternatives:
1. Targeted Morphological Operations
Instead of using generic square kernels, create directional kernels to specifically eliminate horizontal and vertical grid lines without damaging handwritten digits:
# Remove horizontal grid lines horizontal_kernel = np.ones((1, 5), np.uint8) no_horizontal = cv2.morphologyEx(diff, cv2.MORPH_OPEN, horizontal_kernel, iterations=1) # Remove vertical grid lines vertical_kernel = np.ones((5, 1), np.uint8) no_grid = cv2.morphologyEx(no_horizontal, cv2.MORPH_OPEN, vertical_kernel, iterations=1)
Adjust the kernel size (e.g., (1,7) for thicker horizontal lines) based on your scan’s grid thickness. This avoids over-eroding/dilating your digits.
2. Contour Filtering
Grid lines have distinct shape characteristics (long, thin, uniform) that you can filter out using contour analysis:
# After thresholding your diff image contours, _ = cv2.findContours(diff, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) no_grid = np.zeros_like(diff) for cnt in contours: x, y, w, h = cv2.boundingRect(cnt) # Filter out grid lines (adjust ratios based on your grid) aspect_ratio = w / float(h) if (aspect_ratio < 0.05 or aspect_ratio > 20) and cv2.contourArea(cnt) < 500: continue # Skip grid contours cv2.drawContours(no_grid, [cnt], 0, 255, -1)
This keeps only the compact, irregular contours that match handwritten digits.
3. Frequency Domain Filtering (FFT)
Regular grids create sharp peaks in the frequency domain—you can eliminate them with FFT:
import numpy as np import cv2 def remove_grid_fft(img): # Convert to float32 for FFT f = np.fft.fft2(img) fshift = np.fft.fftshift(f) # Create mask to block grid frequency peaks (adjust coordinates based on your spectrum) rows, cols = img.shape crow, ccol = rows//2, cols//2 mask = np.ones((rows, cols), np.uint8) # Block horizontal grid peaks (adjust size as needed) mask[crow-2:crow+2, :] = 0 # Block vertical grid peaks mask[:, ccol-2:ccol+2] = 0 # Apply mask and inverse FFT fshift_filtered = fshift * mask f_ishift = np.fft.ifftshift(fshift_filtered) img_filtered = np.fft.ifft2(f_ishift) img_filtered = np.abs(img_filtered) # Convert back to 8-bit img_filtered = cv2.normalize(img_filtered, None, 0, 255, cv2.NORM_MINMAX, dtype=cv2.CV_8U) return img_filtered
This method is highly stable for regular grids and requires minimal calibration once you set the mask correctly.
Will SURF Improve Feature Matching?
Short answer: Probably yes, but with caveats.
- SURF (Speeded-Up Robust Features) produces more robust descriptors than ORB, especially for blurry or rotated scans—this should reduce the need for frequent recalibration since alignment will be more accurate.
- Note: SURF is part of OpenCV’s non-free
xfeatures2dmodule, so you’ll need to ensure your OpenCV installation includes it (check withcv2.xfeatures2d.SURF_create()). - To swap ORB for SURF in your code:
# Replace ORB with SURF surf = cv2.xfeatures2d.SURF_create(MAX_FEATURES) kp1, des1 = surf.detectAndCompute(img_preprocessed, None) kp2, des2 = surf.detectAndCompute(template_img, None) # Use FLANN matcher instead of brute-force for better performance with SURF FLANN_INDEX_KDTREE = 1 index_params = dict(algorithm=FLANN_INDEX_KDTREE, trees=5) search_params = dict(checks=50) matcher = cv2.FlannBasedMatcher(index_params, search_params) matches = matcher.knnMatch(des1, des2, k=2) # Apply Lowe's ratio test to filter good matches good_matches = [] for m, n in matches: if m.distance < 0.7 * n.distance: good_matches.append(m)
SURF’s descriptors are floating-point, so FLANN is faster than brute-force matching here.
Optimizations for Your Current Code
I noticed a couple of tweaks that could help:
- Remove redundant code: You’re calculating
numGoodMatchestwice—delete the duplicate line. - Fix thresholding: Your current adaptive threshold call is incorrect (you’re using
cv2.thresholdinstead ofcv2.adaptiveThreshold). Replace it with:
diff = cv2.adaptiveThreshold(diff, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY, 11, 2)
- Align preprocessing: Apply Gaussian blur to both your input and template images to ensure consistent feature detection.
Final Recommendation
Start with FFT filtering for grid removal—it’s the most hands-off solution for regular grids. Then swap ORB for SURF to improve alignment accuracy, which should reduce calibration needs. Combine these with the morphological/contour tweaks, and your digit extraction accuracy should jump significantly.
内容的提问来源于stack exchange,提问作者Zahnbrecher

