基于水平线检测的图像定向校正方案优化求助
Hey Ahmad, nice work getting your initial solution to 98% accuracy— that's already really impressive! Let's iron out those last edge cases so your rotation logic works perfectly across your entire dataset. The core issue with your current approach is that it relies on a fixed pixel slice for detecting the white line, which breaks when resizing shifts the line's position, or if the line isn't perfectly aligned in that narrow range. Here are a few more robust strategies to try:
1. Use Hough Line Transform for Reliable Line Detection
Instead of checking a fixed region, use OpenCV's Hough Line Transform to detect all horizontal/vertical lines in the image. This handles variations in the line's position caused by resizing, and it’s more tolerant of minor line imperfections.
Here's how to adapt your code:
import cv2 import numpy as np from PIL import Image def adjust_palette_position(img_path): # Load and resize image img = Image.open(img_path).resize((128, 128), Image.BILINEAR) img_np = np.array(img) gray = cv2.cvtColor(img_np, cv2.COLOR_RGB2GRAY) # Edge detection (tweaked params for better robustness) edges = cv2.Canny(gray, 50, 150, apertureSize=3) # Detect lines using Hough Transform # Adjust rho, theta, threshold based on your dataset's line characteristics lines = cv2.HoughLinesP(edges, rho=1, theta=np.pi/180, threshold=20, minLineLength=80, maxLineGap=5) horizontal_lines = 0 vertical_lines = 0 if lines is not None: for line in lines: x1, y1, x2, y2 = line[0] # Classify line as horizontal (minimal y difference) if abs(y1 - y2) < 5: horizontal_lines += 1 # Classify line as vertical (minimal x difference) elif abs(x1 - x2) < 5: vertical_lines += 1 # Decision logic: # Dominant horizontal lines mean palette is at the bottom → rotate to move it to the right if horizontal_lines > vertical_lines: img_np = np.rot90(img_np) # Dominant vertical lines mean palette is already on the right → no rotation needed return Image.fromarray(img_np)
This approach finds lines across the entire image, so it doesn’t care if the separator line shifts a few pixels during resizing.
2. Fallback: Color Distribution Analysis
If some images lack a clear separator line, use color variance to identify the palette. Palettes are made of solid, uniform color blocks, so they’ll have much lower color variance than the main image content (which tends to have more color variation).
Add this check to handle edge cases where line detection fails:
def get_color_variance(region): # Calculate average variance across all color channels return np.mean(np.var(region, axis=(0,1))) # After line detection, if no clear line signal is found: if horizontal_lines == 0 and vertical_lines == 0: # Check bottom 20% and right 20% of the resized image bottom_region = img_np[-25:, :] # ~20% of 128px height right_region = img_np[:, -25:] # ~20% of 128px width bottom_variance = get_color_variance(bottom_region) right_variance = get_color_variance(right_region) # Lower variance means a more uniform region (aka the palette) if bottom_variance < right_variance: # Palette is at the bottom → rotate to move it to the right img_np = np.rot90(img_np)
3. Combine Both Methods for Maximum Reliability
For the best results, use line detection first, then fall back to color variance if line detection doesn’t find a clear signal. This covers both cases where the separator line exists and where it’s missing or blurry.
A quick tip: Test these methods on your failed cases first— adjust the Hough line parameters (like minLineLength or threshold) and variance region size to match your dataset’s specific characteristics.
内容的提问来源于stack exchange,提问作者Ahmad Moussa

