基于pytesseract优化OCR提取文本的对齐问题
问题描述
提取图像文本时出现文本对齐错乱的情况,推测问题出在格式处理环节而非文本提取环节。想请教是否可以利用边界框坐标来优化文本对齐效果?原图像中价格列的起始位置是统一的,但提取后的输出文本存在对齐错乱问题。
当前使用的代码如下:
def extract_text(self, image): gray_image = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY) text_regions = [] retval, labels, stats, centroids = cv2.connectedComponentsWithStats(gray_image) section_groups = {} section_threshold = 20 for i in range(1, retval): x, y, w, h, area = stats[i] section_groups.setdefault(y // section_threshold, []).append((x, y, w, h)) for section_key, components in section_groups.items(): section_image = gray_image[components[0][1]:components[-1][1] + components[-1][3], :] extracted_text, section_boxes = self._extract_text_and_boxes(section_image) text_regions.append((extracted_text, section_boxes)) return text_regions def _extract_text_and_boxes(self, component): extracted_text = pytesseract.image_to_string(component, config=self.config) boxes = self.get_text_boxes(component) return extracted_text, boxes def get_text_boxes(self, image): results = pytesseract.image_to_data(image, output_type=Output.DICT) boxes = [] for i in range(0, len(results["text"])): x = results["left"][i] y = results["top"][i] w = results["width"][i] h = results["height"][i] text = results["text"][i] conf = float(results["conf"][i]) if conf > 0: boxes.append((x, y, w, h)) return boxes
解决方案
完全可以通过边界框坐标优化文本对齐,核心思路是基于每个文本块的x坐标(横向位置)对文本进行分组或对齐排版,而非直接使用Tesseract返回的无格式纯文本。
现有代码的问题
- 仅按行分割区域提取文本,未利用边界框位置信息排版,Tesseract的
image_to_string输出是无格式纯文本,自然丢失原列对齐结构。 - 连通分量分组逻辑存在缩进错误:
for循环内的append语句被放在循环外,导致仅最后一个连通分量被加入分组,行分割不准确。
优化步骤
1. 修复连通分量分组的缩进错误
修正extract_text方法的循环缩进,确保每个连通分量正确分组,并按行的y坐标排序保证处理顺序:
def extract_text(self, image): gray_image = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY) text_regions = [] retval, labels, stats, centroids = cv2.connectedComponentsWithStats(gray_image) section_groups = {} section_threshold = 20 # 修复缩进:将append放入循环内部 for i in range(1, retval): x, y, w, h, area = stats[i] section_groups.setdefault(y // section_threshold, []).append((x, y, w, h)) # 按行的y坐标排序,保证从上到下处理 sorted_sections = sorted(section_groups.items(), key=lambda item: item[0]) for section_key, components in sorted_sections: # 获取当前行的上下边界 min_y = min(comp[1] for comp in components) max_y = max(comp[1] + comp[3] for comp in components) section_image = gray_image[min_y:max_y, :] # 提取当前行的文本块与对应边界框 line_text_boxes = self._extract_text_with_boxes(section_image) text_regions.append(line_text_boxes) return text_regions
2. 修改文本提取方法,保留文本与边界框的对应关系
替换原_extract_text_and_boxes方法,同时返回每个文本块的内容和位置,并按x坐标排序保证行内顺序:
def _extract_text_with_boxes(self, component): results = pytesseract.image_to_data(component, output_type=Output.DICT) text_boxes = [] for i in range(0, len(results["text"])): x = results["left"][i] y = results["top"][i] w = results["width"][i] h = results["height"][i] text = results["text"][i].strip() conf = float(results["conf"][i]) # 过滤空文本和低置信度结果 if conf > 0 and text: text_boxes.append((x, y, w, h, text)) # 按x坐标排序,确保同一行内从左到右的顺序 text_boxes.sort(key=lambda item: item[0]) return text_boxes
3. 基于边界框实现列对齐排版
通过聚类所有文本块的x坐标确定列基准位置,再对每行文本按基准位置填充空格实现对齐:
def format_aligned_text(self, text_regions): # 收集所有文本块的x坐标用于确定列基准 all_x_coords = [] for line_boxes in text_regions: for box in line_boxes: all_x_coords.append(box[0]) if not all_x_coords: return "" # 对x坐标聚类,生成列基准位置(相邻x差小于10归为同一列) column_threshold = 10 all_x_coords.sort() columns = [all_x_coords[0]] for x in all_x_coords[1:]: if x - columns[-1] > column_threshold: columns.append(x) # 按列基准对齐每行文本 aligned_lines = [] char_width = 5 # 假设每个字符占5像素宽度,可根据实际调整 for line_boxes in text_regions: line_text = "" prev_col_end = 0 for col_x in columns: # 匹配当前列的文本块 col_text = "" for box in line_boxes: if abs(box[0] - col_x) <= column_threshold: col_text = box[4] break # 计算需要填充的空格数 space_needed = max(0, (col_x - prev_col_end) // char_width) line_text += " " * space_needed + col_text prev_col_end = col_x + len(col_text) * char_width aligned_lines.append(line_text) return "\n".join(aligned_lines)
使用示例
# 提取带边界框的文本区域 text_regions = your_instance.extract_text(your_image) # 生成对齐后的文本 aligned_text = your_instance.format_aligned_text(text_regions) print(aligned_text)
内容的提问来源于stack exchange,提问作者code_comm
相关产品推荐
相关产品推荐

