如何在Python中存储Tesseract识别结果里指定字符的首尾坐标?
问题与解决方案
需求
从Tesseract识别出的字符坐标数据中,提取CNIC里**第一个'C'和最后一个'C'**的坐标信息。
原始数据
myboxes = 'C 0 0 19 42 0\nN 23 8 40 26 0\nI 43 9 46 25 0\nC 49 9 68 27 0\n/ 71 8 81 28 0\nN 84 11 100 30 0\nT 103 13 113 31 0\nN 116 13 132 32 0\n/ 136 12 146 33 0\nP 149 15 161 33 0\na 164 16 181 30 0\ns 183 16 192 31 0\ns 194 16 202 31 0\np 205 7 222 32 0\no 224 18 240 32 0\nr 242 18 249 33 0\nt 251 19 256 38 0\n'
之前的代码及输出
你用来处理数据的代码:
r[3] = 'CNIC/' temp = [] for i, b in enumerate(myboxes.splitlines()): temp.append(b.split(' ')) print(b.split(' ')) print(r[3]+"_"+str(b[0])+'-'+str(i)) if temp[i][0] == '/': break
代码输出:
['C', '0', '0', '19', '42', '0'] CNIC_C-0 ['N', '23', '8', '40', '26', '0'] CNIC_N-1 ['I', '43', '9', '46', '25', '0'] CNIC_I-2 ['C', '49', '9', '68', '27', '0'] CNIC_C-3 ['/', '71', '8', '81', '28', '0'] CNIC_/-4
你之前遍历temp提取坐标时只拿到最后一个字符的信息,是因为每次循环都覆盖了char, x, y, width, height变量,没有专门保存目标C的坐标。
解决代码
直接筛选出所有字符为C的条目,再取首尾即可:
# 补全r的定义避免报错,处理数据生成temp列表 r = ['', '', '', 'CNIC/'] temp = [] for i, b in enumerate(myboxes.splitlines()): temp.append(b.split(' ')) if temp[i][0] == '/': break # 收集所有字符为'C'的条目 c_chars = [item for item in temp if item[0] == 'C'] # 提取第一个和最后一个C的坐标 first_c = c_chars[0] last_c = c_chars[-1] # 打印结果 print("第一个C的坐标:") print(f"字符: {first_c[0]}, x: {first_c[1]}, y: {first_c[2]}, width: {first_c[3]}, height: {first_c[4]}") print("最后一个C的坐标:") print(f"字符: {last_c[0]}, x: {last_c[1]}, y: {last_c[2]}, width: {last_c[3]}, height: {last_c[4]}")
输出结果
第一个C的坐标: 字符: C, x: 0, y: 0, width: 19, height: 42 最后一个C的坐标: 字符: C, x: 49, y: 9, width: 68, height: 27
内容的提问来源于stack exchange,提问作者Jawad Mansoor
相关产品推荐
相关产品推荐

