如何使用Python或Tesseract OCR识别输入图片中的文本语言及所属书写系统
图片文本书写系统检测方案(基于Python + Tesseract OCR)
检测图片中文本所属书写系统的核心逻辑分为两步:首先通过OCR从图片中提取有效文本信息,再通过预设规则或原生检测能力对文本所属的书写系统做分类,以下是两种可直接落地的实现方案:
方案1:直接调用Tesseract OSD原生脚本检测能力
Tesseract内置了方向与脚本检测(Orientation and Script Detection, OSD)模式,无需额外规则配置即可直接输出识别到的书写系统类型和置信度。
前置依赖
- 安装Tesseract OCR本体,需勾选全语言支持组件(否则无法识别非拉丁类书写系统)
- 安装Python依赖包:
pip install pytesseract pillow
实现代码
from PIL import Image import pytesseract # Windows用户若未将Tesseract加入环境变量,需手动指定执行路径 # pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe' def detect_script_with_osd(img_path: str): img = Image.open(img_path) # 调用OSD模式返回结构化检测结果 osd_output = pytesseract.image_to_osd(img, output_type=pytesseract.Output.DICT) return osd_output['script'], osd_output['script_confidence'] # 测试调用 if __name__ == "__main__": script, conf = detect_script_with_osd("test_sample.png") print(f"检测到的书写系统:{script}") print(f"检测置信度:{conf}")
优缺点说明
- 优点:调用简单,原生支持拉丁、西里尔、天城文、阿拉伯文等数十种常见书写系统,无需额外配置规则
- 缺点:当图片文本量过少、模糊或倾斜严重时,检测准确率会出现明显下降,建议提前做图片预处理提升效果
方案2:OCR提取文本后通过Unicode区间匹配判断书写系统
如果原生OSD检测准确率不符合要求,可以先通过Tesseract提取全量文本,再基于不同书写系统对应固定Unicode编码区间的规则,统计字符占比最高的书写系统作为结果,准确率更可控。
实现代码
from PIL import Image import pytesseract from collections import defaultdict # 常用书写系统Unicode区间映射,可根据需求自行扩充覆盖更多书写系统 SCRIPT_RANGES = { "拉丁字母": [(0x0000, 0x007F), (0x0080, 0x024F)], "西里尔字母": [(0x0400, 0x052F)], "天城文": [(0x0900, 0x097F)], "阿拉伯字母": [(0x0600, 0x077F)], "韩文": [(0xAC00, 0xD7AF), (0x1100, 0x11FF)], "日文假名": [(0x3040, 0x30FF)], "汉字": [(0x4E00, 0x9FFF), (0x3400, 0x4DBF)] } def detect_script_with_unicode(img_path: str): img = Image.open(img_path) # lang参数按需添加目标语言的Tesseract语言码,覆盖需要识别的书写系统 raw_text = pytesseract.image_to_string(img, lang="osd+eng+rus+hin+ara+jpn+kor+chi_sim").strip() if not raw_text: return "未检测到有效文本", 0 script_count = defaultdict(int) total_valid = 0 for char in raw_text: char_code = ord(char) for script, ranges in SCRIPT_RANGES.items(): for start, end in ranges: if start <= char_code <= end: script_count[script] += 1 total_valid += 1 break else: continue break if total_valid == 0: return "未识别到已知书写系统的字符", 0 # 取字符占比最高的书写系统作为结果 top_script, top_count = max(script_count.items(), key=lambda x: x[1]) confidence = round(top_count / total_valid * 100, 2) return top_script, confidence # 测试调用 if __name__ == "__main__": script, conf = detect_script_with_unicode("test_sample.png") print(f"检测到的书写系统:{script}") print(f"置信度:{conf}%")
优缺点说明
- 优点:检测准确率高,只要OCR能正确提取字符就可以准确分类,规则可灵活调整,支持自定义添加小众书写系统
- 缺点:需要提前维护书写系统对应的Unicode区间映射表
优化建议
如果输入图片质量较差,可先通过OpenCV对图片做灰度化、二值化、去噪、倾斜校正预处理,能大幅提升两种方案的检测准确率。
内容的提问来源于stack exchange,提问作者Gokul NC
相关产品推荐
相关产品推荐

