You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python或Tesseract OCR识别输入图片中的文本语言及所属书写系统

图片文本书写系统检测方案(基于Python + Tesseract OCR)

检测图片中文本所属书写系统的核心逻辑分为两步:首先通过OCR从图片中提取有效文本信息,再通过预设规则或原生检测能力对文本所属的书写系统做分类,以下是两种可直接落地的实现方案:


方案1:直接调用Tesseract OSD原生脚本检测能力

Tesseract内置了方向与脚本检测(Orientation and Script Detection, OSD)模式,无需额外规则配置即可直接输出识别到的书写系统类型和置信度。

前置依赖

  • 安装Tesseract OCR本体,需勾选全语言支持组件(否则无法识别非拉丁类书写系统)
  • 安装Python依赖包:
    pip install pytesseract pillow
    

实现代码

from PIL import Image
import pytesseract

# Windows用户若未将Tesseract加入环境变量,需手动指定执行路径
# pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe'

def detect_script_with_osd(img_path: str):
    img = Image.open(img_path)
    # 调用OSD模式返回结构化检测结果
    osd_output = pytesseract.image_to_osd(img, output_type=pytesseract.Output.DICT)
    return osd_output['script'], osd_output['script_confidence']

# 测试调用
if __name__ == "__main__":
    script, conf = detect_script_with_osd("test_sample.png")
    print(f"检测到的书写系统:{script}")
    print(f"检测置信度:{conf}")

优缺点说明

  • 优点:调用简单,原生支持拉丁、西里尔、天城文、阿拉伯文等数十种常见书写系统,无需额外配置规则
  • 缺点:当图片文本量过少、模糊或倾斜严重时,检测准确率会出现明显下降,建议提前做图片预处理提升效果

方案2:OCR提取文本后通过Unicode区间匹配判断书写系统

如果原生OSD检测准确率不符合要求,可以先通过Tesseract提取全量文本,再基于不同书写系统对应固定Unicode编码区间的规则,统计字符占比最高的书写系统作为结果,准确率更可控。

实现代码

from PIL import Image
import pytesseract
from collections import defaultdict

# 常用书写系统Unicode区间映射,可根据需求自行扩充覆盖更多书写系统
SCRIPT_RANGES = {
    "拉丁字母": [(0x0000, 0x007F), (0x0080, 0x024F)],
    "西里尔字母": [(0x0400, 0x052F)],
    "天城文": [(0x0900, 0x097F)],
    "阿拉伯字母": [(0x0600, 0x077F)],
    "韩文": [(0xAC00, 0xD7AF), (0x1100, 0x11FF)],
    "日文假名": [(0x3040, 0x30FF)],
    "汉字": [(0x4E00, 0x9FFF), (0x3400, 0x4DBF)]
}

def detect_script_with_unicode(img_path: str):
    img = Image.open(img_path)
    # lang参数按需添加目标语言的Tesseract语言码,覆盖需要识别的书写系统
    raw_text = pytesseract.image_to_string(img, lang="osd+eng+rus+hin+ara+jpn+kor+chi_sim").strip()
    if not raw_text:
        return "未检测到有效文本", 0

    script_count = defaultdict(int)
    total_valid = 0
    for char in raw_text:
        char_code = ord(char)
        for script, ranges in SCRIPT_RANGES.items():
            for start, end in ranges:
                if start <= char_code <= end:
                    script_count[script] += 1
                    total_valid += 1
                    break
            else:
                continue
            break
    
    if total_valid == 0:
        return "未识别到已知书写系统的字符", 0
    # 取字符占比最高的书写系统作为结果
    top_script, top_count = max(script_count.items(), key=lambda x: x[1])
    confidence = round(top_count / total_valid * 100, 2)
    return top_script, confidence

# 测试调用
if __name__ == "__main__":
    script, conf = detect_script_with_unicode("test_sample.png")
    print(f"检测到的书写系统:{script}")
    print(f"置信度:{conf}%")

优缺点说明

  • 优点:检测准确率高,只要OCR能正确提取字符就可以准确分类,规则可灵活调整,支持自定义添加小众书写系统
  • 缺点:需要提前维护书写系统对应的Unicode区间映射表

优化建议

如果输入图片质量较差,可先通过OpenCV对图片做灰度化、二值化、去噪、倾斜校正预处理,能大幅提升两种方案的检测准确率。

内容的提问来源于stack exchange,提问作者Gokul NC

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 08:45:04