You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Tesseract OCR+Tika Core添加Latin脚本语言报错求助

解决Tika TesseractOCRConfig添加Latin脚本语言报错的问题

问题原因

Tika的TesseractOCRConfig类的setLanguage方法内部有语言代码校验逻辑,它默认只识别不带路径前缀的普通语言代码(如eng),不支持script/Latin这种带脚本路径的格式,因此会抛出IllegalArgumentException: Invalid language code。而命令行调用Tesseract时没有这个校验限制,所以可以正常使用-l script/Latin。

解决方案:绕过Tika的语言校验,直接传递Tesseract命令行参数

不要使用setLanguage方法,而是通过addTesseractArgs方法直接添加Tesseract的原生命令行参数,这样就能绕过Tika的语言代码校验。

代码示例

import org.apache.tika.parser.ocr.TesseractOCRConfig;
import org.apache.tika.parser.AutoDetectParser;
import org.apache.tika.metadata.Metadata;
import org.apache.tika.io.TikaInputStream;
import java.io.File;

public class TikaOcrExample {
    public static void main(String[] args) throws Exception {
        // 初始化TesseractOCRConfig
        TesseractOCRConfig config = new TesseractOCRConfig();
        
        // 设置tessdata路径
        config.setTessdataPath("E:\\Program Files\\tesseractOCR\\tessdata");
        
        // 直接添加Tesseract命令行参数,指定Latin脚本
        config.addTesseractArgs("-l", "script/Latin");
        
        // 如果需要同时识别英文和Latin脚本,参数改为:
        // config.addTesseractArgs("-l", "eng+script/Latin");
        
        // 配置Parser并执行OCR
        AutoDetectParser parser = new AutoDetectParser();
        parser.setOcrConfig(config);
        
        File file = new File("your-input-file.pdf"); // 替换为你的文件路径
        Metadata metadata = new Metadata();
        try (TikaInputStream stream = TikaInputStream.get(file)) {
            parser.parse(stream, null, metadata);
            String content = metadata.get("X-TIKA:OCR_CONTENT");
            System.out.println(content);
        }
    }
}

验证方法

运行代码后,检查是否能正常识别使用Latin脚本的内容,同时不会抛出语言代码无效的异常。

内容的提问来源于stack exchange,提问作者Kipoti

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 23:24:31