You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PHP中Tesseract命令失效问题及替代工具咨询(含PDF图片文本提取)

替代Tesseract的PHP方案:提取PDF内嵌图片文本

核心思路

先从PDF中提取内嵌图片,再对图片执行OCR识别文本,避开Tesseract命令行调用问题。


步骤1:提取PDF中的内嵌图片

使用Ghostscript命令行工具完成图片提取,PHP通过exec()调用:

Ghostscript命令(终端直接执行)

gs -dNOPAUSE -dBATCH -sDEVICE=jpeg -r300 -sOutputFile=extracted_img_%d.jpg input.pdf

参数说明:

  • -r300:设置输出图片分辨率300DPI(OCR识别最佳分辨率)
  • extracted_img_%d.jpg:按页码生成图片文件(如extracted_img_1.jpg)

PHP调用示例

$pdfPath = '/path/to/target.pdf';
$outputDir = '/path/to/save/images/';
// 确保输出目录存在
if (!is_dir($outputDir)) mkdir($outputDir, 0755, true);

$command = sprintf(
    'gs -dNOPAUSE -dBATCH -sDEVICE=jpeg -r300 -sOutputFile=%sextracted_img_%%d.jpg %s',
    escapeshellarg($outputDir),
    escapeshellarg($pdfPath)
);

exec($command, $output, $returnCode);
if ($returnCode !== 0) {
    die('Ghostscript执行失败:'.implode(PHP_EOL, $output));
}

步骤2:OCR识别图片文本(替代Tesseract的方案)

方案A:调用Google Cloud Vision API(云端OCR)

安装PHP客户端:

composer require google/cloud-vision

编写识别代码:

require 'vendor/autoload.php';

use Google\Cloud\Vision\V1\ImageAnnotatorClient;

// 初始化客户端(需配置Google Cloud服务账号密钥)
putenv('GOOGLE_APPLICATION_CREDENTIALS=/path/to/service-account-key.json');
$imageAnnotator = new ImageAnnotatorClient();

// 遍历提取的图片
$imageFiles = glob($outputDir.'*.jpg');
foreach ($imageFiles as $file) {
    $imageContent = file_get_contents($file);
    $response = $imageAnnotator->textDetection($imageContent);
    $textAnnotations = $response->getTextAnnotations();

    if (!empty($textAnnotations)) {
        // 第一个结果是完整文本,后续是单字/词组
        echo "页面 ".pathinfo($file, PATHINFO_FILENAME)." 提取文本:\n";
        echo $textAnnotations[0]->getDescription()."\n\n";
    }
}

$imageAnnotator->close();

方案B:Python EasyOCR + PHP调用(本地离线OCR)

适合无法访问云端服务的场景,需服务器安装Python及EasyOCR:

  1. 安装EasyOCR:
pip install easyocr
  1. 编写Python脚本ocr_process.py:
import sys
import easyocr

# 初始化识别器,指定语言(如中文+英文)
reader = easyocr.Reader(['ch_sim', 'en'])
image_path = sys.argv[1]

# 执行识别,返回文本列表
result = reader.readtext(image_path, detail=0)
# 输出拼接后的文本
print(' '.join(result))
  1. PHP调用脚本:
$imageFile = '/path/to/extracted_img_1.jpg';
$command = sprintf(
    'python3 /path/to/ocr_process.py %s',
    escapeshellarg($imageFile)
);

$ocrResult = shell_exec($command);
if ($ocrResult) {
    echo "提取文本:\n".$ocrResult;
} else {
    echo "OCR识别失败";
}

方案C:本地部署OCR服务+HTTP请求

部署离线OCR服务(如百度OCR离线版、PaddleOCR服务),PHP通过HTTP请求调用:

$imageFile = '/path/to/extracted_img_1.jpg';
$base64Image = base64_encode(file_get_contents($imageFile));

$curl = curl_init();
curl_setopt_array($curl, [
    CURLOPT_URL => 'http://127.0.0.1:8080/api/ocr', // OCR服务地址
    CURLOPT_POST => true,
    CURLOPT_POSTFIELDS => json_encode(['image' => $base64Image]),
    CURLOPT_HTTPHEADER => ['Content-Type: application/json'],
    CURLOPT_RETURNTRANSFER => true
]);

$response = curl_exec($curl);
curl_close($curl);

$result = json_decode($response, true);
if ($result && isset($result['text'])) {
    echo "提取文本:\n".$result['text'];
} else {
    echo "OCR服务调用失败";
}

内容的提问来源于stack exchange,提问作者PentaD

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 19:16:35