PHP中Tesseract命令失效问题及替代工具咨询(含PDF图片文本提取)
替代Tesseract的PHP方案:提取PDF内嵌图片文本
核心思路
先从PDF中提取内嵌图片,再对图片执行OCR识别文本,避开Tesseract命令行调用问题。
步骤1:提取PDF中的内嵌图片
使用Ghostscript命令行工具完成图片提取,PHP通过exec()调用:
Ghostscript命令(终端直接执行)
gs -dNOPAUSE -dBATCH -sDEVICE=jpeg -r300 -sOutputFile=extracted_img_%d.jpg input.pdf
参数说明:
-r300:设置输出图片分辨率300DPI(OCR识别最佳分辨率)extracted_img_%d.jpg:按页码生成图片文件(如extracted_img_1.jpg)
PHP调用示例
$pdfPath = '/path/to/target.pdf'; $outputDir = '/path/to/save/images/'; // 确保输出目录存在 if (!is_dir($outputDir)) mkdir($outputDir, 0755, true); $command = sprintf( 'gs -dNOPAUSE -dBATCH -sDEVICE=jpeg -r300 -sOutputFile=%sextracted_img_%%d.jpg %s', escapeshellarg($outputDir), escapeshellarg($pdfPath) ); exec($command, $output, $returnCode); if ($returnCode !== 0) { die('Ghostscript执行失败:'.implode(PHP_EOL, $output)); }
步骤2:OCR识别图片文本(替代Tesseract的方案)
方案A:调用Google Cloud Vision API(云端OCR)
安装PHP客户端:
composer require google/cloud-vision
编写识别代码:
require 'vendor/autoload.php'; use Google\Cloud\Vision\V1\ImageAnnotatorClient; // 初始化客户端(需配置Google Cloud服务账号密钥) putenv('GOOGLE_APPLICATION_CREDENTIALS=/path/to/service-account-key.json'); $imageAnnotator = new ImageAnnotatorClient(); // 遍历提取的图片 $imageFiles = glob($outputDir.'*.jpg'); foreach ($imageFiles as $file) { $imageContent = file_get_contents($file); $response = $imageAnnotator->textDetection($imageContent); $textAnnotations = $response->getTextAnnotations(); if (!empty($textAnnotations)) { // 第一个结果是完整文本,后续是单字/词组 echo "页面 ".pathinfo($file, PATHINFO_FILENAME)." 提取文本:\n"; echo $textAnnotations[0]->getDescription()."\n\n"; } } $imageAnnotator->close();
方案B:Python EasyOCR + PHP调用(本地离线OCR)
适合无法访问云端服务的场景,需服务器安装Python及EasyOCR:
- 安装EasyOCR:
pip install easyocr
- 编写Python脚本
ocr_process.py:
import sys import easyocr # 初始化识别器,指定语言(如中文+英文) reader = easyocr.Reader(['ch_sim', 'en']) image_path = sys.argv[1] # 执行识别,返回文本列表 result = reader.readtext(image_path, detail=0) # 输出拼接后的文本 print(' '.join(result))
- PHP调用脚本:
$imageFile = '/path/to/extracted_img_1.jpg'; $command = sprintf( 'python3 /path/to/ocr_process.py %s', escapeshellarg($imageFile) ); $ocrResult = shell_exec($command); if ($ocrResult) { echo "提取文本:\n".$ocrResult; } else { echo "OCR识别失败"; }
方案C:本地部署OCR服务+HTTP请求
部署离线OCR服务(如百度OCR离线版、PaddleOCR服务),PHP通过HTTP请求调用:
$imageFile = '/path/to/extracted_img_1.jpg'; $base64Image = base64_encode(file_get_contents($imageFile)); $curl = curl_init(); curl_setopt_array($curl, [ CURLOPT_URL => 'http://127.0.0.1:8080/api/ocr', // OCR服务地址 CURLOPT_POST => true, CURLOPT_POSTFIELDS => json_encode(['image' => $base64Image]), CURLOPT_HTTPHEADER => ['Content-Type: application/json'], CURLOPT_RETURNTRANSFER => true ]); $response = curl_exec($curl); curl_close($curl); $result = json_decode($response, true); if ($result && isset($result['text'])) { echo "提取文本:\n".$result['text']; } else { echo "OCR服务调用失败"; }
内容的提问来源于stack exchange,提问作者PentaD
相关产品推荐
相关产品推荐

