如何从PDF页面指定区域提取图片?PHP实现求助
从PDF指定区域提取图片的PHP解决方案
你遇到的核心问题是Ghostscript的CropBox仅隐藏内容而非真正移除,导致pdfimages仍会提取页面所有图片。以下是几个可程序化实现的解决方案:
方案1:让Ghostscript真正生成裁剪后的PDF
通过添加参数让Ghostscript在输出新PDF时仅保留CropBox内的内容,而非仅隐藏。修改后的Ghostscript命令结合PHP执行:
$originalPdf = 'path/to/your/original.pdf'; $croppedPdf = 'path/to/save/cropped.pdf'; $targetPage = 12; // 裁剪区域:[左 下 右 上],对应PDF坐标 $cropBox = "[31.46 690.22 560.54 839]"; // 构建Ghostshell命令,确保参数转义安全 $gsCmd = sprintf( 'gs -o %s -sDEVICE=pdfwrite -dUseCropBox -c "[/CropBox %s /PAGES pdfmark" -sPageList=%d -dPDFSETTINGS=/prepress -f %s', escapeshellarg($croppedPdf), $cropBox, $targetPage, escapeshellarg($originalPdf) ); // 执行命令并检查结果 exec($gsCmd, $outputLogs, $exitCode); if ($exitCode !== 0) { die("PDF裁剪失败:" . implode("\n", $outputLogs)); } // 用pdfimages提取裁剪后PDF内的图片 $outputDir = 'path/to/image/output'; mkdir($outputDir, 0777, true); $imgCmd = sprintf( 'pdfimages -j %s %s/extracted_image', escapeshellarg($croppedPdf), escapeshellarg($outputDir) ); exec($imgCmd, $imgLogs, $imgExitCode); if ($imgExitCode !== 0) { die("图片提取失败:" . implode("\n", $imgLogs)); }
原理说明
-dUseCropBox参数让Ghostscript在渲染PDF时严格遵循裁剪框,生成的新PDF仅包含裁剪区域内的内容;-dPDFSETTINGS=/prepress确保输出质量,避免图片失真。
方案2:纯PHP判断图片坐标后提取
借助PDF处理库直接获取页面内图片的位置信息,判断是否在目标区域后提取,无需依赖外部命令。
步骤1:安装依赖
composer require setasign/fpdi setasign/fpdf
步骤2:编写提取代码
require_once 'vendor/autoload.php'; use setasign\Fpdi\Fpdi; $pdfPath = 'path/to/your/original.pdf'; $targetPage = 12; // 目标区域的PDF坐标(左、下、右、上) $targetArea = [ 'left' => 31.46, 'bottom' => 690.22, 'right' => 560.54, 'top' => 839 ]; $outputDir = 'path/to/save/images'; mkdir($outputDir, 0777, true); $pdf = new Fpdi(); $pdf->setSourceFile($pdfPath); $pageTemplate = $pdf->importPage($targetPage); $pageSize = $pdf->getTemplateSize($pageTemplate); // 获取当前页面所有图片 $pageImages = $pdf->getImages($targetPage); foreach ($pageImages as $index => $imageInfo) { // 获取图片的边界框坐标(左下、右上) $imgBbox = $imageInfo['bbox']; $imgLeft = $imgBbox[0]; $imgBottom = $imgBbox[1]; $imgRight = $imgBbox[2]; $imgTop = $imgBbox[3]; // 判断图片是否完全在目标区域内(可根据需求调整为部分重叠判断) $isInTarget = ( $imgLeft >= $targetArea['left'] && $imgBottom >= $targetArea['bottom'] && $imgRight <= $targetArea['right'] && $imgTop <= $targetArea['top'] ); if ($isInTarget) { // 提取图片数据并保存 $imgData = $pdf->extractImage($imageInfo); $fileExt = $imageInfo['ext']; $savePath = $outputDir . "/image_{$index}.{$fileExt}"; file_put_contents($savePath, $imgData); } }
注意事项
PDF的坐标原点在页面左下角,y轴向上,判断时需注意坐标系规则;若允许提取部分重叠的图片,可调整判断逻辑。
方案3:提取所有图片后裁剪目标区域
先提取页面所有图片,再根据PDF坐标转换为图片像素坐标,用ImageMagick裁剪出目标区域的内容。
$originalPdf = 'path/to/your/original.pdf'; $outputDir = 'path/to/output/images'; mkdir($outputDir, 0777, true); // 提取所有图片 $imgExtractCmd = sprintf('pdfimages -j %s %s/raw_image', escapeshellarg($originalPdf), escapeshellarg($outputDir)); exec($imgExtractCmd, $extractLogs, $extractCode); if ($extractCode !== 0) { die("图片提取失败:" . implode("\n", $extractLogs)); } // PDF默认DPI为72,页面高度按实际PDF调整(示例为A4) $pdfDpi = 72; $pageHeight = 841.89; // 目标区域PDF坐标 $targetLeft = 31.46; $targetBottom = 690.22; $targetRight = 560.54; $targetTop = 839; // 遍历提取的原始图片 $rawImages = glob($outputDir . '/raw_image-*.jpg'); foreach ($rawImages as $imgPath) { list($imgWidth, $imgHeight) = getimagesize($imgPath); // 转换PDF坐标到图片像素坐标(图片原点在左上,需反转y轴) $cropLeft = round($targetLeft * ($imgWidth / $pageSize['w'])); $cropTop = round(($pageHeight - $targetTop) * ($imgHeight / $pageHeight)); $cropWidth = round(($targetRight - $targetLeft) * ($imgWidth / $pageSize['w'])); $cropHeight = round(($targetTop - $targetBottom) * ($imgHeight / $pageHeight)); // 用ImageMagick裁剪图片 $cropCmd = sprintf( 'convert %s -crop %dx%d+%d+%d %s_cropped.jpg', escapeshellarg($imgPath), $cropWidth, $cropHeight, $cropLeft, $cropTop, escapeshellarg(pathinfo($imgPath, PATHINFO_FILENAME)) ); exec($cropCmd, $cropLogs, $cropCode); if ($cropCode !== 0) { echo "裁剪图片{$imgPath}失败:" . implode("\n", $cropLogs) . "\n"; } }
内容的提问来源于stack exchange,提问作者Saumini Navaratnam
相关产品推荐
相关产品推荐

