如何获取PDF中各内容的具体坐标?基于Smalot\PdfParser的疑问
使用Smalot\PdfParser提取PDF文本及坐标信息
问题描述
作为新手,使用Smalot\PdfParser提取PDF内容时,尝试了getText()、getDetails()、getPages()等基础方法,发现调用$page->getDataTm()返回的数组中包含疑似坐标的数据,但不清楚数组中哪个是X、Y值,也不知道如何针对性提取。已开启DataTm字体信息配置,期望提取后生成包含页面信息、内容坐标(x、y、w、h)及对应文本的TXT输出,目标PDF为多页带表格结构的文档。
已知返回数据示例
0 => array:4 [▼ 0 => array:6 [▼ 0 => "1.00055" 1 => "0" 2 => "0" 3 => "1" 4 => "70.8" 5 => "760.24" ] 1 => " " 2 => "R8" 3 => "12" ] 1 => array:4 [▼ 0 => array:6 [▼ 0 => "1.00055" 1 => "0" 2 => "0" 3 => "1" 4 => "70.8" 5 => "745.72" ] 1 => "Column1 Column2 Column3 " 2 => "R10" 3 => "12" ]
现有代码
use Smalot\PdfParser\Parser; use Smalot\PdfParser\Config; /* ... */ protected function getCoordinates($pdfPath) { // get font details by config $config = new Config(); $config->setDataTmFontInfoHasToBeIncluded(true); // get PDF structure $parser = new Parser([], $config); $pdf = $parser->parseFile($pdfPath); $coordinates = []; //dd($pdf->getPages()[1]->getDataTm()); foreach ($pdf->getPages() as $page) { $page->getDataTm(); $text = $page->getText(); //$coordinates = ; // This is where I want to extract it } return $coordinates; }
期望输出格式
[Page : 1, width = 1, height = 2] [x:0, y:3, w: 4, h:5]Column1 Column2 Column3 [x:6, y:7, w: 8, h:5]L1C1 L1C2 L1C3 [x:6, y:7, w: 8, h:5]L2C1 L2C2 L2C3 [x:6, y:7, w: 8, h:5]L3C1 L3C2 L3C3 [x:6, y:7, w: 8, h:5]L4C1 L4C2 L4C3
解决方案
1. getDataTm()数组的坐标含义
getDataTm()返回的每个子数组中,第一个元素是PDF文本定位用的变换矩阵数组,各索引对应参数:
- 索引4:文本左下角的X坐标
- 索引5:文本左下角的Y坐标
- 索引0:X方向缩放因子(用于计算文本宽度)
- 子数组第3位(索引3):Y方向缩放因子(通常等于字体大小,即文本高度
h) - 数组第3个元素(
$item[3]):字体大小数值
2. 修改后的提取代码
use Smalot\PdfParser\Parser; use Smalot\PdfParser\Config; protected function getCoordinates($pdfPath) { $config = new Config(); $config->setDataTmFontInfoHasToBeIncluded(true); $parser = new Parser([], $config); $pdf = $parser->parseFile($pdfPath); $output = ''; foreach ($pdf->getPages() as $pageIndex => $page) { $pageNum = $pageIndex + 1; // 获取页面宽高(从MediaBox参数提取) $pageDetails = $page->getDetails(); $pageWidth = $pageDetails['MediaBox'][2] ?? '未知'; $pageHeight = $pageDetails['MediaBox'][3] ?? '未知'; $output .= "[Page : {$pageNum}, width = {$pageWidth}, height = {$pageHeight}]\n"; // 遍历当前页的DataTm数据 $dataTmItems = $page->getDataTm(); foreach ($dataTmItems as $item) { $transformMatrix = $item[0]; $text = trim($item[1]); // 跳过空文本条目 if (empty($text)) continue; // 提取基础坐标 $x = $transformMatrix[4]; $y = $transformMatrix[5]; // 获取字体大小,计算文本宽高 $fontSize = (float)$item[3]; $xScale = (float)$transformMatrix[0]; // 近似计算文本宽度:字符数 × 字体大小 × X缩放因子 $textWidth = strlen($text) * $fontSize * $xScale; $textHeight = $fontSize; // 格式化坐标行 $output .= sprintf( "[x:%s, y:%s, w: %.2f, h:%.2f]%s\n", $x, $y, $textWidth, $textHeight, $text ); } } // 将结果写入TXT文件 file_put_contents('pdf_content_with_coords.txt', $output); return $output; }
3. 代码说明
- 通过
$page->getDetails()获取页面的MediaBox参数,得到页面实际宽高 - 遍历
getDataTm()的每个条目,提取变换矩阵、文本内容和字体大小 - 文本宽度采用近似计算方式(PDF中精确宽度需考虑字符间距,此为简化方案)
- 过滤空文本,避免生成无效输出行
- 最终将结果写入指定TXT文件,同时返回输出内容便于后续处理
4. 实际输出示例
[Page : 1, width = 612, height = 792] [x:70.8, y:745.72, w: 324.18, h:12.00]Column1 Column2 Column3 [x:70.8, y:730.20, w: 324.18, h:12.00]L1C1 L1C2 L1C3 [x:70.8, y:714.68, w: 324.18, h:12.00]L2C1 L2C2 L2C3 [x:70.8, y:699.16, w: 324.18, h:12.00]L3C1 L3C2 L3C3
内容的提问来源于stack exchange,提问作者Keith Lè
相关产品推荐
相关产品推荐

