PHP提取含希伯来语的PDF文本显示乱码,设UTF-8仍无效如何解决
问题解决方法
核心问题说明
你当前的操作存在两个核心错误:
- 直接使用
file_get_contents()读取PDF二进制文件:PDF并非纯文本格式,内部包含压缩块、字体映射、结构化标记等非文本内容,直接读取必然得到乱码,和页面编码设置无关 - 错误调用
utf8_encode():该函数仅支持将ISO-8859-1编码的内容转为UTF-8,你读取的PDF二进制内容不属于该编码范畴,调用该函数会进一步破坏字符编码,加重乱码问题
适配表格+希伯来语的解决方案
因为你不能使用pdfParser,优先推荐使用pdftotext工具(属于poppler-utils组件),它对表格布局的保留效果、希伯来语右到左文本的支持都优于纯PHP实现的解析库:
前置配置
服务器安装poppler-utils组件:
- Linux系统执行安装命令:
apt install poppler-utils - Windows系统可下载poppler预编译二进制包,配置好环境变量即可
修复后的代码示例
<?php header('Content-type: text/html; charset=UTF-8'); $formReturn = $_POST["formReturn"] ?? ''; $text = ''; if ($formReturn && !empty($_FILES["gradesPdf"]["tmp_name"])) { $tmpPdfPath = $_FILES["gradesPdf"]["tmp_name"]; $tmpTxtPath = tempnam(sys_get_temp_dir(), 'pdf_extract_'); // 调用pdftotext,指定UTF-8编码,-layout参数保留表格排版 $cmd = sprintf( "/usr/bin/pdftotext -enc UTF-8 -layout %s %s", escapeshellarg($tmpPdfPath), escapeshellarg($tmpTxtPath) ); exec($cmd, $output, $returnVar); if ($returnVar === 0) { $text = file_get_contents($tmpTxtPath); } unlink($tmpTxtPath); } $html = ' <!DOCTYPE html> <html lang="he"> <head> <meta charset="utf-8" /> <title>נסיון</title> <style> /* 适配希伯来语右到左阅读习惯 */ .extract-content { direction: rtl; text-align: right; font-family: "Arial Hebrew", sans-serif; } </style> </head> <body> <form enctype="multipart/form-data" method="post"> <input type="file" name="gradesPdf" id="gradesPdf"> <br><br> <button type="submit">run</button> <input type="hidden" name="formReturn" value="1"> </form> <div class="extract-content">'. htmlspecialchars($text) .'</div> </body> </html> '; echo $html;
额外注意事项
- 如果pdftotext执行失败,可先在服务器命令行手动执行提取命令排查问题,确认二进制路径是否正确
- 输出文本时加
htmlspecialchars()避免PDF中的特殊字符破坏HTML结构 - 可根据显示效果调整字体配置,确保使用支持希伯来语的字体渲染文本
内容的提问来源于stack exchange,提问作者Yair Cohen
相关产品推荐
相关产品推荐

