Laravel 9中拆分DOCX/PDF页面并按页入库(适配右向语言)
多格式书籍按页提取入库解决方案(适配RTL语言)
一、PDF文件处理(解决波斯语/阿拉伯语乱码+分页提取)
使用smalot/pdfparser库,它支持自定义字体映射,能有效处理RTL语言:
- 安装依赖:
composer require smalot/pdfparser
- 配置RTL字体支持:
- 下载支持波斯语/阿拉伯语的字体(如Vazirmatn、Noto Naskh Arabic),放在Laravel项目的
storage/fonts目录下。 - 初始化PDF解析器时指定字体路径:
use Smalot\PdfParser\Parser; use Smalot\PdfParser\Config; $config = new Config(); $config->setFontDir(storage_path('fonts')); $config->setFontCache(storage_path('fonts/cache')); // 添加字体映射,覆盖默认缺失的RTL字体 $config->setFonts([ 'Arial' => 'Vazirmatn-Regular.ttf', 'Times New Roman' => 'NotoNaskhArabic-Regular.ttf', ]); $parser = new Parser($config); - 下载支持波斯语/阿拉伯语的字体(如Vazirmatn、Noto Naskh Arabic),放在Laravel项目的
- 按页提取内容:
$pdf = $parser->parseFile($uploadedFile->getPathname()); $pages = $pdf->getPages(); foreach ($pages as $index => $page) { $content = $page->getText(); // 存入pages表 Page::create([ 'book_id' => $book->id, 'page_number' => $index + 1, 'content' => $content, 'language' => 'fa' // 或'ar',根据书籍语言设置 ]); }
二、DOCX文件分页提取
DOCX的分页标记在XML中是<w:br w:type="page"/>,用phpoffice/phpword库可以准确识别:
- 安装依赖:
composer require phpoffice/phpword
- 提取分页内容:
use PhpOffice\PhpWord\IOFactory; $phpWord = IOFactory::load($uploadedFile->getPathname()); $sections = $phpWord->getSections(); $currentPageContent = ''; $pageNumber = 1; foreach ($sections as $section) { foreach ($section->getElements() as $element) { // 检测分页符 if ($element instanceof \PhpOffice\PhpWord\Element\Break && $element->getType() === \PhpOffice\PhpWord\Element\Break::TYPE_PAGE) { // 保存当前页内容 Page::create([ 'book_id' => $book->id, 'page_number' => $pageNumber++, 'content' => $currentPageContent, ]); $currentPageContent = ''; continue; } // 追加段落内容 if ($element instanceof \PhpOffice\PhpWord\Element\TextRun) { foreach ($element->getElements() as $textElement) { $currentPageContent .= $textElement->getText() . ' '; } } elseif ($element instanceof \PhpOffice\PhpWord\Element\Text) { $currentPageContent .= $element->getText() . ' '; } } } // 保存最后一页内容 if (!empty(trim($currentPageContent))) { Page::create([ 'book_id' => $book->id, 'page_number' => $pageNumber, 'content' => $currentPageContent, ]); }
三、TXT文件分页处理
TXT无原生分页标记,可自定义分页规则:
- 按固定行数分页:
$content = file_get_contents($uploadedFile->getPathname()); $lines = explode("\n", $content); $linesPerPage = 50; // 可配置,根据需求调整 $pages = array_chunk($lines, $linesPerPage); foreach ($pages as $index => $pageLines) { $pageContent = implode("\n", $pageLines); Page::create([ 'book_id' => $book->id, 'page_number' => $index + 1, 'content' => $pageContent, ]); }
- 或按字符数分页(适合RTL语言,避免截断单词):
$content = file_get_contents($uploadedFile->getPathname()); $charsPerPage = 2000; $pageNumber = 1; $start = 0; while ($start < strlen($content)) { // 从末尾向前找空格,避免截断单词 $end = min($start + $charsPerPage, strlen($content)); if ($end < strlen($content)) { $end = strrpos(substr($content, 0, $end), ' ') + 1; } $pageContent = substr($content, $start, $end - $start); Page::create([ 'book_id' => $book->id, 'page_number' => $pageNumber++, 'content' => $pageContent, ]); $start = $end; }
四、Laravel集成优化建议
- 异步处理:书籍上传后用Laravel队列处理文件解析入库,避免请求超时:
// 上传控制器中 ProcessBookPages::dispatch($book, $uploadedFile); - 数据库字符集:确保
pages表的content字段使用utf8mb4_unicode_ci字符集,支持RTL语言的所有字符。 - 语言识别:可集成语言检测库(如
patrikstarlinger/php-language-detector)自动识别书籍语言,方便后续搜索和显示。
内容的提问来源于stack exchange,提问作者Kmaj
相关产品推荐
相关产品推荐

