You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PHP如何读取解析Word DOC文件批量提取含公式问答存入数据库

PHP解析Word批量问答(保留数学公式)实现方案

现存问题说明

你当前手写的docx解析代码存在三个核心缺陷,无法满足需求:

  • 仅提取word/document.xml后直接调用strip_tags过滤所有标签,会直接删除Word中存储数学公式的<m:oMath>、<m:oMathPara>OMML节点,导致公式、特殊格式完全丢失
  • 没有做问答结构的规则匹配,无法拆分独立的问题、答案字段
  • 代码仅支持.docx(OpenXML压缩格式)文件,无法处理旧版.doc二进制格式文件

内容示例:
Word问答内容示例

原有参考代码:

function read_file_docx($filename){
    $striped_content = '';
    $content = '';
    if(!$filename || !file_exists($filename)) return false;
    $zip = zip_open($filename);
    if (!$zip || is_numeric($zip)) return false;
    while ($zip_entry = zip_read($zip)) {
        if (zip_entry_open($zip, $zip_entry) == FALSE) continue;
        if (zip_entry_name($zip_entry) != "word/document.xml") continue;
        $content .= zip_entry_read($zip_entry, zip_entry_filesize($zip_entry));
        zip_entry_close($zip_entry);
    }
    zip_close($zip);
    $content = str_replace('</w:r></w:p></w:tc><w:tc>', " ", $content);
    $content = str_replace('</w:r></w:p>', "\r\n", $content);
    $striped_content = strip_tags($content);
    return $striped_content;
}

$filename =$r->file;
$content = $this->read_file_docx($filename);
if($content !== false) {
    echo nl2br($content);
}
else {
    echo 'Couldn\'t the file. Please check that file.';
}
exit;

具体实现步骤

1. 依赖选型

放弃手写原生zip解压解析逻辑,使用成熟开源库降低格式丢失、解析异常概率:

  • docx解析:使用phpoffice/phpword,原生支持OMML公式节点读取,兼容所有Word标准格式
  • 旧版doc处理:服务器安装LibreOffice,先将doc文件批量转换为docx格式后统一解析,不要直接解析二进制doc文件

安装依赖:

composer require phpoffice/phpword

doc转docx命令(服务器端执行):

libreoffice --headless --convert-to docx 源文件.doc --outdir 输出目录

2. 解析核心逻辑

禁止直接过滤所有标签,单独识别、处理公式节点,同时按照文档的排版规则匹配问答边界:

  • 普通文本段落直接提取UTF-8编码内容,特殊符号不会乱码
  • 公式节点提取OMML内容,可直接存储原始XML,或转换为LaTeX格式(前端可用KaTeX/MathJax直接渲染,兼容性更好)
  • 根据文档实际排版规则匹配问答边界:如果是编号类问答,按序号前缀匹配;如果是表格排版的问答,直接按表格列拆分,准确率更高

完整参考代码:

require 'vendor/autoload.php';
use PhpOffice\PhpWord\IOFactory;

/**
 * 从docx文件提取结构化问答数据
 * @param string $filePath docx文件路径
 * @return array 结构化问答列表
 */
function extractQaList(string $filePath): array
{
    $phpWord = IOFactory::load($filePath);
    $qaList = [];
    $currentQuestion = '';
    $currentAnswer = '';

    foreach ($phpWord->getSections() as $section) {
        // 先处理表格类排版:如果问答存在表格里,优先遍历表格元素
        foreach ($section->getTables() as $table) {
            foreach ($table->getRows() as $row) {
                $cells = $row->getCells();
                // 约定第一列是问题、第二列是答案的场景可直接用这段逻辑
                if (count($cells) >= 2) {
                    $qText = trim(collectCellContent($cells[0]));
                    $aText = trim(collectCellContent($cells[1]));
                    if ($qText && $aText) {
                        $qaList[] = ['question' => $qText, 'answer' => $aText];
                    }
                }
            }
        }

        // 处理段落类排版
        foreach ($section->getElements() as $element) {
            $nodeContent = '';
            // 处理普通文本
            if (method_exists($element, 'getText')) {
                $nodeContent = trim($element->getText());
            }
            // 处理OMML公式节点
            if (get_class($element) === 'PhpOffice\PhpWord\Element\OMML') {
                $ommlXml = $element->getXml();
                // 可在此处加OMML转LaTeX逻辑,这里先做标记存储
                $nodeContent = '[FORMULA]' . base64_encode($ommlXml) . '[/FORMULA]';
            }
            if (!$nodeContent) continue;

            // 匹配问题开头:根据自身文档规则调整正则,比如匹配"1.xxx"、"问:xxx"类前缀
            if (preg_match('/^(问|\d+[\.、]|Question)\s*[::]?\s*/u', $nodeContent)) {
                // 存入上一组完整问答
                if ($currentQuestion !== '') {
                    $qaList[] = [
                        'question' => trim($currentQuestion),
                        'answer' => trim($currentAnswer)
                    ];
                    $currentAnswer = '';
                }
                $currentQuestion = preg_replace('/^(问|\d+[\.、]|Question)\s*[::]?\s*/u', '', $nodeContent);
            }
            // 匹配答案开头
            elseif (preg_match('/^(答|答案|Answer)\s*[::]?\s*/u', $nodeContent)) {
                $currentAnswer = preg_replace('/^(答|答案|Answer)\s*[::]?\s*/u', '', $nodeContent);
            }
            // 普通内容归属到当前块
            else {
                if ($currentAnswer !== '' || $currentQuestion === '') {
                    $currentAnswer .= $nodeContent;
                } else {
                    $currentQuestion .= $nodeContent;
                }
            }
        }
    }

    // 存入最后一组问答
    if ($currentQuestion !== '') {
        $qaList[] = [
            'question' => trim($currentQuestion),
            'answer' => trim($currentAnswer)
        ];
    }
    return $qaList;
}

/**
 * 递归提取单元格/嵌套元素内所有内容(含文本、公式)
 */
function collectCellContent($element): string
{
    $content = '';
    if (method_exists($element, 'getText')) {
        $content .= $element->getText();
    }
    if (get_class($element) === 'PhpOffice\PhpWord\Element\OMML') {
        $content .= '[FORMULA]' . base64_encode($element->getXml()) . '[/FORMULA]';
    }
    if (method_exists($element, 'getElements')) {
        foreach ($element->getElements() as $child) {
            $content .= collectCellContent($child);
        }
    }
    return $content;
}

// 调用示例
$qaData = extractQaList('你的问答文件.docx');
// 批量写入数据库,注意数据库字符集使用utf8mb4,支持所有特殊符号存储
foreach ($qaData as $qa) {
    // 执行数据库插入逻辑即可
}

3. 适配优化点

  • 正则匹配规则需要和你实际的文档格式对齐,固定排版的文档解析准确率可达100%
  • 如果需要网页端展示公式,引入OMML转LaTeX的转换类,存储LaTeX字符串后前端用KaTeX/MathJax可直接渲染,不需要额外处理格式
  • 批量处理大文件时,适当调大PHP内存限制,单文件解析完成后手动释放对象,避免内存溢出
  • 数据库字段使用utf8mb4字符集,可完整存储所有特殊符号、公式文本,不会出现乱码

内容的提问来源于stack exchange,提问作者Pawan Verma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 07:36:23