You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修复PHP解析PDF时特定换行导致的货币条目拆分问题

修复PDF解析中被错误拆分的货币条目

你可以通过分析文本行的特征,手动合并被错误拆分的条目。以下是具体实现方案:

核心思路

正常货币条目格式为「3位大写货币代码 + 货币名称 + 两个数值」,损坏的PDF将部分条目拆分为多行。我们可以通过识别行的开头特征(3位大写字母、数值开头),将拆分的行合并为完整条目。

代码实现

$parser = new Parser();
$pdf = $parser->parseFile($path . '/' . $filename);
$pdf_text = $pdf->getText();

// 拆分原始文本为行数组,过滤空行
$lines = array_filter(array_map('trim', explode("\n", $pdf_text)));

$result = [];
$currentEntry = '';

foreach ($lines as $line) {
    // 匹配3位大写货币代码开头的行(新条目起始)
    if (preg_match('/^[A-Z]{3}$/', $line)) {
        // 若有未完成的条目,先存入结果
        if ($currentEntry) {
            $result[] = $currentEntry;
            $currentEntry = '';
        }
        $currentEntry = $line;
    }
    // 匹配数值开头的行(条目结尾)
    elseif (preg_match('/^[0-9,]+/', $line)) {
        $currentEntry .= "\t" . $line;
        $result[] = $currentEntry;
        $currentEntry = '';
    }
    // 匹配名称/发行机构行,合并到当前条目
    else {
        if ($currentEntry) {
            $currentEntry .= "\t" . $line;
        } else {
            $result[] = $line;
        }
    }
}

// 处理最后一个未完成的条目
if ($currentEntry) {
    $result[] = $currentEntry;
}

// 生成修复后的文本(用制表符分隔字段,可根据需求替换为空格)
$fixed_pdf_text = implode("\n", $result);
// 如需转成<br>格式
$fixed_pdf_with_line_breaks = nl2br($fixed_pdf_text);

效果说明

运行上述代码后,原本拆分的XAF条目会被合并为单行:

XAF    Central African CFA Franc    BEAC    655,957000     0,001524

与其他正常格式的货币条目保持一致,后续可正常按行处理。

内容的提问来源于stack exchange,提问作者Carlo Pokker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 20:46:34