You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

日志逐行处理异常:部分关键词条目丢失的排查与修复

问题描述

需要处理框架生成的日志文件,日志格式示例如下:
log-2023-05-23.log:

INFO - 2023-05-23 09:10:45 --> CSRF token verified.
CRITICAL - 2023-05-23 09:24:54 --> Undefined array key 158
in APPPATH/Views/tables/components/record_year__months_regions.php on line 18.
 1 APPPATH/Views/tables/components/record_year__months_regions.php(18): CodeIgniter\Debug\Exceptions->errorHandler(2, 'Undefined array key 158', 'FILEPATH', 18)
...
17 FCPATH/index.php(67): CodeIgniter\CodeIgniter->run()
CRITICAL - 2023-05-23 09:26:33 --> Undefined array key 158
in APPPATH/Views/tables/components/record_year__months_regions.php on line 18.
...
Many more entries starting with CRITICAL/INFO/... are following after this.

需求是逐行读取文件,根据每行首词(INFO、CRITICAL等)判断:要么作为新条目存入数组,要么追加到前一个条目末尾。

尝试方案

编写了如下PHP函数,但仅能保存第1、3、5……个条目,其余条目全部丢失:
function.php:

while (! $file->eof()) {
    $line      = $file->fgets();
    $words     = explode(' ', $line);
    $firstWord = $words[0] ?? '';
    $timeWord  = $words[3] ?? '';
    if ($isCollecting) {
        if (in_array($firstWord, $this->keyWords, true)) {
            // Found the start of a new entry, so finalize the current entry
            $output[$date][$keyWord][$time] = $entryLines;
            $isCollecting                   = false;
            $keyWord                        = null;
            $time                           = null;
            $entryLines                     = [];
        } else {
            // Continue collecting lines for the current entry
            $entryLines[] = $line;
        }
    } elseif (in_array($firstWord, $this->keyWords, true)) {
        // Found the start of a new entry
        $isCollecting = true;
        $keyWord      = $firstWord;
        $time         = $timeWord;
        $entryLines[] = $line;
    }
}
// Finalize the last entry if still collecting lines
if ($isCollecting) {
    $output[$date][$keyWord][$time] = $entryLines;
}

生成的数组结果(仅展示部分):

array(1) {
  ["2023-05-23"]=>
  array(2) {
    ["INFO"]=>
    array(1) {
      ["09:10:45"]=>
      array(1) {
        [0]=>
        string(52) "INFO - 2023-05-23 09:10:45 --> CSRF token verified."
      }
    }
    ["CRITICAL"]=>
    array(4) {
      ["09:26:33"]=> ...
问题分析与修复

问题出在遇到新条目时,代码仅完成了前一个条目的保存,却没有立即处理当前这一行的新条目。当前行的首词属于关键词,但此时$isCollecting被设为false,这一行会跳过收集,直到下一次循环才会重新开始,导致这个新条目的第一行丢失,最终只有奇数位置的条目能被完整保存。

修复方法是:在保存完前一个条目后,直接初始化当前新条目的收集流程,无需关闭收集状态:

while (! $file->eof()) {
    $line      = $file->fgets();
    // 处理空行,避免explode出错
    if (trim($line) === '') {
        continue;
    }
    $words     = explode(' ', $line);
    $firstWord = $words[0] ?? '';
    $timeWord  = $words[3] ?? '';
    
    if ($isCollecting) {
        if (in_array($firstWord, $this->keyWords, true)) {
            // 保存当前条目
            $output[$date][$keyWord][$time] = $entryLines;
            // 直接初始化新条目,不跳过当前行
            $keyWord      = $firstWord;
            $time         = $timeWord;
            $entryLines   = [$line];
            // 保持收集状态,无需设置为false
        } else {
            $entryLines[] = $line;
        }
    } elseif (in_array($firstWord, $this->keyWords, true)) {
        $isCollecting = true;
        $keyWord      = $firstWord;
        $time         = $timeWord;
        $entryLines[] = $line;
    }
}
// 保存最后一个条目
if ($isCollecting) {
    $output[$date][$keyWord][$time] = $entryLines;
}

同时添加了空行处理逻辑,避免空行导致的$firstWord为空的情况,防止不必要的错误。

内容的提问来源于stack exchange,提问作者Simon Weber

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 11:00:20