日志逐行处理异常:部分关键词条目丢失的排查与修复
问题描述
需要处理框架生成的日志文件,日志格式示例如下:
log-2023-05-23.log:
INFO - 2023-05-23 09:10:45 --> CSRF token verified. CRITICAL - 2023-05-23 09:24:54 --> Undefined array key 158 in APPPATH/Views/tables/components/record_year__months_regions.php on line 18. 1 APPPATH/Views/tables/components/record_year__months_regions.php(18): CodeIgniter\Debug\Exceptions->errorHandler(2, 'Undefined array key 158', 'FILEPATH', 18) ... 17 FCPATH/index.php(67): CodeIgniter\CodeIgniter->run() CRITICAL - 2023-05-23 09:26:33 --> Undefined array key 158 in APPPATH/Views/tables/components/record_year__months_regions.php on line 18. ... Many more entries starting with CRITICAL/INFO/... are following after this.
需求是逐行读取文件,根据每行首词(INFO、CRITICAL等)判断:要么作为新条目存入数组,要么追加到前一个条目末尾。
尝试方案
编写了如下PHP函数,但仅能保存第1、3、5……个条目,其余条目全部丢失:
function.php:
while (! $file->eof()) { $line = $file->fgets(); $words = explode(' ', $line); $firstWord = $words[0] ?? ''; $timeWord = $words[3] ?? ''; if ($isCollecting) { if (in_array($firstWord, $this->keyWords, true)) { // Found the start of a new entry, so finalize the current entry $output[$date][$keyWord][$time] = $entryLines; $isCollecting = false; $keyWord = null; $time = null; $entryLines = []; } else { // Continue collecting lines for the current entry $entryLines[] = $line; } } elseif (in_array($firstWord, $this->keyWords, true)) { // Found the start of a new entry $isCollecting = true; $keyWord = $firstWord; $time = $timeWord; $entryLines[] = $line; } } // Finalize the last entry if still collecting lines if ($isCollecting) { $output[$date][$keyWord][$time] = $entryLines; }
生成的数组结果(仅展示部分):
array(1) { ["2023-05-23"]=> array(2) { ["INFO"]=> array(1) { ["09:10:45"]=> array(1) { [0]=> string(52) "INFO - 2023-05-23 09:10:45 --> CSRF token verified." } } ["CRITICAL"]=> array(4) { ["09:26:33"]=> ...
问题分析与修复
问题出在遇到新条目时,代码仅完成了前一个条目的保存,却没有立即处理当前这一行的新条目。当前行的首词属于关键词,但此时$isCollecting被设为false,这一行会跳过收集,直到下一次循环才会重新开始,导致这个新条目的第一行丢失,最终只有奇数位置的条目能被完整保存。
修复方法是:在保存完前一个条目后,直接初始化当前新条目的收集流程,无需关闭收集状态:
while (! $file->eof()) { $line = $file->fgets(); // 处理空行,避免explode出错 if (trim($line) === '') { continue; } $words = explode(' ', $line); $firstWord = $words[0] ?? ''; $timeWord = $words[3] ?? ''; if ($isCollecting) { if (in_array($firstWord, $this->keyWords, true)) { // 保存当前条目 $output[$date][$keyWord][$time] = $entryLines; // 直接初始化新条目,不跳过当前行 $keyWord = $firstWord; $time = $timeWord; $entryLines = [$line]; // 保持收集状态,无需设置为false } else { $entryLines[] = $line; } } elseif (in_array($firstWord, $this->keyWords, true)) { $isCollecting = true; $keyWord = $firstWord; $time = $timeWord; $entryLines[] = $line; } } // 保存最后一个条目 if ($isCollecting) { $output[$date][$keyWord][$time] = $entryLines; }
同时添加了空行处理逻辑,避免空行导致的$firstWord为空的情况,防止不必要的错误。
内容的提问来源于stack exchange,提问作者Simon Weber
相关产品推荐
相关产品推荐

