Array Exploding问题求助:附data.txt及SIP通话相关数据
解决Array Exploding问题:正确拆分带空格字段的通话日志
我猜你在处理这份Asterisk通话日志(data.txt)时,遇到了数组拆分的麻烦——直接按空格拆分的话,会把( 0.00%)这种内部带空格的字段拆成多个零散元素,导致数组结构完全混乱。先把你的日志内容贴出来方便参考:
Peer Call ID Duration Recv: Pack Lost ( %) Jitter Send: Pack Lost ( %) Jitter 139.59.232.196 0bb9262d6a1 00:01:12 0000003558 0000000000 ( 0.00%) 0.0000 0000001177 0000000000 ( 0.00%) 0.0200 139.59.232.196 41283499492 00:00:00 0000000000 0000000000 ( 0.00%) 0.0000 0000000000 0000000000 ( 0.00%) 0.0000 139.59.232.196 7033a541240 00:00:08 0000000000 0000000000 ( 0.00%) 0.0000 0000000019 0000000000 ( 0.00%) 0.0000 3 active SIP chan...
问题的核心很明确:日志是固定列数的结构化数据,但部分字段(比如丢失率)内部包含空格,普通的全空格拆分逻辑会把这些字段拆碎。下面给你几种不同场景下的解决方案:
1. PHP 处理方案
用正则表达式精准匹配每一列,把带空格的字段作为整体捕获,避免拆分:
$fileHandle = fopen('data.txt', 'r'); // 跳过第一行表头 fgets($fileHandle); while (($line = fgets($fileHandle)) !== false) { $line = trim($line); // 跳过最后一行的提示信息 if (str_contains($line, 'active SIP chan')) continue; // 正则匹配所有列,将带空格的丢失率字段作为整体提取 if (preg_match('/^(\S+) (\S+) (\S+) (\S+) (\S+) (\( \d+\.\d+%\)) (\S+) (\S+) (\S+) (\( \d+\.\d+%\)) (\S+)$/', $line, $matches)) { array_shift($matches); // 移除匹配的整行字符串 // 整理成可读性更高的关联数组 $callLog = [ 'peer' => $matches[0], 'call_id' => $matches[1], 'duration' => $matches[2], 'recv_packets' => $matches[3], 'recv_lost' => $matches[4], 'recv_lost_percent' => $matches[5], 'recv_jitter' => $matches[6], 'send_packets' => $matches[7], 'send_lost' => $matches[8], 'send_lost_percent' => $matches[9], 'send_jitter' => $matches[10] ]; // 这里可以根据需求处理数据,比如存入数据库或输出 print_r($callLog); } } fclose($fileHandle);
2. Python 处理方案
适合做数据分析或批量脚本的场景,同样用正则提取完整字段:
import re # 读取日志文件 with open('data.txt', 'r', encoding='utf-8') as f: all_lines = f.readlines() # 跳过表头和最后一行提示,过滤空行 data_lines = [line.strip() for line in all_lines[1:-1] if line.strip()] # 正则匹配规则 pattern = r'^(\S+) (\S+) (\S+) (\S+) (\S+) (\( \d+\.\d+%\)) (\S+) (\S+) (\S+) (\( \d+\.\d+%\)) (\S+)$' call_records = [] for line in data_lines: match_result = re.match(pattern, line) if match_result: groups = match_result.groups() call_records.append({ 'peer': groups[0], 'call_id': groups[1], 'duration': groups[2], 'recv_packets': groups[3], 'recv_lost': groups[4], 'recv_lost_percent': groups[5], 'recv_jitter': groups[6], 'send_packets': groups[7], 'send_lost': groups[8], 'send_lost_percent': groups[9], 'send_jitter': groups[10] }) # 打印处理后的结果 for record in call_records: print(record)
3. Bash/Awk 处理方案
如果是在Linux服务器上做批量处理,用Awk脚本最省心,直接按字段位置合并带空格的部分:
#!/bin/bash # 跳过表头,处理数据行,忽略最后一行提示 tail -n +2 data.txt | head -n -1 | awk '{ printf "=== 通话记录 ===\n" printf "Peer: %s\n", $1 printf "Call ID: %s\n", $2 printf "Duration: %s\n", $3 printf "接收数据包: %s\n", $4 printf "接收丢失数: %s\n", $5 printf "接收丢失率: %s %s\n", $6, $7 # 合并带空格的丢失率字段 printf "接收抖动: %s\n", $8 printf "发送数据包: %s\n", $9 printf "发送丢失数: %s\n", $10 printf "发送丢失率: %s %s\n", $11, $12 # 合并带空格的丢失率字段 printf "发送抖动: %s\n", $13 printf "\n" }'
这样处理后,就能得到结构完全正确的数组/字典,不会再出现字段被意外拆分的问题。
内容的提问来源于stack exchange,提问作者Asterisk Developer
相关产品推荐
相关产品推荐

