如何用sscanf解析utmpdump输出并去除[]内末尾空格
用sscanf解析utmpdump日志行:适配空列并自动去除末尾空格
我们需要解析utmpdump /var/log/wtmp输出的日志行,格式示例如下:
[8] [13420] [ ] [ ] [pts/3 ] [ ] [0.0.0.0 ] [2024-07-22T11:18:29,836564+00:00] [7] [13611] [ts/3] [john ] [pts/3 ] [192.168.1.38 ] [192.168.1.38 ] [2024-07-22T11:21:30,065856+00:00] [8] [13611] [ ] [ ] [pts/3 ] [ ] [0.0.0.0 ] [2024-07-22T11:21:41,814051+00:00]
该格式核心特点:
- 每行固定8列,所有列用
[]包裹 - 列内容可为全空格(空列),或无前导空格但带末尾空格的文本
- 文本列长度不超过63字符(加上
\0共64)
要求用**单行sscanf**完成解析,直接去除列内末尾空格,无需后续处理。但现有代码在遇到空列时失效,代码及输出如下:
失效的原有代码
#include <stdio.h> #define BUF_LEN (64) int main (void) { const char* const line = "[8] [13420] [ ] [ ] [pts/3 ] [ ] [0.0.0.0 ] [2024-07-22T11:18:29,836564+00:00]"; size_t record_num = 0; size_t pid = 0; char session_type[BUF_LEN] = { 0 }; char username[BUF_LEN] = { 0 }; char terminal[BUF_LEN] = { 0 }; char source_ip[BUF_LEN] = { 0 }; char dest_ip[BUF_LEN] = { 0 }; char timestamp[BUF_LEN] = { 0 }; sscanf(line, "[%zu] [%zu] [%64[^] ] ] [%64[^] ] ] [%64[^] ] ] [%64[^] ] ] [%64[^] ] ] [%64[^] ] ]", &record_num, &pid, session_type, username, terminal, source_ip, dest_ip, timestamp); printf("Record number: %zu \n", record_num); printf("Pid: %zu \n", pid); printf("Session type: %s \n", session_type); printf("User name: %s \n", username); printf("Terminal: %s \n", terminal); printf("Source IP: %s \n", source_ip); printf("Destination IP: %s \n", dest_ip); printf("Timestamp: %s \n", timestamp); return 0; }
原有代码输出
Record number: 8 Pid: 13420 Session type: User name: Terminal: Source IP: Destination IP: Timestamp:
解决方案:修正sscanf格式串
原有代码的问题在于格式串写法错误([^]是无效的字符集),且未处理空列和末尾空格。我们可以利用sscanf的%s自动跳过开头空格、读取到第一个空格为止的特性,结合%*[ ]跳过列内末尾空格,实现单行解析适配空列并自动去尾空格。
修正后的代码
#include <stdio.h> #define BUF_LEN (64) int main (void) { const char* const line = "[8] [13420] [ ] [ ] [pts/3 ] [ ] [0.0.0.0 ] [2024-07-22T11:18:29,836564+00:00]"; size_t record_num = 0; size_t pid = 0; char session_type[BUF_LEN] = { 0 }; char username[BUF_LEN] = { 0 }; char terminal[BUF_LEN] = { 0 }; char source_ip[BUF_LEN] = { 0 }; char dest_ip[BUF_LEN] = { 0 }; char timestamp[BUF_LEN] = { 0 }; // 核心修正:每个字符串列使用 [ %64s%*[ ]] 格式 // %64s:读取非空格内容(自动跳过列内开头空格,适配空列) // %*[ ]:跳过列内末尾所有空格,不写入变量 int match_count = sscanf(line, "[%zu] [%zu] [ %64s%*[ ]] [ %64s%*[ ]] [ %64s%*[ ]] [ %64s%*[ ]] [ %64s%*[ ]] [ %64s%*[ ]]", &record_num, &pid, session_type, username, terminal, source_ip, dest_ip, timestamp); printf("匹配字段数:%d\n", match_count); printf("Record number: %zu \n", record_num); printf("Pid: %zu \n", pid); printf("Session type: '%s' \n", session_type); printf("User name: '%s' \n", username); printf("Terminal: '%s' \n", terminal); printf("Source IP: '%s' \n", source_ip); printf("Destination IP: '%s' \n", dest_ip); printf("Timestamp: '%s' \n", timestamp); return 0; }
修正后输出
匹配字段数:8 Record number: 8 Pid: 13420 Session type: '' User name: '' Terminal: 'pts/3' Source IP: '' Destination IP: '0.0.0.0' Timestamp: '2024-07-22T11:18:29,836564+00:00'
格式串说明
[%zu]:解析前两个数字列(记录号、PID)[ %64s%*[ ]]:处理每个字符串列:[:匹配列开头的[和后续空格%64s:读取列内的非空格内容(空列时无内容,变量保持空字符串)%*[ ]:跳过列内末尾的所有空格(*表示只匹配不写入变量)]:匹配列结尾的]
这样既适配了空列的情况,又在解析时直接去除了列内的末尾空格,完全满足需求。
内容的提问来源于stack exchange,提问作者Łukasz Przeniosło
相关产品推荐
相关产品推荐

