You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用sscanf解析utmpdump输出并去除[]内末尾空格

用sscanf解析utmpdump日志行:适配空列并自动去除末尾空格

我们需要解析utmpdump /var/log/wtmp输出的日志行,格式示例如下:

[8] [13420] [    ] [        ] [pts/3       ] [                    ] [0.0.0.0        ] [2024-07-22T11:18:29,836564+00:00]
[7] [13611] [ts/3] [john    ] [pts/3       ] [192.168.1.38        ] [192.168.1.38   ] [2024-07-22T11:21:30,065856+00:00]
[8] [13611] [    ] [        ] [pts/3       ] [                    ] [0.0.0.0        ] [2024-07-22T11:21:41,814051+00:00]

该格式核心特点:

  • 每行固定8列,所有列用[]包裹
  • 列内容可为全空格(空列),或无前导空格但带末尾空格的文本
  • 文本列长度不超过63字符(加上\0共64)

要求用**单行sscanf**完成解析,直接去除列内末尾空格,无需后续处理。但现有代码在遇到空列时失效,代码及输出如下:


失效的原有代码

#include <stdio.h>

#define BUF_LEN     (64)

int main (void) 
{
    const char* const line = "[8] [13420] [    ] [        ] [pts/3       ] [                    ] [0.0.0.0        ] [2024-07-22T11:18:29,836564+00:00]";
    size_t record_num = 0;
    size_t pid = 0;
    char session_type[BUF_LEN] = { 0 };
    char username[BUF_LEN] = { 0 };
    char terminal[BUF_LEN] = { 0 };
    char source_ip[BUF_LEN] = { 0 };
    char dest_ip[BUF_LEN] = { 0 };
    char timestamp[BUF_LEN] = { 0 };

    sscanf(line, "[%zu] [%zu] [%64[^] ] ] [%64[^] ] ] [%64[^] ] ] [%64[^] ] ] [%64[^] ] ] [%64[^] ] ]", 
        &record_num, &pid, session_type, username, terminal, source_ip, dest_ip, timestamp);
    
    printf("Record number: %zu \n", record_num);
    printf("Pid: %zu \n", pid);
    printf("Session type: %s \n", session_type);
    printf("User name: %s \n", username);
    printf("Terminal: %s \n", terminal);
    printf("Source IP: %s \n", source_ip);
    printf("Destination IP: %s \n", dest_ip);
    printf("Timestamp: %s \n", timestamp);

    return 0;
}

原有代码输出

Record number: 8
Pid: 13420
Session type:
User name:
Terminal:
Source IP:
Destination IP:
Timestamp:

解决方案:修正sscanf格式串

原有代码的问题在于格式串写法错误([^]是无效的字符集),且未处理空列和末尾空格。我们可以利用sscanf的%s自动跳过开头空格、读取到第一个空格为止的特性,结合%*[ ]跳过列内末尾空格,实现单行解析适配空列并自动去尾空格。

修正后的代码

#include <stdio.h>

#define BUF_LEN     (64)

int main (void) 
{
    const char* const line = "[8] [13420] [    ] [        ] [pts/3       ] [                    ] [0.0.0.0        ] [2024-07-22T11:18:29,836564+00:00]";
    size_t record_num = 0;
    size_t pid = 0;
    char session_type[BUF_LEN] = { 0 };
    char username[BUF_LEN] = { 0 };
    char terminal[BUF_LEN] = { 0 };
    char source_ip[BUF_LEN] = { 0 };
    char dest_ip[BUF_LEN] = { 0 };
    char timestamp[BUF_LEN] = { 0 };

    // 核心修正:每个字符串列使用 [ %64s%*[ ]] 格式
    // %64s:读取非空格内容(自动跳过列内开头空格,适配空列)
    // %*[ ]:跳过列内末尾所有空格,不写入变量
    int match_count = sscanf(line, "[%zu] [%zu] [ %64s%*[ ]] [ %64s%*[ ]] [ %64s%*[ ]] [ %64s%*[ ]] [ %64s%*[ ]] [ %64s%*[ ]]", 
        &record_num, &pid, session_type, username, terminal, source_ip, dest_ip, timestamp);
    
    printf("匹配字段数:%d\n", match_count);
    printf("Record number: %zu \n", record_num);
    printf("Pid: %zu \n", pid);
    printf("Session type: '%s' \n", session_type);
    printf("User name: '%s' \n", username);
    printf("Terminal: '%s' \n", terminal);
    printf("Source IP: '%s' \n", source_ip);
    printf("Destination IP: '%s' \n", dest_ip);
    printf("Timestamp: '%s' \n", timestamp);

    return 0;
}

修正后输出

匹配字段数:8
Record number: 8 
Pid: 13420 
Session type: '' 
User name: '' 
Terminal: 'pts/3' 
Source IP: '' 
Destination IP: '0.0.0.0' 
Timestamp: '2024-07-22T11:18:29,836564+00:00' 

格式串说明

  • [%zu]:解析前两个数字列(记录号、PID)
  • [ %64s%*[ ]]:处理每个字符串列:
    1. [ :匹配列开头的[和后续空格
    2. %64s:读取列内的非空格内容(空列时无内容,变量保持空字符串)
    3. %*[ ]:跳过列内末尾的所有空格(*表示只匹配不写入变量)
    4. ]:匹配列结尾的]

这样既适配了空列的情况,又在解析时直接去除了列内的末尾空格,完全满足需求。

内容的提问来源于stack exchange,提问作者Łukasz Przeniosło

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 11:15:55