You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用sscanf处理长字符串分析?以长MAC地址转数组为例

Handling Extremely Long MAC-like Hex Strings in C

Great question! Hardcoding %x-%x format specifiers works for short MAC addresses, but it's totally impractical for super long strings with repeated hex pairs. Let's walk through two clean, scalable solutions to this problem.

The Core Issue

Your original approach relies on static format strings, which can't adapt to arbitrary lengths. The real challenge is parsing each pair of hex characters (separated by -) into a single unsigned char—something memcpy alone can't do because it just copies bytes, not converts hex text to binary values.


Solution 1: Manual Parsing with a Helper Function

This approach gives you full control over the parsing logic, and it works with any length of input string. We'll write a small helper function to convert two hex characters to a byte, then loop through the input string to process each pair.

#include <stdio.h>
#include <stdlib.h>
#include <ctype.h>

// Convert two hex characters (e.g., "ff" → 0xff) to an unsigned char
unsigned char hex_pair_to_byte(const char *hex) {
    unsigned char byte = 0;
    for (int i = 0; i < 2; i++) {
        char c = toupper(hex[i]); // Handle lowercase hex chars gracefully
        if (c >= '0' && c <= '9') {
            byte = (byte << 4) | (c - '0');
        } else if (c >= 'A' && c <= 'F') {
            byte = (byte << 4) | (c - 'A' + 10);
        } else {
            fprintf(stderr, "Error: Invalid hex character '%c'\n", c);
            exit(EXIT_FAILURE);
        }
    }
    return byte;
}

int main() {
    // Example ultra-long hex string (repeat as many times as needed)
    char long_hex_str[] = "ff-13-a9-1f-b0-88-ff-13-a9-1f-b0-88-ff-13-a9-1f-b0-88-ff-00";
    unsigned char *byte_array = NULL;
    int byte_count = 0;
    char *ptr = long_hex_str;

    // First pass: count how many hex pairs we have
    char *temp_ptr = ptr;
    while (*temp_ptr != '\0') {
        if (*temp_ptr != '-') {
            temp_ptr += 2;
            byte_count++;
        } else {
            temp_ptr++;
        }
    }

    // Allocate memory for our result array
    byte_array = malloc(byte_count * sizeof(unsigned char));
    if (!byte_array) {
        perror("Failed to allocate memory");
        exit(EXIT_FAILURE);
    }

    // Second pass: parse each hex pair into a byte
    int idx = 0;
    while (*ptr != '\0') {
        if (*ptr == '-') {
            ptr++;
            continue;
        }
        byte_array[idx++] = hex_pair_to_byte(ptr);
        ptr += 2;
    }

    // Print the result to verify
    printf("Parsed %d bytes:\n", byte_count);
    for (int i = 0; i < byte_count; i++) {
        printf("%02x ", byte_array[i]);
        if ((i + 1) % 6 == 0) { // Group into MAC-style 6-byte chunks
            printf("\n");
        }
    }
    printf("\n");

    free(byte_array);
    return 0;
}

Why This Works:

  • Scalable: No matter how long your input string is, the loop handles every hex pair automatically.
  • Robust: The helper function handles both uppercase and lowercase hex characters, and includes basic error checking for invalid input.
  • No Static Format Strings: No more copying %x-%x a hundred times!

Solution 2: Use strtok and strtoul (Standard Library Shortcut)

If you prefer leveraging C's standard library instead of writing custom parsing logic, this method splits the input string with strtok (using - as the delimiter), then converts each hex segment to a byte with strtoul.

#include <stdio.h>
#include <stdlib.h>
#include <string.h>

int main() {
    char long_hex_str[] = "ff-13-a9-1f-b0-88-ff-13-a9-1f-b0-88-ff-00";
    unsigned char *byte_array = NULL;
    int byte_count = 0;

    // First pass: count the number of hex segments
    char *token = strtok(long_hex_str, "-");
    while (token != NULL) {
        byte_count++;
        token = strtok(NULL, "-");
    }

    // Make a copy of the original string (strtok modifies the input)
    char *str_copy = strdup(long_hex_str);
    if (!str_copy) {
        perror("Failed to copy string");
        exit(EXIT_FAILURE);
    }

    // Allocate memory for the result
    byte_array = malloc(byte_count * sizeof(unsigned char));
    if (!byte_array) {
        perror("Failed to allocate memory");
        free(str_copy);
        exit(EXIT_FAILURE);
    }

    // Second pass: convert each segment to a byte
    int idx = 0;
    token = strtok(str_copy, "-");
    while (token != NULL) {
        unsigned long val = strtoul(token, NULL, 16);
        // Ensure the value fits in a single byte
        if (val > 0xFF) {
            fprintf(stderr, "Error: '%s' is too large for a byte\n", token);
            free(byte_array);
            free(str_copy);
            exit(EXIT_FAILURE);
        }
        byte_array[idx++] = (unsigned char)val;
        token = strtok(NULL, "-");
    }

    // Print verification
    printf("Parsed %d bytes:\n", byte_count);
    for (int i = 0; i < byte_count; i++) {
        printf("%02x ", byte_array[i]);
        if ((i + 1) % 6 == 0) {
            printf("\n");
        }
    }
    printf("\n");

    free(byte_array);
    free(str_copy);
    return 0;
}

Why This Works:

  • Concise: Uses standard library functions to avoid writing custom hex conversion code.
  • Readable: The logic is straightforward—split the string, convert each part.
  • Note: strtok modifies the input string, so we make a copy with strdup to preserve the original data.

Which One Should You Choose?

  • Use Solution 1 if you need fine-grained control over parsing (e.g., custom error handling, non-standard hex formats).
  • Use Solution 2 if you want a shorter, more maintainable implementation that relies on trusted standard library functions.

Both approaches will handle your ultra-long repeated hex strings without any messy hardcoded format specifiers.

内容的提问来源于stack exchange,提问作者Michael

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 10:09:11