You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

排查HackerRank文档结构化程序的EXC_BAD_ACCESS及word_count异常

Fixing Array Out-of-Bounds & EXC_BAD_ACCESS in HackerRank's Structuring the Document Problem

Let's break down exactly what's causing your weird word_count value and the crash, then fix those issues step by step. You're spot-on that array out-of-bounds access is the root problem—here's where things go wrong and how to fix them:


Key Issues in Your Code

1. Local Stack Arrays Have Expired Lifetimes

Your code uses local stack arrays like struct paragraph para[doc.paragraph_count], struct sentence sen[periods], and struct word word[spaces + periods] to build parts of the document. When the get_document function finishes, these stack arrays are destroyed, but you're still pointing to them in the returned document struct. This leads to undefined behavior—random values like 7428912 for word_count and the EXC_BAD_ACCESS crash are direct results of accessing memory that's no longer valid.

2. Incorrect Counting & Array Indexing

Your logic for counting sentences/words and assigning indices is flawed:

  • When a paragraph ends without a trailing newline, your sentence count is undercounted.
  • You're resetting indices like parano at the wrong times, leading to accidental overwrites or out-of-bounds access.
  • The way you split words/sentences/paragraphs doesn't handle edge cases (like trailing spaces or final sentences properly).

3. Unvalidated Memory Allocation

While you're using malloc, you don't check if allocations succeed, which can lead to null pointer dereferences in rare cases.


Fixed Implementation of get_document

Instead of relying on stack arrays, we'll build the document structure directly using dynamically allocated memory, and split the text in a clear, hierarchical way (paragraphs → sentences → words):

struct document get_document(char* text) {
    struct document doc;
    doc.paragraph_count = 0;

    // First pass: count total paragraphs (split by '\n')
    char* p_ptr = text;
    while (*p_ptr) {
        char* newline = strchr(p_ptr, '\n');
        if (!newline) newline = p_ptr + strlen(p_ptr);
        doc.paragraph_count++;
        p_ptr = (*newline == '\n') ? newline + 1 : newline;
        if (*p_ptr == '\0') break;
    }

    // Allocate memory for paragraphs
    doc.data = malloc(doc.paragraph_count * sizeof(struct paragraph));
    if (!doc.data) exit(EXIT_FAILURE);

    // Process each paragraph
    p_ptr = text;
    for (int p_idx = 0; p_idx < doc.paragraph_count; p_idx++) {
        struct paragraph* curr_para = &doc.data[p_idx];
        curr_para->sentence_count = 0;
        char* para_end = strchr(p_ptr, '\n');
        if (!para_end) para_end = p_ptr + strlen(p_ptr);

        // Count sentences in current paragraph (split by '.')
        char* s_ptr = p_ptr;
        while (s_ptr < para_end) {
            char* period = strchr(s_ptr, '.');
            if (!period || period >= para_end) break;
            curr_para->sentence_count++;
            s_ptr = period + 1;
        }

        // Allocate memory for sentences in this paragraph
        curr_para->data = malloc(curr_para->sentence_count * sizeof(struct sentence));
        if (!curr_para->data) exit(EXIT_FAILURE);

        // Process each sentence in the paragraph
        s_ptr = p_ptr;
        for (int s_idx = 0; s_idx < curr_para->sentence_count; s_idx++) {
            struct sentence* curr_sen = &curr_para->data[s_idx];
            curr_sen->word_count = 0;
            char* sen_end = strchr(s_ptr, '.');
            if (!sen_end) sen_end = para_end;

            // Count words in current sentence (split by ' ')
            char* w_ptr = s_ptr;
            while (w_ptr < sen_end) {
                // Skip leading spaces
                while (w_ptr < sen_end && *w_ptr == ' ') w_ptr++;
                if (w_ptr >= sen_end) break;
                curr_sen->word_count++;
                // Find next space or end of sentence
                w_ptr = strchr(w_ptr, ' ') ?: sen_end;
            }

            // Allocate memory for words in this sentence
            curr_sen->data = malloc(curr_sen->word_count * sizeof(struct word));
            if (!curr_sen->data) exit(EXIT_FAILURE);

            // Process each word in the sentence
            w_ptr = s_ptr;
            for (int w_idx = 0; w_idx < curr_sen->word_count; w_idx++) {
                struct word* curr_word = &curr_sen->data[w_idx];
                // Skip leading spaces
                while (w_ptr < sen_end && *w_ptr == ' ') w_ptr++;
                // Find end of word
                char* word_end = strchr(w_ptr, ' ') ?: sen_end;
                int word_len = word_end - w_ptr;

                // Allocate and copy word data
                curr_word->data = malloc(word_len + 1);
                if (!curr_word->data) exit(EXIT_FAILURE);
                strncpy(curr_word->data, w_ptr, word_len);
                curr_word->data[word_len] = '\0';

                w_ptr = word_end + 1;
            }

            s_ptr = sen_end + 1;
        }

        p_ptr = para_end + 1;
    }

    return doc;
}

Why This Fixes Your Problems

  1. No More Stack Memory Issues: We build the entire document structure using dynamically allocated memory (malloc) directly within the document, paragraph, and sentence structs. This memory persists after get_document returns.
  2. Accurate Counting: We count each level (paragraphs, sentences, words) explicitly before allocating memory, ensuring we never access array indices beyond their allocated size.
  3. Clear Hierarchical Processing: Splitting the text one level at a time makes the logic easier to follow and debug, eliminating the index confusion from your original code.

内容的提问来源于stack exchange,提问作者Divyanshu Varma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 18:48:16