排查HackerRank文档结构化程序的EXC_BAD_ACCESS及word_count异常
Let's break down exactly what's causing your weird word_count value and the crash, then fix those issues step by step. You're spot-on that array out-of-bounds access is the root problem—here's where things go wrong and how to fix them:
Key Issues in Your Code
1. Local Stack Arrays Have Expired Lifetimes
Your code uses local stack arrays like struct paragraph para[doc.paragraph_count], struct sentence sen[periods], and struct word word[spaces + periods] to build parts of the document. When the get_document function finishes, these stack arrays are destroyed, but you're still pointing to them in the returned document struct. This leads to undefined behavior—random values like 7428912 for word_count and the EXC_BAD_ACCESS crash are direct results of accessing memory that's no longer valid.
2. Incorrect Counting & Array Indexing
Your logic for counting sentences/words and assigning indices is flawed:
- When a paragraph ends without a trailing newline, your sentence count is undercounted.
- You're resetting indices like
paranoat the wrong times, leading to accidental overwrites or out-of-bounds access. - The way you split words/sentences/paragraphs doesn't handle edge cases (like trailing spaces or final sentences properly).
3. Unvalidated Memory Allocation
While you're using malloc, you don't check if allocations succeed, which can lead to null pointer dereferences in rare cases.
Fixed Implementation of get_document
Instead of relying on stack arrays, we'll build the document structure directly using dynamically allocated memory, and split the text in a clear, hierarchical way (paragraphs → sentences → words):
struct document get_document(char* text) { struct document doc; doc.paragraph_count = 0; // First pass: count total paragraphs (split by '\n') char* p_ptr = text; while (*p_ptr) { char* newline = strchr(p_ptr, '\n'); if (!newline) newline = p_ptr + strlen(p_ptr); doc.paragraph_count++; p_ptr = (*newline == '\n') ? newline + 1 : newline; if (*p_ptr == '\0') break; } // Allocate memory for paragraphs doc.data = malloc(doc.paragraph_count * sizeof(struct paragraph)); if (!doc.data) exit(EXIT_FAILURE); // Process each paragraph p_ptr = text; for (int p_idx = 0; p_idx < doc.paragraph_count; p_idx++) { struct paragraph* curr_para = &doc.data[p_idx]; curr_para->sentence_count = 0; char* para_end = strchr(p_ptr, '\n'); if (!para_end) para_end = p_ptr + strlen(p_ptr); // Count sentences in current paragraph (split by '.') char* s_ptr = p_ptr; while (s_ptr < para_end) { char* period = strchr(s_ptr, '.'); if (!period || period >= para_end) break; curr_para->sentence_count++; s_ptr = period + 1; } // Allocate memory for sentences in this paragraph curr_para->data = malloc(curr_para->sentence_count * sizeof(struct sentence)); if (!curr_para->data) exit(EXIT_FAILURE); // Process each sentence in the paragraph s_ptr = p_ptr; for (int s_idx = 0; s_idx < curr_para->sentence_count; s_idx++) { struct sentence* curr_sen = &curr_para->data[s_idx]; curr_sen->word_count = 0; char* sen_end = strchr(s_ptr, '.'); if (!sen_end) sen_end = para_end; // Count words in current sentence (split by ' ') char* w_ptr = s_ptr; while (w_ptr < sen_end) { // Skip leading spaces while (w_ptr < sen_end && *w_ptr == ' ') w_ptr++; if (w_ptr >= sen_end) break; curr_sen->word_count++; // Find next space or end of sentence w_ptr = strchr(w_ptr, ' ') ?: sen_end; } // Allocate memory for words in this sentence curr_sen->data = malloc(curr_sen->word_count * sizeof(struct word)); if (!curr_sen->data) exit(EXIT_FAILURE); // Process each word in the sentence w_ptr = s_ptr; for (int w_idx = 0; w_idx < curr_sen->word_count; w_idx++) { struct word* curr_word = &curr_sen->data[w_idx]; // Skip leading spaces while (w_ptr < sen_end && *w_ptr == ' ') w_ptr++; // Find end of word char* word_end = strchr(w_ptr, ' ') ?: sen_end; int word_len = word_end - w_ptr; // Allocate and copy word data curr_word->data = malloc(word_len + 1); if (!curr_word->data) exit(EXIT_FAILURE); strncpy(curr_word->data, w_ptr, word_len); curr_word->data[word_len] = '\0'; w_ptr = word_end + 1; } s_ptr = sen_end + 1; } p_ptr = para_end + 1; } return doc; }
Why This Fixes Your Problems
- No More Stack Memory Issues: We build the entire document structure using dynamically allocated memory (
malloc) directly within thedocument,paragraph, andsentencestructs. This memory persists afterget_documentreturns. - Accurate Counting: We count each level (paragraphs, sentences, words) explicitly before allocating memory, ensuring we never access array indices beyond their allocated size.
- Clear Hierarchical Processing: Splitting the text one level at a time makes the logic easier to follow and debug, eliminating the index confusion from your original code.
内容的提问来源于stack exchange,提问作者Divyanshu Varma

