C程序统计超5000数值时数据丢失问题排查求助
Hey there! Let's dig into why your program's count of values over 5000 doesn't match Excel's result. The difference of 28 (45460 vs 45432) suggests some values are slipping through the cracks in your counting logic, so let's break down the likely causes and fixes.
Most Likely Root Causes
1. Integer Overflow Issues
If your input.txt contains numbers larger than the maximum value of a 32-bit int (which is 2147483647), fscanf will read them incorrectly—often wrapping around to negative values. These negative numbers won't trigger your newData > 5000 check, so they get excluded from your count, leading to the lower number you're seeing.
2. Incomplete File Reading
Your current loop condition uses fscanf(...) != EOF, which doesn't account for failed reads (like encountering non-numeric characters in the file). If fscanf hits a non-digit, it returns 0 (not EOF), and your loop will continue running but won't update newData. This could cause you to miss valid numbers that come after the invalid character, or even count old values repeatedly.
3. Unhandled Memory Leak (Not Directly Related to Counting, But Worth Fixing)
In your InsertNode function, when you detect a duplicate key, you return 0 but don't free the node you malloc'd. This will cause memory leaks over time, so it's a good idea to fix this alongside the counting issue.
Step-by-Step Fixes
Fix 1: Use a Larger Integer Type to Avoid Overflow
Change newData to long long to handle larger numbers, and update your fscanf format string to match. We also updated the count variable to long long to prevent overflow if the total count grows very large.
Fix 2: Correct the Loop Condition for Reliable Reading
Instead of checking against EOF, verify that fscanf successfully read exactly one integer (returns 1). This ensures you only process valid numbers and stop if you hit invalid data or the end of the file. We also added a check to alert you if the loop exited due to a read error instead of reaching EOF.
Fix 3: Fix the Memory Leak in InsertNode
Free the malloc'd node when you detect a duplicate key before returning 0 to avoid unnecessary memory usage.
Modified Code with Fixes
#include<stdio.h> #include<stdlib.h> #include<limits.h> // For INT_MAX reference typedef struct Node { int key; struct Node* link; } listNode; typedef struct Head { struct Node* head; } headNode; int NodeCount = 0; long long Data_morethan5000_Count = 0; // Use long long to avoid count overflow headNode* initialize(headNode* rheadnode); int DeleteList(headNode* rheadnode); void GetData(headNode* rheadnode); int InsertNode(headNode* rheadnode, int key); void PrintResult(headNode* rheadnode); int main() { headNode* headnode = NULL; headnode = initialize(headnode); GetData(headnode); PrintResult(headnode); DeleteList(headnode); return 0; } headNode* initialize(headNode* rheadnode) { headNode* temp = (headNode*)malloc(sizeof(headNode)); temp->head = NULL; return temp; } int DeleteList(headNode* rheadnode) { listNode* p = rheadnode->head; listNode* prev = NULL; while (p != NULL) { prev = p; p = p->link; free(prev); } free(rheadnode); return 0; } void GetData(headNode* rheadnode) { int dataType; long long newData; // Handle large numbers without overflow FILE* fp = NULL; fp = fopen("input.txt", "r"); if (fp != NULL) { // Only process if we successfully read an integer while (fscanf(fp, "%lld", &newData) == 1) { // Skip insertion if number exceeds int range, but still count it if >5000 if (newData > INT_MAX || newData < INT_MIN) { printf("Warning: Number %lld exceeds int range, skipping insertion\n", newData); if (newData > 5000) { Data_morethan5000_Count++; } continue; } int key = (int)newData; dataType = InsertNode(rheadnode, key); if (newData > 5000) { Data_morethan5000_Count++; } switch (dataType) { case 0: break; case 1: NodeCount++; } } // Alert if loop exited due to invalid data, not EOF if (!feof(fp)) { printf("Warning: Encountered invalid data or read error in input.txt\n"); } fclose(fp); } else { printf("Error: Could not open input.txt\n"); } } int InsertNode(headNode* rheadnode, int key) { listNode* search, * previous; listNode* node = (listNode*)malloc(sizeof(listNode)); node->key = key; search = rheadnode->head; previous = NULL; while (search != NULL) { if (node->key < search->key) { previous = search; search = search->link; } else if (node->key == search->key) { free(node); // Fix memory leak for duplicates return 0; } else { break; } } if (previous == NULL) { node->link = rheadnode->head; rheadnode->head = node; } else { node->link = search; previous->link = node; } return 1; } void PrintResult(headNode* rheadnode) { printf("The total number of nodes: %d\n", NodeCount); printf("More than 5000 values: %lld\n", Data_morethan5000_Count); printf("Execution time: sec"); }
Additional Checks
- Open your
input.txtand scan for numbers larger than2147483647—these are the likely candidates causing the overflow. - Check for any non-numeric characters (like letters, symbols, or malformed numbers) that might be causing
fscanfto stop reading early.
内容的提问来源于stack exchange,提问作者student12343

