使用fread分块读取大文件时数据丢失及行数统计错误求助
问题:分块读取大文件时行数统计与Notepad++不一致
我是C新手,正在使用纽约证券交易所(NYSE)的交易数据进行量化分析。单份日数据文件约10GB,我拥有数百份此类文件。第一步我希望确保能正确分块读取数据,之后再进行数据处理。听闻Python处理速度较慢,因此尝试使用C,通过fread()分块读取单个文件。我通过统计行数来验证代码,但得到的行数与Notepad++统计的结果不一致,希望有人能帮我解决这个问题,谢谢。
数据示例
Q,14340,EUR/NZD,1.65027,1,1.6504,1 T,14340,EUR/NZD,1,1.65034,@,70,X Q,14340,AUD/NZD,1.03427,1,1.03437,1 T,14340,AUD/NZD,1,1.03432,@,70,X Q,14340,CAD/CHF,0.75142,1,0.75146,1 T,14340,CAD/CHF,1,0.75144,@,70,X Q,14340,GBP/NZD,1.90908,1,1.90927,1 T,14340,GBP/NZD,1,1.90918,@,70,X Q,14340,GBP/CHF,1.312,1,1.31208,1 T,14340,GBP/CHF,1,1.31204,@,70,X Q,13724,#6S,0.9928,12,0.9929,29
统计结果
- 我的代码统计行数:279 174 248
- Notepad++统计行数:279 485 508
错误代码
#include <iostream> #include <cstdio> using namespace std; int main() { clock_t start = clock(); char buffer[100000]="\0"; int cursor = sizeof(buffer); FILE* fp; int Judge; int offset = 0; Judge = fopen_s(&fp, "E:\\feedRec\\TFD20190227", "r"); int count = 0; int num = 0; while (1) { //read one chunk num=fread(buffer, sizeof(char), (sizeof(buffer) - 1), fp); // null terminated the buffer buffer[num] = '\0'; char* ptr=buffer; while (*ptr!='\0') { if (*ptr == '\n') { count++; } ptr++; } //Since the lines are not formatted. The buffer will end at //the middle of the line. So I need to move back the pointer //back to the beginning of the sliced line to let it be read in //the next loop. //T,14340,AUD/NZD,1,1.03432,@,70,X //if buffer end here| //pointer need to be move back to the beginning. //T,14340,AUD/NZD,1,1.03432,@,70,X //| //calculte the offset offset = 0; cursor = sizeof(buffer)-1; while (buffer[cursor] != '\n') { cursor--; offset++; } //move back the pointer if (offset > 0 && (num == sizeof(buffer) - 1)) { fseek(fp, -offset, SEEK_CUR); } //clear the buffer memset(buffer, '\0', sizeof(buffer) - 1); // deal with last chunk: if the num less than buffer size, stop if (num < sizeof(buffer)-1) { fclose(fp); break; } } //count the last chunk char* ptr = buffer; while (*ptr != '\0') { if (*ptr == '\n') { count++; } ptr++; } cout << "count: " << count << endl; clock_t end = clock(); cout << "time : " << ((double)end - start) / CLOCKS_PER_SEC << "s\n"; fclose(fp); return 0; }
问题分析与修复
你的代码存在几个关键问题,导致行数统计错误:
1. 文本模式读取的换行符转换问题
用"r"模式打开文件是文本模式,Windows系统下会自动将文件中的\r\n(CRLF)转换为\n,但fread按字节读取的特性会让这种转换破坏字节计数的准确性,进而干扰fseek的偏移计算。
修复:改用二进制模式打开文件,避免换行符自动转换:
Judge = fopen_s(&fp, "E:\\feedRec\\TFD20190227", "rb");
2. 偏移计算逻辑错误
你从sizeof(buffer)-1开始向前找\n,但如果读取的最后一块数据长度小于缓冲区大小,或者整块都没有换行符,会导致越界访问或找错位置。
修复:从实际读取的最后一个字节(num-1)开始向前查找,并添加边界检查:
offset = 0; cursor = num - 1; // 从实际读取的最后一个字节开始 while (cursor >= 0 && buffer[cursor] != '\n') { cursor--; offset++; }
3. 重复统计最后一块
循环内已经统计了当前块的换行符,循环结束后又再次统计缓冲区内容,导致最后一块行数被重复计算。
修复:删除循环外的重复统计代码。
4. 重复关闭文件
循环内已经调用fclose(fp),循环外再次调用会触发未定义行为。
修复:仅在循环结束后关闭一次文件。
修复后的完整代码
#include <iostream> #include <cstdio> #include <cstring> using namespace std; int main() { clock_t start = clock(); const int BUFFER_SIZE = 100000; char buffer[BUFFER_SIZE]; FILE* fp; int Judge; Judge = fopen_s(&fp, "E:\\feedRec\\TFD20190227", "rb"); if (Judge != 0) { cerr << "Failed to open file!" << endl; return 1; } long long count = 0; // 行数可能很大,用long long避免溢出 size_t num = 0; while (true) { // 读取一块数据 num = fread(buffer, sizeof(char), BUFFER_SIZE, fp); if (num == 0) break; // 读取结束或出错 // 统计当前块的换行符 for (size_t i = 0; i < num; ++i) { if (buffer[i] == '\n') { count++; } } // 处理块末尾的不完整行:如果没读到文件末尾,且最后一个字节不是换行符 if (num == BUFFER_SIZE && buffer[num-1] != '\n') { // 向前找最近的换行符 int offset = 0; int cursor = num - 1; while (cursor >= 0 && buffer[cursor] != '\n') { cursor--; offset++; } // 调整文件指针,让下一次读取从不完整行的开头开始 if (cursor >= 0) { fseek(fp, -(num - cursor - 1), SEEK_CUR); } // 整块都是一行时,回退整个块 else { fseek(fp, -num, SEEK_CUR); } } } // 处理文件最后一行无换行符的情况(Notepad++会统计为一行) fseek(fp, -1, SEEK_END); char last_char; if (fread(&last_char, 1, 1, fp) == 1 && last_char != '\n') { count++; } fclose(fp); cout << "count: " << count << endl; clock_t end = clock(); cout << "time : " << ((double)end - start) / CLOCKS_PER_SEC << "s\n"; return 0; }
额外说明
- 使用
long long存储行数:10GB文件的行数可能超过int最大值(约20亿),避免溢出。 - 处理无换行符的最后一行:Notepad++会将文件末尾无
\n的内容算作一行,修复后的代码添加了对应逻辑补全统计。 - 二进制模式的必要性:确保读取的字节与文件完全一致,避免文本模式的换行符转换干扰偏移计算。
内容的提问来源于stack exchange,提问作者DongNie
相关产品推荐
相关产品推荐

