C语言多线程读取类JSON大文件的技术问题咨询
C语言多线程读取大文件的问题解答
问题背景
在C语言中尝试用多线程读取约100MB的大文件,按文件大小分块时部分块会起始/结束于行中间,尝试调整块大小但因各行长度不一致无法解决。当前方案是将块的起始和结束位置调整至行边界,线程数大于行数时跳过对应线程,但偶尔出现漏读、重复读取或行内容被拆分的问题,同时有几个疑问需要解答。
1. 分块读取+调整行边界的方案是否可行?还是多线程仅适合行处理?
这种分块读取并对齐行边界的方案完全可行,但你的代码逻辑存在缺陷,导致了漏读/重复读取的问题:
- 代码中先调整了
start和end的位置,之后又执行thread_data[i].start = last_end,这会覆盖之前的行边界调整结果,容易导致块重叠或间隙。 - 调整
end位置时的逻辑有问题:当fgetc读到EOF时,thread_data[i].end = ftell(file)已经是文件末尾,不需要额外end++;如果读到\n,ftell的位置已经是\n之后,也不需要调整。 - 主线程在调整所有块的边界时,多次对同一个
FILE*执行fseek和fgetc,如果后续线程复用这个FILE*,可能因为文件指针的共享导致定位混乱。
建议优化逻辑:先基于last_end确定当前块的初始start,再计算初始end,然后调整end到下一个行尾,最后更新last_end。这样能保证块之间完全连续无重叠。
2. 不同线程同时读取同一文件是否存在问题?
在POSIX标准下,stdio库的函数(如fseek、fgetc、fgets)是线程安全的——每个操作会对FILE结构体内部的锁进行加锁/解锁,不会出现数据竞争导致的崩溃。但共享同一个FILE*会导致线程间的串行化:一个线程在操作文件指针时,其他线程会被阻塞,无法真正并行读取。
更高效的做法是:每个线程单独打开同一个文件(使用fopen),这样每个线程有独立的文件指针,互不干扰,能真正实现并行读取。
3. 是否应该使用JSON库而非假设每行都是类JSON对象?
强烈建议使用成熟的JSON库,原因如下:
- 你的“类JSON”只是把逗号换成分号,但本质还是JSON结构,用JSON库能直接解析,无需自己处理语法细节。
- 假设每行是完整JSON对象的可靠性极低:如果数据中出现字符串包含换行、转义字符等情况,自己的解析逻辑会直接失效,而JSON库能正确处理这些边界情况。
- 效率上,成熟的JSON库(如cJSON、Jansson)经过优化,性能不会比自己写的简单解析差,同时能保证正确性和健壮性。
附用户提供的代码
块划分逻辑代码
long chunk_size = file_size / num_threads; pthread_t threads[num_threads]; ThreadData thread_data[num_threads]; long last_end = 0; for (uint32_t i = 0; i < num_threads; ++i) { thread_data[i].stats = stats; thread_data[i].thread_tweets = NULL; thread_data[i].failed = 0; thread_data[i].file = file; thread_data[i].start = i * chunk_size; thread_data[i].end = (i == num_threads - 1)? file_size : (i + 1) * chunk_size; if (i > 0) { if (thread_data[i].end < thread_data[i - 1].start) { thread_data[i].failed = 1; continue; } } int ch; // Adjust start position to the beginning of the next line if (!is_start_at_line_boundary(file, thread_data[i].start)) { fseek(file, thread_data[i].start, SEEK_SET); while ((ch = fgetc(file))!= '\n' && ch!= EOF); thread_data[i].start = ftell(file); } // Adjust end position to the end of the line fseek(file, thread_data[i].end, SEEK_SET); while ((ch = fgetc(file))!= '\n' && ch!= EOF); thread_data[i].end = ftell(file); if (ch!= '\n' && ch!= EOF) { thread_data[i].end++; } // If they coincide, the chunk was inside a line and the thread shoudnt run if (thread_data[i].end == thread_data[i].start) { thread_data[i].failed = 1; continue; } if (i > 0) { thread_data[i].start = last_end; } if (pthread_create(&threads[i], NULL, read_file_chunk, &thread_data[i])) { fprintf(stderr, "Error creating thread\n"); exit(EXIT_FAILURE); } last_end = thread_data[i].end; }
行边界判断函数
int is_start_at_line_boundary(FILE *file, long start) { if (start == 0) { return 1; // Start of the file } fseek(file, start - 1, SEEK_SET); if (fgetc(file) == '\n') { return 1; // Start is at the beginning of a line } return 0; }
示例行格式
{"created_at": "2020-01-14 12:00:00"; "hashtags": ["A", "B"]; "id": 546542; "uid": 1500}
内容的提问来源于stack exchange,提问作者thuruk9
相关产品推荐
相关产品推荐

