C语言处理100GB大文件内存优化咨询:避免大内存分配
大文件行处理优化方案(无需大量内存)
核心思路是不预存所有行,利用行模式重复(32767行后重复)的特性,只缓存前32767行,之后直接复用缓存内容,同时分两次遍历输入文件:第一次缓存模式行,第二次按需拼接写入。
具体实现步骤
1. 缓存模式行(前32767行)
- 打开输入文件,逐行读取前32767行,用动态数组缓存每行内容(每行单独分配内存,避免二维数组的大内存占用)。
- 保留每行的换行符,确保缓存内容与原文件完全一致。
- 读取完成后关闭输入文件。
2. 按需拼接并写入输出文件
- 重新打开输入文件,同时创建并打开输出文件。
- 逐行读取输入文件,对每行做如下处理:
- 若为前32767行内的目标行,直接从缓存中取出对应内容拼接;
- 若为32767行之后的目标行,计算其在模式中的索引(
(行号-1) % 32767,适配行号从1开始的场景),从缓存中取对应内容拼接。
- 每完成一段目标内容的拼接,立即写入输出文件,无需暂存所有拼接结果。
示例代码片段
#include <stdio.h> #include <stdlib.h> #include <string.h> #define PATTERN_LINE_COUNT 32767 // 存储单行缓存的结构体 typedef struct { char *content; size_t len; } LineCache; int main() { FILE *in_fp = fopen("input.txt", "r"); FILE *out_fp = fopen("output.txt", "w"); if (!in_fp || !out_fp) { perror("File open failed"); return 1; } // 第一步:缓存前32767行 LineCache *cache = malloc(sizeof(LineCache) * PATTERN_LINE_COUNT); if (!cache) { perror("Malloc cache failed"); fclose(in_fp); fclose(out_fp); return 1; } char *line_buf = NULL; size_t buf_cap = 0; ssize_t read_len; int line_idx = 0; while (line_idx < PATTERN_LINE_COUNT && (read_len = getline(&line_buf, &buf_cap, in_fp)) != -1) { cache[line_idx].len = read_len; cache[line_idx].content = malloc(read_len + 1); if (!cache[line_idx].content) { perror("Malloc line content failed"); // 释放已分配的缓存资源 for (int i = 0; i < line_idx; i++) free(cache[i].content); free(cache); free(line_buf); fclose(in_fp); fclose(out_fp); return 1; } memcpy(cache[line_idx].content, line_buf, read_len); cache[line_idx].content[read_len] = '\0'; line_idx++; } free(line_buf); fclose(in_fp); // 第二步:重新读取输入文件,按需拼接写入 in_fp = fopen("input.txt", "r"); if (!in_fp) { perror("Reopen input file failed"); // 释放缓存资源 for (int i = 0; i < PATTERN_LINE_COUNT; i++) free(cache[i].content); free(cache); fclose(out_fp); return 1; } // 示例目标行列表:可根据需求修改 int target_lines[] = {1, 32768, 32769, 65535}; int target_count = sizeof(target_lines)/sizeof(target_lines[0]); int current_line = 0; line_buf = NULL; buf_cap = 0; while ((read_len = getline(&line_buf, &buf_cap, in_fp)) != -1) { current_line++; // 检查当前行是否为目标行 for (int i = 0; i < target_count; i++) { if (current_line == target_lines[i]) { int cache_idx = (current_line - 1) % PATTERN_LINE_COUNT; fwrite(cache[cache_idx].content, 1, cache[cache_idx].len, out_fp); // 如需添加拼接分隔符,可在此处写入 break; } } } // 释放所有资源 free(line_buf); fclose(in_fp); fclose(out_fp); for (int i = 0; i < PATTERN_LINE_COUNT; i++) free(cache[i].content); free(cache); return 0; }
关键优化点
- 内存占用可控:仅缓存32767行,内存占用仅与单行长有关,远低于存储100GB文件所有行的需求。
- 复用模式缓存:利用行重复特性,后续行直接调用缓存内容,无需重复解析原文件。
- 边处理边写入:拼接完成后立即写入输出文件,避免暂存大量拼接结果,进一步压缩内存开销。
内容的提问来源于stack exchange,提问作者John Black
相关产品推荐
相关产品推荐

