You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C语言处理100GB大文件内存优化咨询:避免大内存分配

大文件行处理优化方案(无需大量内存)

核心思路是不预存所有行,利用行模式重复(32767行后重复)的特性,只缓存前32767行,之后直接复用缓存内容,同时分两次遍历输入文件:第一次缓存模式行,第二次按需拼接写入。

具体实现步骤

1. 缓存模式行(前32767行)

  • 打开输入文件,逐行读取前32767行,用动态数组缓存每行内容(每行单独分配内存,避免二维数组的大内存占用)。
  • 保留每行的换行符,确保缓存内容与原文件完全一致。
  • 读取完成后关闭输入文件。

2. 按需拼接并写入输出文件

  • 重新打开输入文件,同时创建并打开输出文件。
  • 逐行读取输入文件,对每行做如下处理:
    • 若为前32767行内的目标行,直接从缓存中取出对应内容拼接;
    • 若为32767行之后的目标行,计算其在模式中的索引((行号-1) % 32767,适配行号从1开始的场景),从缓存中取对应内容拼接。
  • 每完成一段目标内容的拼接,立即写入输出文件,无需暂存所有拼接结果。

示例代码片段

#include <stdio.h>
#include <stdlib.h>
#include <string.h>

#define PATTERN_LINE_COUNT 32767

// 存储单行缓存的结构体
typedef struct {
    char *content;
    size_t len;
} LineCache;

int main() {
    FILE *in_fp = fopen("input.txt", "r");
    FILE *out_fp = fopen("output.txt", "w");
    if (!in_fp || !out_fp) {
        perror("File open failed");
        return 1;
    }

    // 第一步:缓存前32767行
    LineCache *cache = malloc(sizeof(LineCache) * PATTERN_LINE_COUNT);
    if (!cache) {
        perror("Malloc cache failed");
        fclose(in_fp);
        fclose(out_fp);
        return 1;
    }

    char *line_buf = NULL;
    size_t buf_cap = 0;
    ssize_t read_len;
    int line_idx = 0;

    while (line_idx < PATTERN_LINE_COUNT && (read_len = getline(&line_buf, &buf_cap, in_fp)) != -1) {
        cache[line_idx].len = read_len;
        cache[line_idx].content = malloc(read_len + 1);
        if (!cache[line_idx].content) {
            perror("Malloc line content failed");
            // 释放已分配的缓存资源
            for (int i = 0; i < line_idx; i++) free(cache[i].content);
            free(cache);
            free(line_buf);
            fclose(in_fp);
            fclose(out_fp);
            return 1;
        }
        memcpy(cache[line_idx].content, line_buf, read_len);
        cache[line_idx].content[read_len] = '\0';
        line_idx++;
    }
    free(line_buf);
    fclose(in_fp);

    // 第二步:重新读取输入文件,按需拼接写入
    in_fp = fopen("input.txt", "r");
    if (!in_fp) {
        perror("Reopen input file failed");
        // 释放缓存资源
        for (int i = 0; i < PATTERN_LINE_COUNT; i++) free(cache[i].content);
        free(cache);
        fclose(out_fp);
        return 1;
    }

    // 示例目标行列表:可根据需求修改
    int target_lines[] = {1, 32768, 32769, 65535};
    int target_count = sizeof(target_lines)/sizeof(target_lines[0]);
    int current_line = 0;

    line_buf = NULL;
    buf_cap = 0;
    while ((read_len = getline(&line_buf, &buf_cap, in_fp)) != -1) {
        current_line++;
        // 检查当前行是否为目标行
        for (int i = 0; i < target_count; i++) {
            if (current_line == target_lines[i]) {
                int cache_idx = (current_line - 1) % PATTERN_LINE_COUNT;
                fwrite(cache[cache_idx].content, 1, cache[cache_idx].len, out_fp);
                // 如需添加拼接分隔符,可在此处写入
                break;
            }
        }
    }

    // 释放所有资源
    free(line_buf);
    fclose(in_fp);
    fclose(out_fp);
    for (int i = 0; i < PATTERN_LINE_COUNT; i++) free(cache[i].content);
    free(cache);

    return 0;
}

关键优化点

  • 内存占用可控:仅缓存32767行,内存占用仅与单行长有关,远低于存储100GB文件所有行的需求。
  • 复用模式缓存:利用行重复特性,后续行直接调用缓存内容,无需重复解析原文件。
  • 边处理边写入:拼接完成后立即写入输出文件,避免暂存大量拼接结果,进一步压缩内存开销。

内容的提问来源于stack exchange,提问作者John Black

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 00:22:29