向fwrite()传入二维数组列数据的最优方法与最佳实践
当数据以二维数组形式存储,且需要将每列数据单独保存到文件时,调用fwrite()是否存在更高效的实现方式?
以下示例代码模拟了从磁盘读取数据到缓冲区,处理后解析为8行×3列的processedData数组,随后通过循环为每列创建文件、逐列拷贝数据到临时缓冲区再写入,最终生成3个各8字节的文件。当前实现需要额外的内存分配和数据拷贝,有没有更高效的方式直接将列数据传入fwrite()?
#include <stdint.h> #include <stdio.h> #include <stdlib.h> #include <string.h> int main(int argc, char *argv[]) { uint8_t t, p, s; // Assume inputFile is a buffer with the contents of a file from disk uint8_t buffer[24] = {2, 1, 3, 3, 2, 4, 4, 3, 5, 5, 4, 6, 6, 5, 7, 7, 6, 8, 8, 7, 9, 9, 8, 0}; uint8_t(*processedData)[8][3] = malloc(sizeof *processedData); memset(processedData, '\0', sizeof(*processedData)); // Assume the contents of the buffer are processed in some way... for (s = 0; s < 24; s++) (*processedData)[s % 8][s % 3] = buffer[s]; // Parse the processed content by column and save in 3 separate files for (t = 0; t < 3; t++) { uint8_t *hld = (uint8_t *)calloc(8, sizeof(uint8_t)); for (p = 0; p < 8; p++) hld[p] = (*processedData)[p][t]; char *fileName = (char *)malloc(sizeof(char) * 30); sprintf(fileName, "%s-%u", "TestFile", t); FILE *tmpFile = fopen(fileName, "w+"); if (tmpFile == NULL) { fputs("File error", stderr); exit(1); } fwrite(hld, 1, 8, tmpFile); fclose(tmpFile); free(hld); } return 0; }
优化方案
1. 跳过临时缓冲区,直接逐行写入对应文件
原代码中为每列分配临时缓冲区并拷贝数据的操作完全可以省略。可以提前打开所有目标文件,然后遍历数组的每一行,将每行的对应列数据直接写入对应的文件,避免额外的内存分配和数据拷贝。
示例代码:
#include <stdint.h> #include <stdio.h> #include <stdlib.h> #include <string.h> int main(int argc, char *argv[]) { uint8_t p, s; uint8_t buffer[24] = {2, 1, 3, 3, 2, 4, 4, 3, 5, 5, 4, 6, 6, 5, 7, 7, 6, 8, 8, 7, 9, 9, 8, 0}; uint8_t(*processedData)[8][3] = malloc(sizeof *processedData); memset(processedData, '\0', sizeof(*processedData)); for (s = 0; s < 24; s++) (*processedData)[s % 8][s % 3] = buffer[s]; // 提前打开所有3个输出文件 FILE *files[3]; char fileName[30]; for (int t = 0; t < 3; t++) { sprintf(fileName, "TestFile-%u", t); files[t] = fopen(fileName, "w+"); if (files[t] == NULL) { fputs("File error", stderr); // 关闭已打开的文件,避免资源泄漏 for (int i = 0; i < t; i++) { fclose(files[i]); } exit(1); } } // 遍历每行,直接将对应列的数据写入文件 for (p = 0; p < 8; p++) { for (int t = 0; t < 3; t++) { fwrite(&(*processedData)[p][t], sizeof(uint8_t), 1, files[t]); } } // 关闭所有文件 for (int t = 0; t < 3; t++) { fclose(files[t]); } free(processedData); return 0; }
这种方式的核心是减少了内存操作开销,直接操作原数组的元素完成写入,在保持代码可读性的同时提升了效率。
2. 调整数组存储为列优先(最高效方案)
如果数据处理阶段可以适配,将数组改为列优先存储(即先存储第一列的所有元素,再存储第二列,以此类推),那么后续写入文件时可以直接用fwrite()一次性写入整列的连续内存块,不需要任何遍历或拷贝操作。
比如将processedData定义为列优先的结构:
// 3列,每列8行,列优先存储 uint8_t(*processedData)[3][8] = malloc(sizeof *processedData);
在数据处理时按列顺序填充,后续写入文件时直接:
for (int t = 0; t < 3; t++) { sprintf(fileName, "TestFile-%u", t); FILE *tmpFile = fopen(fileName, "w+"); if (tmpFile == NULL) { fputs("File error", stderr); exit(1); } // 直接写入整列的连续内存 fwrite(&(*processedData)[t][0], sizeof(uint8_t), 8, tmpFile); fclose(tmpFile); }
这种方式效率最高,因为fwrite()可以一次性完成整块内存的写入,减少了IO操作的次数,也完全消除了数据拷贝的开销。
3. 分散写入(不推荐)
部分系统提供了writev()这类系统调用,支持一次性写入多个不连续的内存块。但对于8字节的小数据来说,需要构建iovec结构体数组来存放每个列元素的地址和长度,代码复杂度高,收益远不如前两种方案,因此一般不推荐使用。
内容的提问来源于stack exchange,提问作者affluentbarnburner

