You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用C实现目录条目流读取文件名?适配超1亿文件场景

Streaming Directory Entry Reading in C & CSV Generation (For 100M+ Files on External Drives)

Hey there! Since Node.js is waiting for the updated libuv release to support streaming directory reads, let's put together a practical C solution that handles streaming directory entries, generates the required <filename>,<timestamp> CSV, and is optimized for the extreme scenario of external drives storing over 100 million files.


Core Idea

The key to handling massive directories is avoiding loading all entries into memory at once. Instead, we'll read one entry at a time, process its timestamp immediately, and write straight to the CSV—keeping memory usage consistently low, even for 100M+ files. We'll use POSIX standard functions first (for Linux/macOS), then cover Windows adaptations.


Step-by-Step Implementation (POSIX Environments)

1. Basic Streaming Directory Read

We'll use opendir()/readdir() to stream entries, no in-memory caching of the full directory list. Here's a working code example:

#include <stdio.h>
#include <dirent.h>
#include <sys/stat.h>
#include <time.h>
#include <string.h>
#include <errno.h>
#include <limits.h>

#define OUTPUT_CSV "file_timestamps.csv"

// Fetch a file's modification timestamp (Unix epoch in seconds)
time_t get_modification_time(const char *dir_path, const char *filename) {
    char full_path[PATH_MAX];
    int snprintf_result = snprintf(full_path, sizeof(full_path), "%s/%s", dir_path, filename);
    if (snprintf_result >= sizeof(full_path) || snprintf_result < 0) {
        fprintf(stderr, "Path too long for file: %s\n", filename);
        return -1;
    }

    struct stat stat_buffer;
    if (stat(full_path, &stat_buffer) == -1) {
        fprintf(stderr, "Failed to fetch stat for %s: %s\n", full_path, strerror(errno));
        return -1;
    }
    return stat_buffer.st_mtime;
}

// Escape filename for CSV (handles commas, quotes, newlines)
void escape_csv_filename(const char *input, char *output, size_t output_size) {
    size_t input_len = strlen(input);
    size_t output_idx = 0;

    // Wrap in quotes if special characters exist
    int needs_quotes = strchr(input, ',') || strchr(input, '"') || strchr(input, '\n');
    if (needs_quotes && output_idx < output_size - 1) {
        output[output_idx++] = '"';
    }

    for (size_t i = 0; i < input_len && output_idx < output_size - 1; i++) {
        if (input[i] == '"') {
            // Escape double quote with two double quotes
            if (output_idx + 1 < output_size - 1) {
                output[output_idx++] = '"';
                output[output_idx++] = '"';
            } else {
                break;
            }
        } else {
            output[output_idx++] = input[i];
        }
    }

    if (needs_quotes && output_idx < output_size - 1) {
        output[output_idx++] = '"';
    }
    output[output_idx] = '\0';
}

int main(int argc, char *argv[]) {
    if (argc != 2) {
        fprintf(stderr, "Usage: %s <target-directory>\n", argv[0]);
        return EXIT_FAILURE;
    }
    const char *target_dir = argv[1];

    // Open target directory
    DIR *dir_stream = opendir(target_dir);
    if (!dir_stream) {
        fprintf(stderr, "Failed to open directory %s: %s\n", target_dir, strerror(errno));
        return EXIT_FAILURE;
    }

    // Open CSV output file
    FILE *csv_file = fopen(OUTPUT_CSV, "w+");
    if (!csv_file) {
        fprintf(stderr, "Failed to create CSV file %s: %s\n", OUTPUT_CSV, strerror(errno));
        closedir(dir_stream);
        return EXIT_FAILURE;
    }
    // Write CSV header
    fprintf(csv_file, "filename,timestamp\n");

    struct dirent *directory_entry;
    unsigned long processed_count = 0;
    // Stream entries one by one
    while ((directory_entry = readdir(dir_stream)) != NULL) {
        // Skip . and .. directories
        if (strcmp(directory_entry->d_name, ".") == 0 || strcmp(directory_entry->d_name, "..") == 0) {
            continue;
        }

        time_t mtime = get_modification_time(target_dir, directory_entry->d_name);
        if (mtime == -1) {
            continue; // Skip files we can't access
        }

        // Escape filename for safe CSV writing
        char escaped_name[PATH_MAX * 2]; // Extra space for escapes
        escape_csv_filename(directory_entry->d_name, escaped_name, sizeof(escaped_name));

        // Write to CSV
        fprintf(csv_file, "%s,%ld\n", escaped_name, (long)mtime);

        // Optional: Flush buffer every 1000 entries to avoid IO backpressure on external drives
        processed_count++;
        if (processed_count % 1000 == 0) {
            fflush(csv_file);
        }
    }

    // Check if readdir exited due to an error
    if (errno != 0) {
        fprintf(stderr, "Error reading directory: %s\n", strerror(errno));
    }

    // Clean up resources
    fclose(csv_file);
    closedir(dir_stream);

    printf("Successfully processed %lu files. CSV saved to %s\n", processed_count, OUTPUT_CSV);
    return EXIT_SUCCESS;
}

2. Windows Adaptation

For Windows, replace opendir()/readdir() with FindFirstFileW()/FindNextFileW() (use wide characters to handle non-ASCII filenames). To get timestamps, use GetFileTime() and convert it to a Unix timestamp. The core streaming logic remains the same—process one entry at a time, no full directory cache.


Critical Optimizations for 100M+ File Scenarios

External drives have slower IO, so these tweaks are non-negotiable:

  • Disable Directory Prefetch: On Linux, use fcntl() to set O_DIRECTORY and disable kernel prefetching, which would otherwise try to load millions of entries into RAM.
  • Batch IO Flushes: As shown in the code, flush the CSV buffer every 1000-10000 entries—this balances IO performance and memory usage.
  • Asynchronous IO: For maximum throughput, use Linux io_uring or Windows IOCP to handle directory reads and CSV writes asynchronously, avoiding blocking on slow external drive IO.
  • Skip Sorting: Never sort directory entries unless absolutely necessary—sorting requires loading all entries into memory, which is impossible for 100M+ files.
  • Robust Error Handling: Add retry logic for IO errors (common with external drives) and log failed entries instead of crashing.

CSV Best Practices

  • Filename Escaping: Always handle commas, quotes, and newlines in filenames as shown—otherwise your CSV will be corrupted.
  • Timestamp Flexibility: The example uses Unix epoch seconds. If you need human-readable timestamps, use strftime() to convert time_t to YYYY-MM-DD HH:MM:SS (just watch out for timezone differences).
  • Large File Support: Ensure your external drive uses a filesystem that supports huge files (EXT4, NTFS, APFS—avoid FAT32 which has 4GB file limits).

内容的提问来源于stack exchange,提问作者Lance Pollard

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 07:54:41