如何用C实现目录条目流读取文件名?适配超1亿文件场景
Hey there! Since Node.js is waiting for the updated libuv release to support streaming directory reads, let's put together a practical C solution that handles streaming directory entries, generates the required <filename>,<timestamp> CSV, and is optimized for the extreme scenario of external drives storing over 100 million files.
Core Idea
The key to handling massive directories is avoiding loading all entries into memory at once. Instead, we'll read one entry at a time, process its timestamp immediately, and write straight to the CSV—keeping memory usage consistently low, even for 100M+ files. We'll use POSIX standard functions first (for Linux/macOS), then cover Windows adaptations.
Step-by-Step Implementation (POSIX Environments)
1. Basic Streaming Directory Read
We'll use opendir()/readdir() to stream entries, no in-memory caching of the full directory list. Here's a working code example:
#include <stdio.h> #include <dirent.h> #include <sys/stat.h> #include <time.h> #include <string.h> #include <errno.h> #include <limits.h> #define OUTPUT_CSV "file_timestamps.csv" // Fetch a file's modification timestamp (Unix epoch in seconds) time_t get_modification_time(const char *dir_path, const char *filename) { char full_path[PATH_MAX]; int snprintf_result = snprintf(full_path, sizeof(full_path), "%s/%s", dir_path, filename); if (snprintf_result >= sizeof(full_path) || snprintf_result < 0) { fprintf(stderr, "Path too long for file: %s\n", filename); return -1; } struct stat stat_buffer; if (stat(full_path, &stat_buffer) == -1) { fprintf(stderr, "Failed to fetch stat for %s: %s\n", full_path, strerror(errno)); return -1; } return stat_buffer.st_mtime; } // Escape filename for CSV (handles commas, quotes, newlines) void escape_csv_filename(const char *input, char *output, size_t output_size) { size_t input_len = strlen(input); size_t output_idx = 0; // Wrap in quotes if special characters exist int needs_quotes = strchr(input, ',') || strchr(input, '"') || strchr(input, '\n'); if (needs_quotes && output_idx < output_size - 1) { output[output_idx++] = '"'; } for (size_t i = 0; i < input_len && output_idx < output_size - 1; i++) { if (input[i] == '"') { // Escape double quote with two double quotes if (output_idx + 1 < output_size - 1) { output[output_idx++] = '"'; output[output_idx++] = '"'; } else { break; } } else { output[output_idx++] = input[i]; } } if (needs_quotes && output_idx < output_size - 1) { output[output_idx++] = '"'; } output[output_idx] = '\0'; } int main(int argc, char *argv[]) { if (argc != 2) { fprintf(stderr, "Usage: %s <target-directory>\n", argv[0]); return EXIT_FAILURE; } const char *target_dir = argv[1]; // Open target directory DIR *dir_stream = opendir(target_dir); if (!dir_stream) { fprintf(stderr, "Failed to open directory %s: %s\n", target_dir, strerror(errno)); return EXIT_FAILURE; } // Open CSV output file FILE *csv_file = fopen(OUTPUT_CSV, "w+"); if (!csv_file) { fprintf(stderr, "Failed to create CSV file %s: %s\n", OUTPUT_CSV, strerror(errno)); closedir(dir_stream); return EXIT_FAILURE; } // Write CSV header fprintf(csv_file, "filename,timestamp\n"); struct dirent *directory_entry; unsigned long processed_count = 0; // Stream entries one by one while ((directory_entry = readdir(dir_stream)) != NULL) { // Skip . and .. directories if (strcmp(directory_entry->d_name, ".") == 0 || strcmp(directory_entry->d_name, "..") == 0) { continue; } time_t mtime = get_modification_time(target_dir, directory_entry->d_name); if (mtime == -1) { continue; // Skip files we can't access } // Escape filename for safe CSV writing char escaped_name[PATH_MAX * 2]; // Extra space for escapes escape_csv_filename(directory_entry->d_name, escaped_name, sizeof(escaped_name)); // Write to CSV fprintf(csv_file, "%s,%ld\n", escaped_name, (long)mtime); // Optional: Flush buffer every 1000 entries to avoid IO backpressure on external drives processed_count++; if (processed_count % 1000 == 0) { fflush(csv_file); } } // Check if readdir exited due to an error if (errno != 0) { fprintf(stderr, "Error reading directory: %s\n", strerror(errno)); } // Clean up resources fclose(csv_file); closedir(dir_stream); printf("Successfully processed %lu files. CSV saved to %s\n", processed_count, OUTPUT_CSV); return EXIT_SUCCESS; }
2. Windows Adaptation
For Windows, replace opendir()/readdir() with FindFirstFileW()/FindNextFileW() (use wide characters to handle non-ASCII filenames). To get timestamps, use GetFileTime() and convert it to a Unix timestamp. The core streaming logic remains the same—process one entry at a time, no full directory cache.
Critical Optimizations for 100M+ File Scenarios
External drives have slower IO, so these tweaks are non-negotiable:
- Disable Directory Prefetch: On Linux, use
fcntl()to setO_DIRECTORYand disable kernel prefetching, which would otherwise try to load millions of entries into RAM. - Batch IO Flushes: As shown in the code, flush the CSV buffer every 1000-10000 entries—this balances IO performance and memory usage.
- Asynchronous IO: For maximum throughput, use Linux
io_uringor Windows IOCP to handle directory reads and CSV writes asynchronously, avoiding blocking on slow external drive IO. - Skip Sorting: Never sort directory entries unless absolutely necessary—sorting requires loading all entries into memory, which is impossible for 100M+ files.
- Robust Error Handling: Add retry logic for IO errors (common with external drives) and log failed entries instead of crashing.
CSV Best Practices
- Filename Escaping: Always handle commas, quotes, and newlines in filenames as shown—otherwise your CSV will be corrupted.
- Timestamp Flexibility: The example uses Unix epoch seconds. If you need human-readable timestamps, use
strftime()to converttime_ttoYYYY-MM-DD HH:MM:SS(just watch out for timezone differences). - Large File Support: Ensure your external drive uses a filesystem that supports huge files (EXT4, NTFS, APFS—avoid FAT32 which has 4GB file limits).
内容的提问来源于stack exchange,提问作者Lance Pollard

