求递归统计Linux目录数量的最优C代码及现有代码优化方案
Great question! Since you only need to count directories (not files) and want to squeeze out more speed than your existing code (which is already 4x faster than find), let's focus on stripping out all unnecessary operations that aren't related to directory traversal and counting. The key insight is that even if you removed file-printing code, the original code still spends cycles checking non-directory entries and doing unnecessary string operations for files. Here's how to fix that:
Key Optimizations to Implement
- Eliminate all file entry handling: Remove any code that checks or processes non-directory entries entirely—we don't care about files, so skip them immediately.
- Remove all I/O operations: Printing paths (even directories) is a slow I/O-bound operation; since we only need a count, we can delete all
printf/putscalls entirely. - Simplify path management: Keep only the minimal path logic needed to recurse into subdirectories, and avoid any extra string manipulation for non-directory entries.
- Use a pointer-based counter: Pass a counter by reference instead of using a global variable (cleaner and thread-safe if needed) to track the directory count.
Optimized Code
#include <errno.h> #include <stdio.h> #include <string.h> #include <sys/types.h> #include <unistd.h> #include <dirent.h> // Recursive function to count directories void count_dirs(char *path, size_t size, unsigned long *dir_count) { DIR *dir; struct dirent *entry; size_t len = strlen(path); if (!(dir = opendir(path))) { fprintf(stderr, "Failed to open directory: %s: %s\n", path, strerror(errno)); return; } // Increment count for the current directory (*dir_count)++; while ((entry = readdir(dir)) != NULL) { char *name = entry->d_name; // Only process directory entries (skip files entirely) if (entry->d_type == DT_DIR) { // Skip . and .. to avoid infinite recursion if (!strcmp(name, ".") || !strcmp(name, "..")) continue; // Check if path will fit in our buffer if (len + strlen(name) + 2 > size) { fprintf(stderr, "Path too long: %s/%s\n", path, name); continue; } // Build the path for the subdirectory path[len] = '/'; strcpy(path + len + 1, name); // Recurse into the subdirectory count_dirs(path, size, dir_count); // Restore the original path path[len] = '\0'; } // No else branch—we completely ignore non-directory entries } closedir(dir); } int main(int argc, char *argv[]) { if (argc != 2) { fprintf(stderr, "Usage: %s <directory_path>\n", argv[0]); return 1; } char path[1024]; unsigned long dir_count = 0; // Safely copy the input path strncpy(path, argv[1], sizeof(path) - 1); path[sizeof(path) - 1] = '\0'; // Ensure null-termination count_dirs(path, sizeof(path), &dir_count); printf("Total directories: %lu\n", dir_count); return 0; }
What's Changed & Why It's Faster
- No file processing: The code now skips non-directory entries immediately, avoiding any conditional logic or string operations for files. This saves a lot of cycles, especially in directories with many files.
- Removed all printing: I/O operations are far slower than in-memory counting, so removing all
printf/putscalls eliminates a major bottleneck. - Simplified counter: Using a pointer to pass the counter avoids global variables and keeps the code clean while ensuring we track the count efficiently.
- Safer path handling: Replaced
memcpywithstrncpyto prevent buffer overflow, which is a good practice without sacrificing speed.
This optimized version will be even faster than your original code because it's doing the absolute minimum work needed to traverse directories and count them—no extra processing for files, no slow I/O, just pure directory traversal and counting.
内容的提问来源于stack exchange,提问作者zlobul

