百万级用户独立目录存储场景下的性能优化方案咨询
Great question—scaling file storage from a small user base to 1 million users hits a classic bottleneck when you start with a flat per-user directory structure. Most file systems struggle with directories containing hundreds of thousands of entries, as lookup times degrade significantly. Here are the most practical optimizations to fix this:
1. Multilevel Hash-Based Directory Splitting
This is the industry standard for scaling file storage without overhauling your entire setup. Instead of putting all user directories directly under /users/, split them into nested subdirectories using a hash of the user ID (or a portion of it) to distribute users evenly.
For example:
- If your user ID is
123456789, take the first two digits12and next two digits34to create a path like/users/12/34/123456789/file1.jpg - For numeric IDs, you can also use modulo operations:
userId % 100gives the first-level directory,(userId // 100) % 100gives the second level. This ensures even distribution, avoiding "hot" directories that get overcrowded.
Why this works:
- Each directory only contains up to 100 (or your chosen split size) subdirectories, keeping lookup times fast.
- You can retain your original URL structure (
{url}/users/userId/file1.jpg) by adding a backend layer that translates theuserIdto the hashed directory path before serving the file.
2. Skip Per-User Directories Altogether (Use UUIDs + Database Mapping)
If you don’t need to organize files by user directories for backend maintenance, you can eliminate user-specific directories entirely:
- Generate a unique UUID for each uploaded file (e.g.,
abc123-def456-ghi789.jpg) - Store files in a hash-split directory structure based on the UUID (e.g.,
/files/ab/c1/abc123-def456-ghi789.jpg) - Maintain a database table that maps
userId,fileName, and the actual storage path/UUID.
When a user requests {url}/users/userId/file1.jpg, your backend queries the database to find the corresponding UUID/storage path, then serves the file. This approach keeps your directory structure extremely flat and avoids any per-user directory scaling issues.
3. Switch to Object Storage
If you’re running in the cloud or can deploy self-hosted tools like MinIO, object storage is built specifically for handling massive numbers of files. Services like AWS S3, Google Cloud Storage, or self-hosted MinIO don’t rely on traditional file system directories—they use object keys (e.g., user-12345/file1.jpg) and internally distribute objects across storage nodes using hashing.
Benefits:
- Zero manual directory management; the storage system handles scaling automatically.
- Built-in high availability, replication, and CDN integration for faster file delivery.
- You can still map your original URL structure to object keys via backend routing.
4. Optimize Your File System Configuration (Last Resort)
If you must stick to a local file system, tweak its settings to better handle large directories:
- For ext4, ensure the
dir_indexfeature is enabled (most modern systems enable this by default). It uses a hash tree to index directory entries, drastically improving lookup speeds for large directories. - Consider switching to XFS, which has better performance for large-scale file storage and handles millions of directory entries more gracefully than ext4.
Note: This is a band-aid solution—hash splitting or object storage will still be more reliable long-term as your user count grows.
Key Implementation Tip
Whichever approach you choose, make sure to keep your public URL structure unchanged if needed. Add a middleware or routing layer that translates the user-facing URL ({url}/users/userId/file1.jpg) to the actual storage path. This way, your users and frontend code don’t need any changes.
内容的提问来源于stack exchange,提问作者mohsenJsh

