如何在Windows中预预热(Pre-warm)磁盘缓存以加速多文件读取
Hey there, let's break down how to tackle this pre-warming disk cache problem for your multi-file directory on Windows—since you've already got the theory down, let's dive into practical, proven approaches and their tradeoffs, starting with core principles to avoid blind testing.
为什么默认预读不够用?
Windows' built-in prefetch mechanism is designed for single files: when you open a file with the FILE_FLAG_SEQUENTIAL_SCAN flag, the system proactively pulls the next few megabytes into cache. But it won't automatically traverse and preload thousands of files in a directory—so we need to manually trigger this behavior.
你提出的两种方案可行性分析
1. Low-priority background thread: Open → (read a tiny chunk) → Close files
This approach is fully feasible and the lowest-effort option, with a few key details to optimize:
- Always add the
FILE_FLAG_SEQUENTIAL_SCANflag when opening files. This tells Windows you'll read sequentially, making it more aggressive about preloading subsequent content (usually 64KB to several MB, depending on system settings). - Don't just open and close—this might not trigger prefetch. Instead, use
ReadFileto grab the first 1KB of each file (you don't even need to process the data). This small read kicks off the system's prefetch logic, automatically caching the rest of the file to RAM. - Set the background thread's priority to
THREAD_PRIORITY_LOWESTorTHREAD_PRIORITY_BELOW_NORMALviaSetThreadPriority. This ensures it only uses I/O and CPU when your main thread is idle, so it won't interfere with your tool's burst-mode tasks. - Use efficient directory enumeration like
FindFirstFileExinstead of basicFindFirstFileto skip unnecessary file attributes and speed up traversal.
2. Memory-mapped files locked to RAM
This is technically possible but not cost-effective for your use case, here's why:
VirtualLock(used to pin mapped memory to RAM) has strict quota limits: by default, 32-bit processes can lock only 20MB, and 64-bit processes have higher but still constrained limits. Locking 10GB would require adjusting the process'sSeLockMemoryPrivilegepermission, adding complexity and potential security risks.- Pinning memory reserves physical RAM that the system can't reallocate to other processes, even if idle. Letting Windows' cache manager handle this is far more flexible—it will automatically reclaim unused cache when memory is tight, whereas pinned memory stays locked.
- While this guarantees files are 100% in RAM, the implementation is way more complex than the first option, and the performance gain for your burst I/O pattern is minimal compared to the effort.
Optimizations tailored to your burst I/O pattern
Since your tool alternates between I/O and CPU tasks, you can leverage CPU idle time to preload the next batch of files, maximizing resource utilization:
- Use asynchronous I/O (with
OVERLAPPEDandReadFileEx) to batch prefetch files. While your main thread is handling CPU tasks, the background thread can initiate multiple async pre-reads—by the time CPU processing finishes, the next set of files will already be in cache. - Enumerate files in physical disk order: On NTFS, each file has a
FileIndex(retrieved viaGetFileInformationByHandle) that maps to its physical location on disk. Sorting by this index minimizes disk seek time, making prefetch far more efficient. - Don't preload all 10GB at once: Split files into batches (e.g., 1GB per batch) aligned with your main thread's processing rhythm, avoiding overwhelming the I/O subsystem.
Key principle reminders
- Windows' file cache is global—prefetched content stays in cache until memory is low and the system reclaims it. As long as your tool accesses the files soon after prewarming, the cache will be effective.
- The core of prewarming is triggering the system's prefetch logic, not manually loading all data into memory. Letting Windows manage the cache is more flexible and efficient than pinning memory yourself.
- Low-priority prefetch threads won't steal I/O resources from your main thread—Windows' I/O scheduler prioritizes high-priority requests, so your tool's burst I/O will always take precedence.
内容的提问来源于stack exchange,提问作者Krumelur

