Linux下C程序编译后缓存对I/O输入量的影响原因探究
Great question—this is a perfect example of how Linux’s page cache (disk caching) drastically impacts application I/O behavior. Let’s break down what’s happening, the underlying mechanism, and where you can dive deeper.
Why the I/O Input Count Differs
The key difference boils down to Linux’s page cache:
- In Case A, after compiling your program, the files needed to run your app and launch Sublime Text (like Sublime’s binary, its shared libraries, and even your compiled
pmultiexecutable) were already stored in the system’s memory cache from prior access (either the compilation process or previous system activity). When you ran./pmulti, most of the data needed was pulled directly from memory—no slow disk reads required, hence the lowru_inblockvalue (632). - In Case B, you cleared the system’s cache with
free && sync && echo 3 > /proc/sys/vm/drop_caches. This wiped all cached file data from memory. When you launched./pmulti, every byte needed to load Sublime Text and run your program had to be read from the physical disk, leading to the much higherru_inblockvalue (1400).
The Page Cache Mechanism Explained
Linux uses free system memory to create a page cache (also called disk cache) that stores recently accessed file content. Here’s how it works for your program:
- When any process (like
gccduring compilation, or a prior Sublime launch) reads a file from disk, the kernel copies the file’s data into memory pages in the page cache. - When your program’s second thread calls
execvpto launch Sublime Text, the kernel first checks if Sublime’s binary and dependent libraries are already in the page cache:- Cache hit: If the data is present, the kernel loads it directly from memory—no disk I/O is performed, so
ru_inblock(which counts physical disk read blocks) doesn’t increase much. - Cache miss: If the data isn’t present (like after clearing the cache), the kernel reads the data from disk into the page cache, then passes it to the process. Each disk read increments
ru_inblock.
- Cache hit: If the data is present, the kernel loads it directly from memory—no disk I/O is performed, so
- The
echo 3 > /proc/sys/vm/drop_cachescommand specifically clears:- Page cache (cached file data)
- Directory entry cache (cached file names/paths)
- Inode cache (cached file metadata)
This forces all subsequent file accesses to go straight to disk.
Where to Deepen Your Knowledge
To learn more about Linux’s caching and I/O subsystems, try these resources:
- Linux Kernel Documentation: Look for files in the kernel source tree like
Documentation/filesystems/pagecache.txt(covers page cache fundamentals) andDocumentation/sysctl/vm.txt(explainsdrop_cachesand other VM tunables). - Core Kernel Books:
- Linux Kernel Development by Robert Love: A practical guide that breaks down memory management and the page cache in easy-to-follow terms.
- Understanding the Linux Kernel by Daniel P. Bovet & Marco Cesati: A deep dive into the kernel’s storage and I/O subsystems.
- Hands-On Experiments:
- Use tools like
vmstatoriostatto monitor cache usage before/after clearingdrop_caches. - Run
strace ./pmultito trace thereadsystem calls—you’ll see which calls result in actual disk I/O vs. cache hits. - Adjust your program to measure cache hits/misses directly (using
/proc/self/stator kernel tracepoints).
- Use tools like
Your Experiment Details
Case A: No Cache Clearing
$ gcc pmulti.c -o pmulti $ ./pmulti <result> I/O Input: 632 Output: 0
Case B: After Cache Clearing
$ gcc pmulti.c -o pmulti # free && sync && echo 3 > /proc/sys/vm/drop_caches && free $ ./pmulti <result> I/O Input: 1400 Output: 0
Your Simplified Code
struct rusage usage; void *func1(void *vargp) { /* Manual System Calls Here */ return NULL; } void *func2(void *vargp) { long pid; int stat_loc; if ((pid = (long) fork()) == 0){ //Format string for execvp here char s[] = "/opt/sublime_text/sublime_text"; char* separator = " "; char* argv[64]; int argc = 0; char* tmp; argv[argc] = strtok_r(s, separator, &tmp); while( argv[argc] != NULL){ argc+=1; argv[argc] = strtok_r(NULL, separator, &tmp); } execvp(argv[0],argv); } else { waitpid(pid, & stat_loc, WUNTRACED); } return NULL; } int main() { pthread_t thread_id[2]; pthread_create(&thread_id[1], NULL, func2, NULL); pthread_create(&thread_id[0], NULL, func1, NULL); pthread_join(thread_id[0], NULL); pthread_join(thread_id[1], NULL); getrusage(RUSAGE_SELF, &usage); printf("Input: %ld Output: %ld\n", usage.ru_inblock, usage.ru_oublock); exit(0); }
Note: This is an operating systems class semester project, attempting to improve application launch speed via multithreaded manual system calls, running on Linux MintMate in VirtualBox (host: Windows 10).
内容的提问来源于stack exchange,提问作者Leonard

