使用System.Reflection.Metadata的程序重复运行性能提升的Windows机制问询
问题背景
有一款命令行工具程序,功能是接收目录路径作为输入,递归扫描其中扩展名为.dll的文件,并验证这些文件是否为有效的.NET程序集(使用System.Reflection.Metadata NuGet包),最终输出所有有效.NET程序集的名称和版本到控制台。
程序功能正常,但存在明显性能波动:首次扫描包含数千个文件的目录时,耗时长达数十秒,磁盘负载极高;但同一目录再次运行时,执行速度极快,仅需数百毫秒。已排除磁盘休眠因素,且程序本身未实现任何缓存逻辑,执行完成后即终止。
现需解答两个问题:
- Windows系统中存在何种机制,使得
var dllFilePaths = new DirectoryInfo(args[0]).GetFiles("*.dll", SearchOption.AllDirectories)代码重复运行时速度更快? - 调用
System.Reflection.Metadata读取器的代码重复运行时速度更快的原因是什么?
程序代码
namespace AssemblyFinder { using System; using System.Collections.Generic; using System.Diagnostics; using System.IO; using System.Linq; using System.Reflection.Metadata; using System.Reflection.PortableExecutable; public static class Program { public static int Main(string[] args) { if (args.Length != 1) { Console.Error.WriteLine("Please specify the path of the directory to be scanned for .NET assemblies."); return 1; } if (!Directory.Exists(args[0])) { Console.Error.WriteLine($"The specified argument \"{args[0]}\" is not an existing directory."); return 2; } Stopwatch stopWatch = new Stopwatch(); stopWatch.Start(); var assemblies = new List<AssemblyInfo>(); var dllFilePaths = new DirectoryInfo(args[0]) .GetFiles("*.dll", SearchOption.AllDirectories) .Select(dllFileInfo => dllFileInfo.FullName) .ToList(); stopWatch.Stop(); Console.WriteLine($"Found {dllFilePaths.Count} dlls in {(long)stopWatch.Elapsed.TotalMilliseconds} milliseconds"); stopWatch.Restart(); foreach (var dllFilePath in dllFilePaths) { if (IsAssembly(dllFilePath, out var assemblyInfo)) { assemblies.Add(assemblyInfo); } } stopWatch.Stop(); foreach (var assembly in assemblies) { Console.WriteLine($"{assembly.FilePath} -> {assembly.AssemblyName}, {assembly.AssemblyVersion}"); } Console.WriteLine($"Found {dllFilePaths.Count} dlls, identified {assemblies.Count} assemblies in {(long)stopWatch.Elapsed.TotalMilliseconds} milliseconds"); return 0; } private static bool IsAssembly(string filePath, out AssemblyInfo assemblyInfo) { using (var fileStream = new FileStream(filePath, FileMode.Open, FileAccess.Read, FileShare.ReadWrite)) using (var peReader = new PEReader(fileStream)) { if (!peReader.PEHeaders.IsDll || !peReader.HasMetadata) { assemblyInfo = null; return false; } var metadataReader = peReader.GetMetadataReader(); if (!metadataReader.IsAssembly) { assemblyInfo = null; return false; } var assemblyDefinition = metadataReader.GetAssemblyDefinition(); assemblyInfo = new AssemblyInfo() { AssemblyName = assemblyDefinition.Name.IsNil ? null : metadataReader.GetString(assemblyDefinition.Name), AssemblyVersion = assemblyDefinition.Version, FilePath = filePath, }; return true; } } } public class AssemblyInfo { public string AssemblyName { get; set; } public Version AssemblyVersion { get; set; } public string FilePath { get; set; } } }
解答
问题1:目录扫描重复运行更快的原因
这是Windows的**文件系统缓存(File System Cache,又称页缓存)**机制导致的:
- 首次扫描时,系统需要从物理磁盘读取目录结构、文件名、路径等元数据,涉及磁盘I/O操作,速度慢且磁盘负载高。
- 首次扫描完成后,这些目录和文件的元数据会被Windows缓存到物理内存中。再次运行时,
DirectoryInfo.GetFiles会直接从内存缓存中获取数据,无需访问物理磁盘,因此速度大幅提升。
问题2:System.Reflection.Metadata读取重复运行更快的原因
核心依然是Windows文件系统缓存,但针对的是DLL文件的内容:
- 首次读取DLL文件时,需要从磁盘读取文件内容到内存;读取完成后,Windows会将这些文件数据缓存到系统缓存中。
- 再次运行时,
FileStream打开文件并读取数据时,系统会直接从内存缓存中返回数据,无需执行磁盘I/O操作。PEReader和MetadataReader处理的是内存中的数据,自然速度更快。
另外,System.Reflection.Metadata本身是高效的元数据读取库,仅读取PE文件中必要的元数据部分,但首次运行的瓶颈始终是磁盘I/O,第二次运行则完全规避了这个瓶颈。
内容的提问来源于stack exchange,提问作者ackh
相关产品推荐
相关产品推荐

