You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用System.Reflection.Metadata的程序重复运行性能提升的Windows机制问询

问题背景

有一款命令行工具程序,功能是接收目录路径作为输入,递归扫描其中扩展名为.dll的文件,并验证这些文件是否为有效的.NET程序集(使用System.Reflection.Metadata NuGet包),最终输出所有有效.NET程序集的名称和版本到控制台。

程序功能正常,但存在明显性能波动:首次扫描包含数千个文件的目录时,耗时长达数十秒,磁盘负载极高;但同一目录再次运行时,执行速度极快,仅需数百毫秒。已排除磁盘休眠因素,且程序本身未实现任何缓存逻辑,执行完成后即终止。

现需解答两个问题:

  1. Windows系统中存在何种机制,使得var dllFilePaths = new DirectoryInfo(args[0]).GetFiles("*.dll", SearchOption.AllDirectories)代码重复运行时速度更快?
  2. 调用System.Reflection.Metadata读取器的代码重复运行时速度更快的原因是什么?

程序代码

namespace AssemblyFinder
{
   using System;
   using System.Collections.Generic;
   using System.Diagnostics;
   using System.IO;
   using System.Linq;
   using System.Reflection.Metadata;
   using System.Reflection.PortableExecutable;

   public static class Program
   {
      public static int Main(string[] args)
      {
         if (args.Length != 1)
         {
            Console.Error.WriteLine("Please specify the path of the directory to be scanned for .NET assemblies.");
            return 1;
         }

         if (!Directory.Exists(args[0]))
         {
            Console.Error.WriteLine($"The specified argument \"{args[0]}\" is not an existing directory.");
            return 2;
         }

         Stopwatch stopWatch = new Stopwatch();
         stopWatch.Start();

         var assemblies = new List<AssemblyInfo>();
         var dllFilePaths = new DirectoryInfo(args[0])
            .GetFiles("*.dll", SearchOption.AllDirectories)
            .Select(dllFileInfo => dllFileInfo.FullName)
            .ToList();

         stopWatch.Stop();
         Console.WriteLine($"Found {dllFilePaths.Count} dlls in {(long)stopWatch.Elapsed.TotalMilliseconds} milliseconds");
         stopWatch.Restart();

         foreach (var dllFilePath in dllFilePaths)
         {
            if (IsAssembly(dllFilePath, out var assemblyInfo))
            {
               assemblies.Add(assemblyInfo);
            }
         }

         stopWatch.Stop();

         foreach (var assembly in assemblies)
         {
            Console.WriteLine($"{assembly.FilePath} -> {assembly.AssemblyName}, {assembly.AssemblyVersion}");
         }

         Console.WriteLine($"Found {dllFilePaths.Count} dlls, identified {assemblies.Count} assemblies in {(long)stopWatch.Elapsed.TotalMilliseconds} milliseconds");
         return 0;
      }

      private static bool IsAssembly(string filePath, out AssemblyInfo assemblyInfo)
      {
         using (var fileStream = new FileStream(filePath, FileMode.Open, FileAccess.Read, FileShare.ReadWrite))
         using (var peReader = new PEReader(fileStream))
         {
            if (!peReader.PEHeaders.IsDll || !peReader.HasMetadata)
            {
               assemblyInfo = null;
               return false;
            }

            var metadataReader = peReader.GetMetadataReader();
            if (!metadataReader.IsAssembly)
            {
               assemblyInfo = null;
               return false;
            }

            var assemblyDefinition = metadataReader.GetAssemblyDefinition();
            assemblyInfo = new AssemblyInfo()
            {
               AssemblyName = assemblyDefinition.Name.IsNil ? null : metadataReader.GetString(assemblyDefinition.Name),
               AssemblyVersion = assemblyDefinition.Version,
               FilePath = filePath,
            };

            return true;
         }
      }
   }

   public class AssemblyInfo
   {
      public string AssemblyName
      {
         get;
         set;
      }

      public Version AssemblyVersion
      {
         get;
         set;
      }

      public string FilePath
      {
         get;
         set;
      }
   }
}

解答

问题1:目录扫描重复运行更快的原因

这是Windows的**文件系统缓存(File System Cache,又称页缓存)**机制导致的:

  • 首次扫描时,系统需要从物理磁盘读取目录结构、文件名、路径等元数据,涉及磁盘I/O操作,速度慢且磁盘负载高。
  • 首次扫描完成后,这些目录和文件的元数据会被Windows缓存到物理内存中。再次运行时,DirectoryInfo.GetFiles会直接从内存缓存中获取数据,无需访问物理磁盘,因此速度大幅提升。

问题2:System.Reflection.Metadata读取重复运行更快的原因

核心依然是Windows文件系统缓存,但针对的是DLL文件的内容:

  • 首次读取DLL文件时,需要从磁盘读取文件内容到内存;读取完成后,Windows会将这些文件数据缓存到系统缓存中。
  • 再次运行时,FileStream打开文件并读取数据时,系统会直接从内存缓存中返回数据,无需执行磁盘I/O操作。PEReader和MetadataReader处理的是内存中的数据,自然速度更快。

另外,System.Reflection.Metadata本身是高效的元数据读取库,仅读取PE文件中必要的元数据部分,但首次运行的瓶颈始终是磁盘I/O,第二次运行则完全规避了这个瓶颈。


内容的提问来源于stack exchange,提问作者ackh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 09:15:46