如何优化C#中Parallel.ForEach循环的ID查询性能?
优化10万条记录遍历与存在性检查的性能问题
我来帮你拆解下性能暴跌的核心原因,再给出针对性的优化方案:
问题根源分析
你的代码里有两个致命的性能瓶颈:
- O(n)的低效存在性检查:用
List.FirstOrDefault(x => x == action)做存在性判断,每次都要遍历整个10万条的列表。10万次循环下来,总操作量是100000*100000次,时间复杂度直接飙升到O(n²),这是耗时从3秒涨到24秒的主因。 - 锁的范围过大:你把整个循环体都套在
lock里,等于让Parallel.ForEach退化成了单线程执行,完全浪费了并行的优势,还额外增加了锁的开销。
另外提一句:如果你的业务场景里是直接遍历collection集合的元素,那这个存在性检查本身是冗余的——action必然是集合里的元素。不过假设真实场景是遍历另一个ID集合、检查是否在collection中,我们就针对这个场景优化。
优化方案
1. 用HashSet替换List做存在性检查
HashSet的Contains方法是O(1)时间复杂度,比List的O(n)快几个数量级。先把原集合转换成HashSet:
HashSet<int> collectionSet = new HashSet<int>(collection);
之后用collectionSet.Contains(action)替代FirstOrDefault的检查逻辑。
2. 缩小锁的范围
只在修改共享集合students的时候加锁,不要把对象创建、存在性检查这些逻辑都放在锁里,这样才能真正发挥并行的优势。
3. 可选:用线程安全集合替代手动锁
如果不想手动处理锁,可以用ConcurrentBag<Student>代替List<Student>——它是专门为并行场景设计的线程安全集合,不需要手动加锁就能安全添加元素。
优化后的代码示例
方案一:HashSet+缩小锁范围
class Student { public int ID { get; set; } public string Name { get; set; } public string Email { get; set; } } class Program { private static object _lockObj = new object(); static void Main(string[] args) { List<int> collection = Enumerable.Range(1, 100000).ToList(); // 转换成HashSet,实现O(1)的存在性检查 HashSet<int> collectionSet = new HashSet<int>(collection); List<Student> students = new List<Student>(100000); // 建议用CPU核心数作为并行度,默认也是这个值,不用硬设为1 var options = new ParallelOptions() { MaxDegreeOfParallelism = Environment.ProcessorCount }; var sp = System.Diagnostics.Stopwatch.StartNew(); Parallel.ForEach(collection, options, action => { // 快速检查ID是否存在 if (collectionSet.Contains(action)) { // 对象创建逻辑放在锁外,避免阻塞其他线程 Student student = new Student(); student.ID = action; student.Name = "Zoyeb"; student.Email = "ShaikhZoyeb@Gmail.com"; // 仅在修改共享列表时加锁 lock (_lockObj) { students.Add(student); } Console.WriteLine(@"value of i = {0}, thread = {1}", action, Thread.CurrentThread.ManagedThreadId); } }); sp.Stop(); double elapsedSeconds = sp.ElapsedMilliseconds / 1000.0; Console.WriteLine($"耗时:{elapsedSeconds}秒"); } }
方案二:HashSet+ConcurrentBag(无需手动锁)
class Student { public int ID { get; set; } public string Name { get; set; } public string Email { get; set; } } class Program { static void Main(string[] args) { List<int> collection = Enumerable.Range(1, 100000).ToList(); HashSet<int> collectionSet = new HashSet<int>(collection); // 线程安全集合,并行场景下无需手动加锁 ConcurrentBag<Student> students = new ConcurrentBag<Student>(); var options = new ParallelOptions() { MaxDegreeOfParallelism = Environment.ProcessorCount }; var sp = System.Diagnostics.Stopwatch.StartNew(); Parallel.ForEach(collection, options, action => { if (collectionSet.Contains(action)) { Student student = new Student(); student.ID = action; student.Name = "Zoyeb"; student.Email = "ShaikhZoyeb@Gmail.com"; // 直接添加,自动处理线程安全 students.Add(student); Console.WriteLine(@"value of i = {0}, thread = {1}", action, Thread.CurrentThread.ManagedThreadId); } }); sp.Stop(); double elapsedSeconds = sp.ElapsedMilliseconds / 1000.0; Console.WriteLine($"耗时:{elapsedSeconds}秒"); } }
额外提示
- 如果你的业务场景中,遍历的元素本身就来自
collection,那存在性检查可以直接去掉,性能还能再上一个台阶。 - 不要把
MaxDegreeOfParallelism硬设为1,这样等于放弃并行,默认值Environment.ProcessorCount能充分利用CPU资源。 Console.WriteLine本身是线程安全但耗时的操作,如果不是调试需要,建议去掉,能进一步提升性能。
内容的提问来源于stack exchange,提问作者Zoyeb Shaikh
相关产品推荐
相关产品推荐

