自定义Where()实现为何比LINQ Where()性能更低?
MyWhere Underperforms LINQ's Built-in Where Great question! Let's dive into exactly why your custom MyWhere implementation isn't as fast as the official LINQ Where, and how their performance stacks up in different scenarios.
Key Issues in Your Implementation
Your code works, but it has a few design choices that hurt performance and deviate from LINQ's behavior:
Unnecessary Empty Sequence Check
Your code immediately callsenumerator.MoveNext()to check if the sequence is empty, and throws an exception if it is. This is a double whammy:- It adds an extra, unneeded
MoveNext()call even when the sequence has elements. - It changes core behavior: LINQ's
Wheredoesn't throw on empty sequences—it simply returns an empty enumerator. Throwing exceptions is also inherently expensive, so this adds overhead when dealing with empty inputs.
- It adds an extra, unneeded
No Specialized Optimizations for Common Collections
The built-inWhereis optimized for specificIEnumerable<T>implementations like arrays,List<T>, orHashSet<T>. For example, if you pass an array, LINQ uses direct index access instead of an enumerator, which avoids the overhead of enumerator state management and method calls. Your implementation uses the generic enumerator for all sources, missing out on these performance gains.Suboptimal Enumerator Flow
Yourdo-whileloop starts with a pre-fetched element (from the initialMoveNext()). While this works, it doesn't align with LINQ's lazy execution model perfectly. LINQ only fetches elements when the consumer callsMoveNext()on the resulting enumerator, which avoids pre-fetching elements that might never be used (e.g., if the consumer stops enumerating early).
Performance Comparison
Let's break down how these differences play out in real scenarios:
1. Empty Sequences
- Your
MyWhere: Immediately throws anInvalidOperationException, which involves exception stack trace generation and cleanup—this is slow. - LINQ
Where: Returns an empty enumerator with almost zero overhead, no exceptions thrown.
2. Small/Medium Sequences
For sequences where you might only consume a few elements (e.g., using First() or Take(5)):
- Your implementation pre-fetches the first element upfront, which wastes cycles if that element doesn't match the predicate and you stop enumerating early.
- LINQ fetches elements on-demand, so it only processes exactly what's needed.
3. Large Sequences (e.g., Arrays/Lists)
Here's where the gap widens:
- LINQ switches to optimized index-based iteration for collections like arrays, which is significantly faster than using an enumerator (fewer method calls, no state tracking).
- Your implementation uses the generic enumerator regardless of the source type, which adds consistent overhead per element processed.
Example Performance Test
If you run a quick test with a large array:
var largeArray = Enumerable.Range(0, 1_000_000).ToArray(); // LINQ Where var linqStopwatch = Stopwatch.StartNew(); var linqCount = largeArray.Where(x => x % 2 == 0).Count(); linqStopwatch.Stop(); // Custom MyWhere var customStopwatch = Stopwatch.StartNew(); var customCount = largeArray.MyWhere(x => x % 2 == 0).Count(); customStopwatch.Stop(); Console.WriteLine($"LINQ Where: {linqStopwatch.ElapsedMilliseconds}ms"); Console.WriteLine($"Custom MyWhere: {customStopwatch.ElapsedMilliseconds}ms");
You'll typically see LINQ finish in half the time (or less) compared to your implementation, thanks to array-specific optimizations.
Fixing Your Implementation
To align closer to LINQ's performance and behavior, you'd want to:
- Remove the empty sequence check (let the consumer handle empty results if needed).
- Add null checks for
sourceandpredicate(throwArgumentNullExceptionupfront, like LINQ does). - Consider adding overloads for specific collection types to leverage optimized iteration.
内容的提问来源于stack exchange,提问作者HelloWorld

