C# Linq性能优化求助:List与Where子句相关代码优化
Hey there! Let's dive into fixing that slow code of yours. The biggest culprit here is the repeated linear search through your codes collection every time the loop runs—if codes is large, this adds up to a lot of unnecessary work. Let's break down the best optimizations step by step:
1. Replace Linear Search with a Prebuilt Lookup Dictionary
Instead of using Where(x => x.Value == code && x.Active).FirstOrDefault() on every iteration, prebuild a dictionary of active codes once before the loop. This turns your O(n) search into an O(1) lookup, which is night-and-day faster for large datasets.
// Do this ONCE, outside your foreach loop var activeCodeLookup = codes .Where(x => x.Active) .ToDictionary(x => x.Value); // Maps code values to their corresponding objects // Then inside your loop: if (!activeCodeLookup.TryGetValue(code, out var cond)) { errors.Add(record); continue; }
Using TryGetValue is also more efficient than FirstOrDefault because it avoids enumerating the collection entirely once the key is found (or not found).
2. Preallocate List Capacities
If you have a rough idea of how many records might end up in errors or the lists inside tempResults, preallocating their capacity prevents expensive memory reallocations as the lists grow. For example:
// If diag has 1000 records, errors can't have more than that errors = new List<Record>(diag.Count); // Similarly, if you know how many unique active codes there are, preallocate tempResults or its inner lists
3. Speed Up String Splitting
Your current Regex.Split call is relatively slow for repeated use. If your line splitting logic is just splitting on one or more whitespace characters, swap it out for string.Split—it's much faster for simple whitespace splits:
// Replace this: // var code = Convert.ToInt16(Regex.Split(record.Line, @"\s{1,}")[4], 16); // With this: var parts = record.Line.Split(new[] {' ', '\t'}, StringSplitOptions.RemoveEmptyEntries); var code = Convert.ToInt16(parts[4], 16);
If you absolutely need regex for more complex splitting, precompile the regex outside the loop to avoid recompiling it every iteration:
// Outside the loop var lineSplitter = new Regex(@"\s{1,}", RegexOptions.Compiled); // Inside the loop var parts = lineSplitter.Split(record.Line); var code = Convert.ToInt16(parts[4], 16);
4. Batch Processing with LINQ (Optional)
If readability is a priority and your dataset isn't extremely large, you can refactor the entire logic into a LINQ pipeline. Just make sure to avoid redundant calculations by projecting to an intermediate type first:
var activeCodeLookup = codes.Where(x => x.Active).ToDictionary(x => x.Value); // First, project each record to its code and matching condition var processedRecords = diag.Select(record => new { Record = record, Condition = activeCodeLookup.TryGetValue( Convert.ToInt16(record.Line.Split(new[] {' ', '\t'}, StringSplitOptions.RemoveEmptyEntries)[4], 16), out var cond ) ? cond : null }); // Split into results and errors var tempResults = processedRecords .Where(x => x.Condition != null) .GroupBy(x => x.Condition) .ToDictionary(g => g.Key, g => g.Select(x => x.Record).ToList()); errors = processedRecords .Where(x => x.Condition == null) .Select(x => x.Record) .ToList();
This keeps everything concise, but keep in mind that for super large datasets, a manual loop might still be slightly faster (due to less overhead from LINQ's delegates).
内容的提问来源于stack exchange,提问作者Filip Procházka

