You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

嵌套Union与链式Union的差异探究

Great question! Let’s break down the similarities and differences between these three approaches clearly.

1. Logical Consistency (Result Content)

You’re correct that all three lines will produce a List<T> containing the same unique elements—since Union is a set operation that adheres to commutativity and associativity for sets. However, there’s one subtle detail to note:

  • The order of elements in the final List will differ between the three approaches.
    • foo preserves the order of elements as they first appear in A, then B (elements not in A), then C (elements not in A∪B), then D (elements not in A∪B∪C).
    • baz reverses this priority: elements first appear in D, then C (not in D), then B (not in D∪C), then A (not in D∪C∪B).
    • bar follows the same order as foo (A → B → C → D) because the nested Union calls don’t change the traversal order of the original collections.

So while the set of elements is identical, the sequence in the List may vary if the input collections have overlapping elements.

2. Performance Differences

The key differences lie in how the Union operations traverse collections and manage their internal tracking of unique elements (via HashSet<T>):

How Union Works Under the Hood

Each call to Enumerable.Union creates a new iterator that uses a HashSet<T> to track elements it has already seen. When you iterate over the result (e.g., with ToList()), the iterator:

  1. Traverses the first input sequence, adding each element to the HashSet and yielding it.
  2. Traverses the second input sequence, yielding elements only if they aren’t already in the HashSet.

Comparing foo and bar

  • foo: A.Union(B).Union(C).Union(D) creates three separate HashSet<T> instances (one for each Union call).
  • bar: A.Union(B.Union(C.Union(D))) also creates three HashSet<T> instances (one for each nested Union).

In terms of traversal: both approaches end up traversing each original collection (A, B, C, D) exactly once. The nested vs. chained syntax doesn’t change the total number of element visits—only the order in which the iterators are constructed. For most cases, their performance will be nearly identical.

Comparing foo and baz

These two differ in the order of collection traversal and how the HashSet is populated:

  • foo starts with A, then adds elements from B, C, D that haven’t been seen yet.
  • baz starts with D, then adds elements from C, B, A that haven’t been seen yet.

The performance gap depends on your data:

  • If D contains many elements that also exist in A, B, or C, baz will populate the HashSet with those duplicate elements early. This means subsequent traversals of C, B, and A will skip more elements, potentially saving time.
  • Conversely, if A has many duplicates in the other collections, foo will be more efficient.
  • If duplicate rates are similar across all collections, the performance difference will be negligible.

Optimization Note

If performance is critical, you can avoid creating multiple HashSet instances by manually combining the collections into a single HashSet first:

var combined = new HashSet<T>(A);
combined.UnionWith(B);
combined.UnionWith(C);
combined.UnionWith(D);
var result = combined.ToList();

This creates only one HashSet and traverses each collection exactly once, which is more efficient than any of the three chained/nested Union approaches.

内容的提问来源于stack exchange,提问作者Brondahl

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:57:24