如何在调用AddTest向List添加KeyValuePair数据前去除重复项
解决方案
快速实现方案(适合数据量较小的场景)
核心逻辑:byte[]属于引用类型,默认的相等判断仅对比内存地址而非实际内容,所以需要借助SequenceEqual方法对比字节数组的实际存储值,同时匹配字符串值判断是否为重复项,需要引入System.Linq命名空间。
using System.Linq; private static List<KeyValuePair<byte[], string>> list = new List<KeyValuePair<byte[], string>>(); public static void AddTest(byte[] myarray, string test) { // 判定重复规则:字符串值相同 + 字节数组内容完全一致 bool isDuplicate = list.Any(item => item.Value == test && item.Key.Length == myarray.Length && item.Key.SequenceEqual(myarray)); if (!isDuplicate) { list.Add(new KeyValuePair<byte[], string>(myarray, test)); } }
高性能实现方案(适合数据量大、频繁去重的场景)
如果数据量较大,每次遍历List的O(n)时间复杂度性能较差,可以自定义相等比较器,配合HashSet实现O(1)时间复杂度的去重判断:
// 自定义KeyValuePair<byte[], string>相等比较器 public class KvpComparer : IEqualityComparer<KeyValuePair<byte[], string>> { public bool Equals(KeyValuePair<byte[], string> x, KeyValuePair<byte[], string> y) { return x.Value == y.Value && x.Key.Length == y.Key.Length && x.Key.SequenceEqual(y.Key); } public int GetHashCode(KeyValuePair<byte[], string> obj) { int hash = obj.Value.GetHashCode(); foreach (byte b in obj.Key) { hash = HashCode.Combine(hash, b); } return hash; } }
使用HashSet替换List,自动处理去重逻辑:
private static HashSet<KeyValuePair<byte[], string>> dataSet = new HashSet<KeyValuePair<byte[], string>>(new KvpComparer()); public static void AddTest(byte[] myarray, string test) { // Add方法自动判断重复,重复项会直接返回false不执行添加 dataSet.Add(new KeyValuePair<byte[], string>(myarray, test)); }
注意事项
- 如果业务中存在
byte[]为null的场景,需要提前添加空值判断,避免调用SequenceEqual时抛出空引用异常 - 字节数组对比前优先判断长度的逻辑,可大幅降低长数组对比的性能开销,建议保留
内容的提问来源于stack exchange,提问作者Andrey Vasiliykov
相关产品推荐
相关产品推荐

