多字段去重技术问询:C# AnswerList模型数据列表处理
Hey there, let's work through how to deduplicate your AnswerList collection based on multiple fields. I've got a few solid approaches you can use depending on your scenario:
This is the quickest approach for small to medium datasets. We group the collection by the fields you want to use for deduplication, then pick the first item from each group. Anonymous types automatically handle equality checks for nullable fields, which is perfect here since your model has int? and nullable strings.
// Define the fields you want to use for deduplication (adjust based on your needs) var deduplicatedAnswers = answerListCollection .GroupBy(item => new { item.Schedule_ID, item.SubItemID, item.ItemName, item.SubItemName }) .Select(group => group.First()) // Pick the first entry in each unique group .ToList();
If you need to prioritize a specific entry (like the most recent one by Load_Date), you can sort the group first:
var deduplicatedLatestAnswers = answerListCollection .GroupBy(item => new { item.Schedule_ID, item.SubItemID }) .Select(group => group.OrderByDescending(x => x.Load_Date).First()) .ToList();
If you need to reuse the same deduplication logic across multiple parts of your code, or if you have complex equality rules (like case-insensitive string comparisons), create a custom equality comparer and use it with Distinct().
First, implement the comparer:
public class AnswerListDeduplicationComparer : IEqualityComparer<AnswerList> { public bool Equals(AnswerList x, AnswerList y) { if (x == null && y == null) return true; if (x == null || y == null) return false; // Define your equality rules here return x.Schedule_ID == y.Schedule_ID && x.SubItemID == y.SubItemID && string.Equals(x.ItemName, y.ItemName, StringComparison.OrdinalIgnoreCase) && string.Equals(x.SubItemName, y.SubItemName, StringComparison.OrdinalIgnoreCase); } public int GetHashCode(AnswerList obj) { if (obj == null) return 0; // Combine hash codes of your target fields (handle nulls safely) int hashSchedule = obj.Schedule_ID.GetHashCode(); int hashSubItem = obj.SubItemID.GetHashCode(); int hashItemName = obj.ItemName?.GetHashCode(StringComparison.OrdinalIgnoreCase) ?? 0; int hashSubItemName = obj.SubItemName?.GetHashCode(StringComparison.OrdinalIgnoreCase) ?? 0; return hashSchedule ^ hashSubItem ^ hashItemName ^ hashSubItemName; } }
Then use it in your code:
var deduplicatedAnswers = answerListCollection .Distinct(new AnswerListDeduplicationComparer()) .ToList();
If your AnswerList data comes directly from a database via Entity Framework Core 3.0+, use DistinctBy() to push the deduplication logic to the database. This is way more efficient for large datasets since it avoids loading duplicate records into memory.
// Example with async query var deduplicatedAnswers = await dbContext.AnswerLists .DistinctBy(item => new { item.Schedule_ID, item.SubItemID, item.ItemName }) .ToListAsync();
Key Notes to Keep in Mind
- Nullable Fields: All the above approaches handle nullable value types (like
int?) and nullable strings correctly, but double-check your equality logic if you have edge cases (e.g., treatingnulland empty strings as equal). - Performance: For large datasets, always prefer database-level deduplication (EF Core's
DistinctBy) over in-memory processing. - Customization: Adjust the fields in each approach to match your actual deduplication requirements—you don't have to use all fields, just the ones that define a "unique" entry for your use case.
内容的提问来源于stack exchange,提问作者Tanwer

