如何从Java多值映射(Map<Enum, List<String>>)中移除‘纵向重复’值?
Removing Vertical Duplicates from Map<Enum, List>
Alright, let's tackle this problem step by step. You need to eliminate "vertical duplicates" from your Map<Enum, List<String>>—meaning any column (where a column is the collection of values at the same index across all enum-keyed lists) that's an exact copy of a column we've already kept should be removed. Since all your lists are the same length, this approach will work smoothly.
Core Approach
- Represent Columns as Comparable Objects: For each column index, create a list containing the value from every enum's list at that index. This list will act as a unique identifier for the column.
- Track Seen Columns: Use a set to keep track of which column identifiers we've already encountered. If a column's identifier is already in the set, it's a duplicate and we skip it.
- Reconstruct the Map: Once we have the indices of all non-duplicate columns, build a new map where each enum's list only includes values from those kept indices.
Full Java Implementation
import java.util.*; import java.util.stream.Collectors; // Replace this with your actual Enum type enum DataField { USER_ID, USER_NAME, ACCOUNT_STATUS, LAST_LOGIN } public class VerticalDuplicateCleaner { public static Map<DataField, List<String>> removeVerticalDuplicates(Map<DataField, List<String>> inputMap) { // Handle edge cases: empty or null input if (inputMap == null || inputMap.isEmpty()) { return Collections.emptyMap(); } // Get all enum keys and confirm column count (all lists are same length) List<DataField> enumKeys = new ArrayList<>(inputMap.keySet()); int totalColumns = inputMap.values().iterator().next().size(); // Track columns we've already seen (using column value lists as identifiers) Set<List<String>> seenColumns = new HashSet<>(); // Store indices of columns we want to keep List<Integer> keptColumnIndices = new ArrayList<>(); for (int colIndex = 0; colIndex < totalColumns; colIndex++) { // Build the identifier for the current column List<String> currentColumn = enumKeys.stream() .map(key -> inputMap.get(key).get(colIndex)) .collect(Collectors.toList()); // If this column hasn't been seen before, keep its index if (seenColumns.add(currentColumn)) { keptColumnIndices.add(colIndex); } } // Reconstruct the map with only non-duplicate columns return enumKeys.stream() .collect(Collectors.toMap( key -> key, key -> keptColumnIndices.stream() .map(colIndex -> inputMap.get(key).get(colIndex)) .collect(Collectors.toList()) )); } // Test the implementation public static void main(String[] args) { Map<DataField, List<String>> testData = new HashMap<>(); testData.put(DataField.USER_ID, Arrays.asList("101", "102", "101", "102")); testData.put(DataField.USER_NAME, Arrays.asList("Alice", "Bob", "Alice", "Bob")); testData.put(DataField.ACCOUNT_STATUS, Arrays.asList("ACTIVE", "ACTIVE", "ACTIVE", "ACTIVE")); testData.put(DataField.LAST_LOGIN, Arrays.asList("2024-01-01", "2024-01-02", "2024-01-01", "2024-01-02")); Map<DataField, List<String>> cleanedData = removeVerticalDuplicates(testData); // Output will show columns 0, 1, 2 are kept (column 3 duplicates column 1) cleanedData.forEach((key, values) -> System.out.println(key + ": " + values)); } }
Key Details & Notes
- Handling Nulls: This code works even if your lists contain
nullvalues, since Java'sList.equals()andhashCode()correctly handlenullelements. - Performance: The time complexity is O(C * K), where C is the number of columns and K is the number of enum keys. This is efficient for most use cases, as we're only iterating through each column's values once.
- Optimization for Large Datasets: If you're working with very large datasets, consider using an immutable list (like Guava's
ImmutableListor Java 9+'sList.of()) for the column identifiers. These cache their hash codes, which speeds up lookups in theHashSet. - Preserving Order: We iterate columns from left to right, so the first occurrence of each unique column is kept, and subsequent duplicates are removed—this preserves the original order of your non-duplicate columns.
内容的提问来源于stack exchange,提问作者AlexMTMorgan
相关产品推荐
相关产品推荐

