You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Java多值映射(Map<Enum, List<String>>)中移除‘纵向重复’值?

Removing Vertical Duplicates from Map<Enum, List>

Alright, let's tackle this problem step by step. You need to eliminate "vertical duplicates" from your Map<Enum, List<String>>—meaning any column (where a column is the collection of values at the same index across all enum-keyed lists) that's an exact copy of a column we've already kept should be removed. Since all your lists are the same length, this approach will work smoothly.

Core Approach

  1. Represent Columns as Comparable Objects: For each column index, create a list containing the value from every enum's list at that index. This list will act as a unique identifier for the column.
  2. Track Seen Columns: Use a set to keep track of which column identifiers we've already encountered. If a column's identifier is already in the set, it's a duplicate and we skip it.
  3. Reconstruct the Map: Once we have the indices of all non-duplicate columns, build a new map where each enum's list only includes values from those kept indices.

Full Java Implementation

import java.util.*;
import java.util.stream.Collectors;

// Replace this with your actual Enum type
enum DataField {
    USER_ID, USER_NAME, ACCOUNT_STATUS, LAST_LOGIN
}

public class VerticalDuplicateCleaner {

    public static Map<DataField, List<String>> removeVerticalDuplicates(Map<DataField, List<String>> inputMap) {
        // Handle edge cases: empty or null input
        if (inputMap == null || inputMap.isEmpty()) {
            return Collections.emptyMap();
        }

        // Get all enum keys and confirm column count (all lists are same length)
        List<DataField> enumKeys = new ArrayList<>(inputMap.keySet());
        int totalColumns = inputMap.values().iterator().next().size();

        // Track columns we've already seen (using column value lists as identifiers)
        Set<List<String>> seenColumns = new HashSet<>();
        // Store indices of columns we want to keep
        List<Integer> keptColumnIndices = new ArrayList<>();

        for (int colIndex = 0; colIndex < totalColumns; colIndex++) {
            // Build the identifier for the current column
            List<String> currentColumn = enumKeys.stream()
                    .map(key -> inputMap.get(key).get(colIndex))
                    .collect(Collectors.toList());

            // If this column hasn't been seen before, keep its index
            if (seenColumns.add(currentColumn)) {
                keptColumnIndices.add(colIndex);
            }
        }

        // Reconstruct the map with only non-duplicate columns
        return enumKeys.stream()
                .collect(Collectors.toMap(
                        key -> key,
                        key -> keptColumnIndices.stream()
                                .map(colIndex -> inputMap.get(key).get(colIndex))
                                .collect(Collectors.toList())
                ));
    }

    // Test the implementation
    public static void main(String[] args) {
        Map<DataField, List<String>> testData = new HashMap<>();
        testData.put(DataField.USER_ID, Arrays.asList("101", "102", "101", "102"));
        testData.put(DataField.USER_NAME, Arrays.asList("Alice", "Bob", "Alice", "Bob"));
        testData.put(DataField.ACCOUNT_STATUS, Arrays.asList("ACTIVE", "ACTIVE", "ACTIVE", "ACTIVE"));
        testData.put(DataField.LAST_LOGIN, Arrays.asList("2024-01-01", "2024-01-02", "2024-01-01", "2024-01-02"));

        Map<DataField, List<String>> cleanedData = removeVerticalDuplicates(testData);
        // Output will show columns 0, 1, 2 are kept (column 3 duplicates column 1)
        cleanedData.forEach((key, values) -> System.out.println(key + ": " + values));
    }
}

Key Details & Notes

  • Handling Nulls: This code works even if your lists contain null values, since Java's List.equals() and hashCode() correctly handle null elements.
  • Performance: The time complexity is O(C * K), where C is the number of columns and K is the number of enum keys. This is efficient for most use cases, as we're only iterating through each column's values once.
  • Optimization for Large Datasets: If you're working with very large datasets, consider using an immutable list (like Guava's ImmutableList or Java 9+'s List.of()) for the column identifiers. These cache their hash codes, which speeds up lookups in the HashSet.
  • Preserving Order: We iterate columns from left to right, so the first occurrence of each unique column is kept, and subsequent duplicates are removed—this preserves the original order of your non-duplicate columns.

内容的提问来源于stack exchange,提问作者AlexMTMorgan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:06:50