You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何简化基于Java Stream的Item列表分类过滤与去重逻辑

Simplify Stream Processing for Item Deduplication & Category Filtering

Hey there! Let's sort this out cleanly—first, let's spot a small issue in your original code: it's accidentally deduplicating T-category items too, but your requirement says we need to keep all T items, no matter if their titles repeat.

Here's how we can fix that and simplify everything into a single stream pipeline (no splitting into two streams and merging results):

Approach 1: Using a HashSet for Tracking K-Titles (Readable & Efficient)

public List<Item> findSuggestions(String req) {
    List<Item> items = service.findSuggestions(req);
    Set<String> seenKeywordTitles = new HashSet<>();

    return items.stream()
        .filter(item -> {
            if ("T".equals(item.getCategory())) {
                // Keep every T-category item, no questions asked
                return true;
            } else if ("K".equals(item.getCategory())) {
                // For K-category items, only keep the first occurrence of each title
                return seenKeywordTitles.add(item.getTitle());
            }
            // If there are other categories, filter them out (adjust as needed)
            return false;
        })
        .collect(Collectors.toList());
}

How this works:

  • We use a HashSet to track which K-item titles we've already added to the result. The add() method returns true if the title wasn't in the set yet (so we keep the item), and false if it was (so we filter it out).
  • T-category items are always kept, exactly as your requirement specifies.
  • This is a single pass through the stream—no extra collection merges needed.

Approach 2: Functional Style with Collectors.toMap

If you prefer a more functional approach without an external set, you can use Collectors.toMap to handle deduplication directly:

public List<Item> findSuggestions(String req) {
    List<Item> items = service.findSuggestions(req);

    return new ArrayList<>(items.stream()
        .collect(Collectors.toMap(
            // For T-items: use unique ID as key to preserve all entries
            // For K-items: use title as key to deduplicate
            item -> "T".equals(item.getCategory()) ? item.getId() : item.getTitle(),
            Function.identity(),
            // When keys collide (only possible for K-items), keep the first occurrence
            (existingItem, newItem) -> existingItem
        ))
        .values());
}

How this works:

  • The toMap collector creates entries where:
    • T-items use their unique id as the key (so every T-item gets its own entry, no duplication).
    • K-items use their title as the key (so duplicate titles will overwrite each other).
  • The merge function ensures we keep the first occurrence of any duplicate K-item title.
  • We convert the map's values to an ArrayList to get our final list.

Either of these approaches will clean up your code and meet your requirements perfectly!

内容的提问来源于stack exchange,提问作者César Castro Aroche

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 20:37:34