You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java 8:基于关联字段高效过滤两个不同Schema的JSON数组

Got it, let's tackle this efficiently—when dealing with huge datasets, naive nested iteration is a non-starter because it’s O(nm)* time complexity, which will crawl to a stop with large arrays. Here’s a streamlined approach using Java 8 Streams and hash-based lookups to get your filtered results in O(n + m) time, which is perfect for big data:

Efficient Filtering of Large JSON Arrays in Java 8

Step 1: Parse JSON to Type-Safe Objects (or Work Directly with JSON Nodes)

First, let’s define simple POJOs to map your JSON structures—this makes the code cleaner and avoids runtime type errors. If you prefer skipping POJOs, I’ll include a direct JSON node approach too.

POJO Definitions

// Maps to elements in JSON Array A
class User {
    private String name;
    private long uid;

    // Constructor, getters, setters (use Lombok to cut down boilerplate if you want)
    public User(String name, long uid) {
        this.name = name;
        this.uid = uid;
    }

    public long getUid() { return uid; }
    public String getName() { return name; }
}

// Maps to elements in JSON Array B
class Group {
    private String group;
    private long admin;

    public Group(String group, long admin) {
        this.group = group;
        this.admin = admin;
    }

    public long getAdmin() { return admin; }
    public String getGroup() { return group; }
}

Step 2: Build a Hash Set for O(1) Lookups

The core of this efficiency is converting one array’s matching field into a HashSet—lookups in a hash set are constant time, which eliminates the need for slow nested loops.

Example 1: Filter Groups Where Admin Exists in User UIDs

// Assume you've parsed JSON Array A into a List<User> usersList
Set<Long> userUidSet = usersList.parallelStream() // Use parallel stream for large datasets
    .map(User::getUid)
    .collect(Collectors.toSet());

// Filter JSON Array B to keep only groups whose admin is in the User UID set
List<Group> filteredGroups = groupsList.parallelStream()
    .filter(group -> userUidSet.contains(group.getAdmin()))
    .collect(Collectors.toList());

Example 2: Filter Users Who Are Group Admins

If you need the reverse (users who act as admins in any group), just flip the logic:

Set<Long> adminIdSet = groupsList.parallelStream()
    .map(Group::getAdmin)
    .collect(Collectors.toSet());

List<User> filteredUsers = usersList.parallelStream()
    .filter(user -> adminIdSet.contains(user.getUid()))
    .collect(Collectors.toList());

Step 3: Working Directly with JSON Nodes (No POJOs)

If you don’t want to create POJOs, use Jackson’s JsonNode to process the JSON directly:

import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.databind.JsonNode;
import java.util.stream.StreamSupport;

ObjectMapper mapper = new ObjectMapper();
JsonNode arrayA = mapper.readTree(jsonAString); // Your JSON Array A as a string
JsonNode arrayB = mapper.readTree(jsonBString); // Your JSON Array B as a string

// Extract UIDs from Array A into a hash set
Set<Long> uidSet = StreamSupport.stream(arrayA.spliterator(), true)
    .map(node -> node.get("uid").asLong())
    .collect(Collectors.toSet());

// Filter Array B to get matching groups
List<JsonNode> filteredGroups = StreamSupport.stream(arrayB.spliterator(), true)
    .filter(node -> uidSet.contains(node.get("admin").asLong()))
    .collect(Collectors.toList());

Why This Works

  • Time Complexity: O(n + m) where n is the size of Array A and m is the size of Array B. Building the hash set takes O(n) time, and filtering the second array takes O(m) time (each lookup is O(1)).
  • Parallel Streams: Using parallelStream() leverages multiple threads to process large datasets faster—ideal when dealing with massive arrays.
  • Memory Efficiency: Storing just the IDs in a hash set uses far less memory than keeping entire objects, which is crucial for large-scale data.

内容的提问来源于stack exchange,提问作者Yash Krishnan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:33:06