You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Stream API转换DataEntity列表为指定JSON?选普通流还是并行流?

使用Java Stream API重构DataEntity列表转指定JSON结构的代码,及流类型选择分析

问题背景

从仓库获取一组DataEntity列表,已通过普通循环实现将其转换为指定结构的JSON,但代码不够优雅,希望用Stream API重构,同时想明确该场景下使用普通流还是并行流更合适。

示例DataEntity列表

[DataEntity(key=0a3e1588-ad59-3586-b071-d5001f5ff9a7, timestamp=2024-05-24 09:48:00.0, value=10.0),
DataEntity(key=0a3e1588-ad59-3586-b071-d5001f5ff9a7, timestamp=2024-05-24 09:46:00.0, value=11.0),
DataEntity(key=0a3e1588-ad59-3586-b071-d5001f5ff9a7, timestamp=2024-05-24 09:44:00.0, value=12.0),

DataEntity(key=b349f0ea-c810-3790-bdc3-a82ed5921a28, timestamp=2024-05-24 09:48:00.0, value=13.0),
DataEntity(key=b349f0ea-c810-3790-bdc3-a82ed5921a28, timestamp=2024-05-24 09:46:00.0, value=14.0),
DataEntity(key=b349f0ea-c810-3790-bdc3-a82ed5921a28, timestamp=2024-05-24 09:44:00.0, value=15.0),

DataEntity(key=63d29aeb-ab91-3d24-b40c-d69f2955a893, timestamp=2024-05-24 09:48:00.0, value=16.0),
DataEntity(key=63d29aeb-ab91-3d24-b40c-d69f2955a893, timestamp=2024-05-24 09:46:00.0, value=17.0),
DataEntity(key=63d29aeb-ab91-3d24-b40c-d69f2955a893, timestamp=2024-05-24 09:44:00.0, value=18.0)]

期望转换后的JSON结构

{
    "0a3e1588-ad59-3586-b071-d5001f5ff9a7": [
        {
            "timestamp" : "2024-05-24 09:48:00.0",
            "value" : 10.0
        },
        {
            "timestamp" : "2024-05-24 09:46:00.0",
            "value" : 11.0
        },
        {
            "timestamp" : "2024-05-24 09:44:00.0",
            "value" : 12.0
        }
    ],

    "b349f0ea-c810-3790-bdc3-a82ed5921a28": [
        {
            "timestamp" : "2024-05-24 09:48:00.0",
            "value" : 13.0
        },
        {
            "timestamp" : "2024-05-24 09:46:00.0",
            "value" : 14.0
        },
        {
            "timestamp" : "2024-05-24 09:44:00.0",
            "value" : 15.0
        }
    ],

    "63d29aeb-ab91-3d24-b40c-d69f2955a893": [
        {
            "timestamp" : "2024-05-24 09:48:00.0",
            "value" : 16.0
        },
        {
            "timestamp" : "2024-05-24 09:46:00.0",
            "value" : 17.0
        },
        {
            "timestamp" : "2024-05-24 09:44:00.0",
            "value" : 18.0
        }
    ]
}

现有普通循环实现代码

List<DataEntity> dataEntityList = repository.getRecordCountForPeriod();

Map response = new HashMap();
for (DataEntity data : dataEntityList) {
    if(response.containsKey(data.getKey())){

        List list = (List) response.get(data.getKey());
        Map map = new HashMap<>();
        map.put("timestamp", data.getTimestamp());
        map.put("value", data.getInstancesValue());
        list.add(map);

    }else{
        List list = new ArrayList();
        Map map = new HashMap<>();
        map.put("timestamp", data.getTimestamp());
        map.put("value", data.getInstancesValue());
        list.add(map);
        response.put(data.getKey(), list);
    }
}

Stream API重构实现

可以利用Collectors.groupingBy按key分组,结合Collectors.mapping将每个DataEntity转换为目标Map结构,代码更简洁优雅:

List<DataEntity> dataEntityList = repository.getRecordCountForPeriod();

Map<String, List<Map<String, Object>>> response = dataEntityList.stream()
    .collect(Collectors.groupingBy(
        DataEntity::getKey,
        Collectors.mapping(data -> {
            Map<String, Object> item = new HashMap<>();
            item.put("timestamp", data.getTimestamp());
            item.put("value", data.getInstancesValue());
            return item;
        }, Collectors.toList())
    ));

如果希望输出的列表保持原数据的时序顺序,可以使用有序集合维持顺序:

Map<String, List<Map<String, Object>>> response = dataEntityList.stream()
    .collect(Collectors.groupingBy(
        DataEntity::getKey,
        LinkedHashMap::new, // 保持key的插入顺序
        Collectors.mapping(data -> {
            Map<String, Object> item = new HashMap<>();
            item.put("timestamp", data.getTimestamp());
            item.put("value", data.getInstancesValue());
            return item;
        }, Collectors.toCollection(LinkedList::new)) // 保持每个key下元素的原顺序
    ));

普通流vs并行流的选择

该场景下优先使用普通流,原因如下:

  1. 操作轻量:这里的分组和对象转换都是简单的getter调用和Map创建,单线程执行的开销远小于并行流带来的线程调度、同步开销。
  2. 数据量通常不大:从仓库获取的这类时序数据,若数据量未达到数万甚至数十万级别,并行流无法体现性能优势,反而可能因为线程切换拖慢速度。
  3. 顺序一致性:并行流的分组过程可能会打乱元素的原有顺序(若未指定有序收集器),如果需要保留原数据的时序顺序,普通流更易保证。

只有当数据量极大(比如百万级以上)且每个元素的处理逻辑复杂(比如涉及IO、计算密集型操作)时,才需要考虑使用并行流,此时线程开销能被处理逻辑的耗时抵消,带来性能提升。

内容的提问来源于stack exchange,提问作者Romillion

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 14:42:05