You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java实现类似R dplyr的filter功能 匹配两个CSV的相同personID行

实现方案

核心思路是提前把活动表全量加载到内存并按personID分组,避免遍历属性表时反复读取活动表产生额外IO开销,类比R中先读取活动表为data.frame再多次filter的操作,效率更高。

步骤1:预加载活动表生成分组索引

先实现方法读取活动表,构造Map<Integer, List<String[]>>结构,key为personID,value为该ID对应的所有活动行拆分后的字段数组:

import java.io.BufferedReader;
import java.io.FileReader;
import java.util.ArrayList;
import java.util.HashMap;
import java.util.List;
import java.util.Map;

// 预加载活动表方法
private Map<Integer, List<String[]>> loadActivityMap(String activityFilePath) throws Exception {
    Map<Integer, List<String[]>> activityMap = new HashMap<>();
    // 用try-with-resources语法自动关闭流,避免资源泄漏
    try (BufferedReader activityReader = new BufferedReader(new FileReader(activityFilePath))) {
        String line;
        while ((line = activityReader.readLine()) != null) {
            String[] fields = line.split(",");
            Integer personId = Integer.parseInt(fields[0]);
            // 如果是该ID的第一条记录,自动初始化存储列表
            activityMap.computeIfAbsent(personId, k -> new ArrayList<>()).add(fields);
        }
    }
    return activityMap;
}

步骤2:改造原有逻辑,直接匹配活动记录

在原有遍历属性表的代码前,先调用上面的方法加载活动分组,之后直接取对应ID的活动列表即可:

// 先加载活动分组索引(只需执行一次)
Map<Integer, List<String[]>> activityMap = loadActivityMap("替换为你的活动表文件路径");

// 用try-with-resources改造属性表读取逻辑,自动关闭流
try (BufferedReader attributeReader = new BufferedReader(new FileReader(attributesFile))) {
    String agent;
    while ((agent = attributeReader.readLine()) != null) {
        String[] attributeSpilted = agent.split(",");
        int attributeAgentID = Integer.parseInt(attributeSpilted[0]);

        // Set attributes for agents
        Person person = populationFactory.createPerson(Id.createPersonId(attributeAgentID));

        // 等价于R dplyr的 filter(activity_df, personID == attributeAgentID)
        List<String[]> matchedActivities = activityMap.get(attributeAgentID);
        if (matchedActivities != null) { // 避免该ID无活动记录时出现空指针
            for (String[] activity : matchedActivities) {
                // 在这里编写你的活动行处理逻辑
                // 例如 activity[6] 为活动类型、activity[7] 为活动时长
            }
        }

        population.addPerson(person);
    }
}

补充说明

如果你的活动表文件特别大(超过JVM可分配内存上限),才需要考虑逐行扫描活动表的方案,常规业务场景下上述预加载分组的方案性能最优。

内容的提问来源于stack exchange,提问作者Jingjun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 18:15:00