Java实现类似R dplyr的filter功能 匹配两个CSV的相同personID行
实现方案
核心思路是提前把活动表全量加载到内存并按personID分组,避免遍历属性表时反复读取活动表产生额外IO开销,类比R中先读取活动表为data.frame再多次filter的操作,效率更高。
步骤1:预加载活动表生成分组索引
先实现方法读取活动表,构造Map<Integer, List<String[]>>结构,key为personID,value为该ID对应的所有活动行拆分后的字段数组:
import java.io.BufferedReader; import java.io.FileReader; import java.util.ArrayList; import java.util.HashMap; import java.util.List; import java.util.Map; // 预加载活动表方法 private Map<Integer, List<String[]>> loadActivityMap(String activityFilePath) throws Exception { Map<Integer, List<String[]>> activityMap = new HashMap<>(); // 用try-with-resources语法自动关闭流,避免资源泄漏 try (BufferedReader activityReader = new BufferedReader(new FileReader(activityFilePath))) { String line; while ((line = activityReader.readLine()) != null) { String[] fields = line.split(","); Integer personId = Integer.parseInt(fields[0]); // 如果是该ID的第一条记录,自动初始化存储列表 activityMap.computeIfAbsent(personId, k -> new ArrayList<>()).add(fields); } } return activityMap; }
步骤2:改造原有逻辑,直接匹配活动记录
在原有遍历属性表的代码前,先调用上面的方法加载活动分组,之后直接取对应ID的活动列表即可:
// 先加载活动分组索引(只需执行一次) Map<Integer, List<String[]>> activityMap = loadActivityMap("替换为你的活动表文件路径"); // 用try-with-resources改造属性表读取逻辑,自动关闭流 try (BufferedReader attributeReader = new BufferedReader(new FileReader(attributesFile))) { String agent; while ((agent = attributeReader.readLine()) != null) { String[] attributeSpilted = agent.split(","); int attributeAgentID = Integer.parseInt(attributeSpilted[0]); // Set attributes for agents Person person = populationFactory.createPerson(Id.createPersonId(attributeAgentID)); // 等价于R dplyr的 filter(activity_df, personID == attributeAgentID) List<String[]> matchedActivities = activityMap.get(attributeAgentID); if (matchedActivities != null) { // 避免该ID无活动记录时出现空指针 for (String[] activity : matchedActivities) { // 在这里编写你的活动行处理逻辑 // 例如 activity[6] 为活动类型、activity[7] 为活动时长 } } population.addPerson(person); } }
补充说明
如果你的活动表文件特别大(超过JVM可分配内存上限),才需要考虑逐行扫描活动表的方案,常规业务场景下上述预加载分组的方案性能最优。
内容的提问来源于stack exchange,提问作者Jingjun
相关产品推荐
相关产品推荐

