如何使用OpenCsv库在Java中实现CSV分批读取?
如何使用OpenCsv库在Java中实现CSV分批读取?
嘿,我来帮你搞定OpenCsv的分批读取问题!你现在的代码是一次性把整个CSV加载到内存里转成DTO,要是遇到超大CSV的话很容易爆内存,分批处理确实是更稳妥的方案。OpenCsv本身没有直接的“分批读取”API,但我们可以用迭代器手动控制批次大小,实现边读边处理的逻辑。
核心思路
不用直接调用parse()一次性获取所有数据,而是通过CsvToBean的迭代器逐行读取DTO对象,凑够指定批次数量后就处理该批次,循环直到所有数据处理完成,最后别忘了处理剩余的不足一个批次的数据。
代码实现示例
下面是基于你原有代码修改的分批处理版本,包含批次处理的核心逻辑:
public void processCsvInBatches(MultipartFile customerCsvMultipartFile, int batchSize) throws IOException, CsvValidationException { // 先执行你原有的CSV文件校验逻辑 validateCsvFile(customerCsvMultipartFile); try (InputStreamReader customerReader = new InputStreamReader(customerCsvMultipartFile.getInputStream())) { // 初始化CsvToBean,和你原来的配置一致 CsvToBean<CustomerCsvDto> customerCsvToBean = new CsvToBeanBuilder<CustomerCsvDto>(customerReader) .withType(CustomerCsvDto.class) .withIgnoreLeadingWhiteSpace(true) .build(); // 获取DTO迭代器,逐行读取数据 Iterator<CustomerCsvDto> dtoIterator = customerCsvToBean.iterator(); // 初始化批次列表,指定容量提升性能 List<CustomerCsvDto> currentBatch = new ArrayList<>(batchSize); while (dtoIterator.hasNext()) { currentBatch.add(dtoIterator.next()); // 当批次达到指定大小,触发处理逻辑 if (currentBatch.size() == batchSize) { processSingleBatch(currentBatch); currentBatch.clear(); // 清空批次,准备下一批 } } // 处理最后一批不足batchSize的数据 if (!currentBatch.isEmpty()) { processSingleBatch(currentBatch); } } catch (Exception e) { throw new RuntimeException("分批处理CSV时出错: " + e.getMessage(), e); } } // 这里替换成你自己的业务处理逻辑,比如数据校验、入库等 private void processSingleBatch(List<CustomerCsvDto> batch) { // 示例:打印批次信息,实际业务中换成你的逻辑 System.out.println("正在处理一批数据,共" + batch.size() + "条记录"); // 比如:customerService.saveBatch(batch); }
关键细节说明
- 迭代器的使用:用
customerCsvToBean.iterator()替代parse(),避免一次性加载所有数据到内存,尤其适合超大CSV文件。 - 批次容量优化:初始化
currentBatch时指定容量为batchSize,减少列表扩容的性能开销。 - 资源管理:用try-with-resources包裹
InputStreamReader,确保IO资源被自动关闭,避免泄漏。 - 剩余数据处理:循环结束后要检查批次是否为空,处理最后一批不足指定大小的数据,避免遗漏。
可选:返回所有批次数据
如果你需要把所有批次的数据收集起来返回(而不是边读边处理),可以修改方法返回List<List<CustomerCsvDto>>,代码如下:
public List<List<CustomerCsvDto>> readCsvInBatches(MultipartFile customerCsvMultipartFile, int batchSize) throws IOException, CsvValidationException { validateCsvFile(customerCsvMultipartFile); List<List<CustomerCsvDto>> allBatches = new ArrayList<>(); try (InputStreamReader customerReader = new InputStreamReader(customerCsvMultipartFile.getInputStream())) { CsvToBean<CustomerCsvDto> customerCsvToBean = new CsvToBeanBuilder<CustomerCsvDto>(customerReader) .withType(CustomerCsvDto.class) .withIgnoreLeadingWhiteSpace(true) .build(); Iterator<CustomerCsvDto> dtoIterator = customerCsvToBean.iterator(); List<CustomerCsvDto> currentBatch = new ArrayList<>(batchSize); while (dtoIterator.hasNext()) { currentBatch.add(dtoIterator.next()); if (currentBatch.size() == batchSize) { // 存入新的列表,避免后续清空影响已保存的批次数据 allBatches.add(new ArrayList<>(currentBatch)); currentBatch.clear(); } } if (!currentBatch.isEmpty()) { allBatches.add(currentBatch); } return allBatches; } catch (Exception e) { throw new RuntimeException("分批读取CSV失败: " + e.getMessage(), e); } }
额外注意事项
- 批次大小选择:根据你的内存情况和业务场景调整,比如100、500或1000条,不要过大导致内存压力,也不要过小增加处理开销。
- 异常处理:如果处理批次时抛出异常,你可能需要添加回滚逻辑,或者记录错误数据,避免整个任务失败。
- 表头处理:OpenCsv的
@CsvBindByName会自动识别表头,迭代器默认会跳过表头行,如果你需要调整跳过行数,可以用withSkipLines(int)方法配置。
备注:内容来源于stack exchange,提问作者Ibrahim Alvi
相关产品推荐
相关产品推荐

