Java读取900万行CSV存入List报堆内存溢出,如何用Buffer方案替代?
问题解决方案
一、堆内存溢出原因
你当前的实现是把900万行转换后的Investidor对象全部存入ArrayList,所有对象都常驻堆内存,数据量超过JVM分配的堆内存上限就会抛出OOM错误。
二、堆内存溢出规避方案
- 取消全量加载逻辑:无需把所有行的对象都存入List,改为边读取边写入随机访问文件,内存中仅保留当前行的对象,堆内存占用始终维持在极低水平。
- 分批批量处理:如果要提升IO写入效率,可以设置批次阈值(比如每5000行作为一个批次),攒满一批后统一写入随机访问文件,写完清空当前批次的临时集合再继续读取,内存仅保留单批次的对象数据。
- 临时调优JVM堆参数:启动程序时增加JVM启动参数
-Xms4g -Xmx8g(根据服务器内存调整),但这是治标方案,数据量继续增长时仍会出现OOM。 - 优化对象内存占用:Investidor类中重复率高的字符串字段(比如状态、地区类字段)可调用
String.intern()减少重复实例,日期字段可转为long型时间戳存储,降低单对象的内存占用。
三、基于缓冲的实现方案(替换List)
你可以配合BufferedReader读文件 + 缓冲批次写入RandomAccessFile的方案,完全不需要全量List存储,改造后的代码示例如下:
改造后的读取写入方法
// 方法不再返回全量List,直接边读边写,传入RandomAccessFile对象用于写入 public void lerDadosCSV(String arquivoCSV, JProgressBar progressBar, JTextField textField, int tipo, RandomAccessFile raf) { long indice = 0; // 批次大小,可根据实际内存调整,1000-10000都可 final int BATCH_SIZE = 5000; List<Object> batchRecords = new ArrayList<>(BATCH_SIZE); numeroTotalLinhas = numeroTotalLinhas(arquivoCSV); DecimalFormat decimalFormat = new DecimalFormat("#,###"); // 替换TextFile为BufferedReader,大文件读取更稳定 try (BufferedReader br = new BufferedReader(new FileReader(arquivoCSV, StandardCharsets.UTF_8))) { String linha; while ((linha = br.readLine()) != null) { if (indice != 0) { Object record = tipo == 0 ? montaEstoque(linha.split(";")) : montaInvestidor(linha.split(";")); if (record != null) { batchRecords.add(record); } // 攒满批次就写入 if (batchRecords.size() >= BATCH_SIZE) { writeBatchToRaf(batchRecords, raf); batchRecords.clear(); } } // Swing UI更新要放到EDT线程避免卡顿 final long currentInd = indice; SwingUtilities.invokeLater(() -> { textField.setText(decimalFormat.format(currentInd)); int progress = (int) (currentInd * 100 / numeroTotalLinhas); progressBar.setValue(progress); progressBar.setString(progress + "%"); }); indice++; } // 写入最后一批不足阈值的剩余数据 if (!batchRecords.isEmpty()) { writeBatchToRaf(batchRecords, raf); } } catch (IOException e) { e.printStackTrace(); } }
批次写入随机访问文件的方法
private void writeBatchToRaf(List<Object> batch, RandomAccessFile raf) throws IOException { for (Object record : batch) { if (record instanceof Investidor) { Investidor inv = (Investidor) record; // 按你随机访问文件的格式写入即可,示例: raf.writeInt(inv.getId()); raf.writeLong(inv.getDataCadastro().getTime()); raf.writeUTF(inv.getNome().trim()); // 其余字段按需求写入 } else if (record instanceof Estoque) { // 处理Estoque类型的写入逻辑 } } }
附加优化提示
你当前montaInvestidor方法里的splitLinha[10].contains("s|S")写法有误,contains方法不支持正则匹配,会直接匹配s|S字符串,建议改为splitLinha[10].matches("[sS]")或splitLinha[10].equalsIgnoreCase("s")。
内容的提问来源于stack exchange,提问作者Matheus William
相关产品推荐
相关产品推荐

